Methodology
Library science has a 150-year tradition of evaluating sources. Cerulean operationalizes that tradition in a search box.
The lineage
Source evaluation predates the web. Library and information science has spent more than a century building frameworks for distinguishing kinds of sources, evaluating authority, and tracing claims back toward original evidence. The Association of College and Research Libraries publishes the ACRL Framework for Information Literacy, which is the current professional standard. The CRAAP test (Currency, Relevance, Authority, Accuracy, and Purpose) is the most widely taught checklist for source evaluation in undergraduate research instruction. The BEAM framework (Background, Exhibit, Argument, and Method) describes how sources function inside a research argument. Underneath all of them sits the foundational distinction between primary, secondary, and tertiary sources, which dates to nineteenth-century historiography.
None of this is novel. What is novel is applying it to web search at scale, and surfacing the results to the user instead of hiding them behind relevance ranking.
What web search lost
Relevance ranking collapsed source-type distinctions because it did not need them. A search engine optimizing for click-through can put a SEO listicle, a peer-reviewed journal article, and a Wikipedia entry side by side as long as all three contain the query terms. The user is expected to evaluate the sources after the click. In practice most users do not. The structural signals a research librarian would foreground (who produced this, how close to the original evidence, with what editorial process) disappeared from the result page.
This was not malicious. Relevance ranking is a tractable engineering problem. Source evaluation is a harder one. The engines that won did the tractable thing, and the harder thing waited. Cerulean is an attempt to do the harder thing and put the result back on the page, in front of the click.
Two axes
Every classified result on Cerulean gets two tags. The axes are orthogonal, and both are load-bearing.
Source role — the journalistic lens
How close is the source to the original evidence? This is the axis a journalist or researcher reaches for first. It answers whether you are looking at the record itself, an account of the record, or a synthesis of other people's accounts.
- Primary. Original materials. Court filings, datasets, government raw reports, original interviews, original research articles in scientific journals, eyewitness accounts, raw artifacts, statutes, original creative works, and official communications from the entity in question.
- Secondary. Analysis or interpretation of primary sources. Most editorial journalism, academic books and review articles, biographies, most textbooks, expert commentary, and analytical pieces.
- Tertiary. Summaries or syntheses of secondary sources. Encyclopedias, Wikipedia, dictionary entries, listicles, "what is X" guides, AI-generated summaries, and most search snippets.
Source type
What kind of entity produced it?
- Primary-source publisher / Official. Government data portals, court databases, academic journal publishers, statute repositories, and raw dataset hosts.
- Journalism. Editorial process, byline, original reporting markers.
- Academic. Research institutions, academic publishers, and peer-reviewed venues.
- Reference. Wikipedia, MDN, Stanford Encyclopedia of Philosophy, and other encyclopedic works.
- Indie. Personal blogs, neocities, github.io, IndieWeb participants, and independent publishing platforms (Medium, Substack, and the like).
- Community. Reddit, Hacker News, Stack Overflow, and other forums for collective discussion.
- Social. Social-media posts, profiles, and feeds.
- Commercial. Corporate sites, product pages, and marketing materials.
- Aggregator. Lyric sites, recipe aggregators, video hosts, and content repackagers.
- SEO farm. AI-generated content farms, listicle factories, and thin affiliate sites.
Why both
Role answers the research question. Type answers the trust question. A government site can be primary (raw dataset release), secondary (policy analysis paper), or tertiary ("about this agency" page). Journalism can be primary (original 2008 financial-crisis reporting) or secondary (2015 retrospective). Collapsing both into a single tag loses information that matters at the moment a user decides whether to click. On the result page, Cerulean leads with role, then shows type, so the evidentiary structure is the first thing you read.
How classification works
Cerulean classifies in layers, most authoritative first. None of them run an LLM at query time.
Authority registry. A curated, library-style authority file: hand-reviewed records for organizations whose classification requires deliberate judgment rather than a domain rule — advocacy groups, think tanks, official outlets, and the like. Each record carries the entity's canonical name, source type, and the editorial fields described below. When a domain matches the registry, its record wins. The registry is git-versioned and publicly auditable; it is the single source of truth for individual-entity judgment, edited by deliberate curation rather than by patching classifier rules.
Bundled index. A curated set of domain entries shipped with the deploy, holding source-type and source-role labels. Lookups are sub-millisecond. Source data includes Tranco top sites, Crossref DOI prefixes, the IndieWeb directory, public .gov and .edu lists, known academic publishers, and hand-curated additions. The index grows from real traffic, not from the benchmark.
Structural heuristics. For domains in neither the registry nor the index, structural features get read at query time: TLD class, subdomain patterns (forum and documentation hosts), URL path patterns, and document-type signals. Genuine forums stay community; personal-publishing platforms read as indie; social hosts read as social; everything unresolved falls back to "unclassified" rather than guessing.
Document-level role refinement. Role is not a function of type. Inside several types the role splits by document: an arxiv preprint or a journal's original article is primary, while that journal's review piece is secondary; a government dataset is primary, while its press release reads as the agency reporting on itself; an encyclopedia entry is tertiary, while a curated library guide is secondary. The classifier reads the URL path and structural markers to make this call instead of stamping one role per type. This refinement is what lifts role accuracy past the ceiling a type-only scheme can reach.
Tier 3: background batch. For long-tail unknowns that keep surfacing in real queries, an LLM classifies them in an offline batch, and the results promote to the bundled index on the next deploy. This is the only place an LLM touches the classification pipeline. It runs offline, on unknowns the deterministic layers could not resolve, and never on a query path.
How the benchmark works
Cerulean is benchmarked two ways, and the numbers are published here, with their limits, rather than rounded up.
URL-level accuracy against a gold set. A curated gold set of real search results carries a reviewed source role and type for every URL. The classifier is scored against it per axis. Because two source classes are genuinely contestable — whether a social-media host is a community or an aggregator, whether an independent platform is indie or commercial — those are carved into a separate, reported segment so they neither inflate nor drag the core number. The core confidence interval is the shipping gate.
Distribution-bound query benchmark. A separate suite generates queries across research, news, technical, commercial, and controversial intents, and checks whether the classifier produces a sensible role/type distribution over each query's top results, not just per-URL correctness. The query set and the prompts are git-versioned.
What the current numbers say. On the gold set, the core classifier reaches roughly 90–96% on type and the low-to-mid 90s on role after the document-level refinement, with social and indie reported separately. That headline is the ceiling on this particular gold set, not a field guarantee — see the limits below. The durable, generalizable result is the role accuracy, which holds up when the gold-seeded domain lists are stripped away; the type number leans partly on a curated index that should grow from live traffic.
Honest limits
The gold set is single-assessor. The reference labels were produced by one assessor following a documented protocol. Library-science and information-retrieval practice want multiple independent assessors with a measured agreement statistic (Cohen's Kappa or Krippendorff's alpha) before benchmark numbers carry full weight. Cerulean's gold set does not have that yet. The headline accuracy is best read as a ceiling on this set, to be re-validated against independent labels.
Source role is partially context-dependent. A 2010 historian writing about the French Revolution is secondary for the Revolution itself and primary for 21st-century historiography. Cerulean classifies at the document level using structural defaults, which is correct often enough to be useful, and wrong sometimes. The methodology cannot eliminate that.
Structural signals do not capture semantic context. A page that looks like editorial journalism by every structural measure can still be a press release with a byline glued on. A page that looks like a SEO farm can be a legitimate small publisher with weak design conventions. The classifier reports its best structural read; it does not verify intent.
The index is not exhaustive. New domains appear constantly. Long-tail entries fall through to the heuristics, and a fraction of those fall through to "unclassified." Surfacing "unclassified" honestly is better than overclaiming.
Current rollout status
Live today: source role and source type. Every classified result carries both axes. The result page leads with the source-role tag (primary, secondary, or tertiary), then the source type. Role is the journalistic lens: it tells you whether you are looking at an original record, reporting and analysis of it, or a synthesis of other people's work, before you click.
Rolling into the live path: the full classifier. The authority registry and the document-level role refinement described above are implemented and benchmarked in the classification pipeline, and are being promoted into the production query path. As that lands, the role shown on each result moves from a structural default toward the path-aware call the benchmark measures.
Rolling out next: editorial assessment. Bias direction, reliability grade, and funding independence are modeled today in the gold set and the authority registry as facets orthogonal to role and type. They will surface as result facets once the single-assessor limitation is resolved with independent review. The point of keeping them orthogonal is that two sources can share a role and type and still differ in stance — an issue-advocacy think tank and a peer-reviewed lab both produce research-shaped documents, and the stance axis is where that difference belongs, not the type taxonomy.
How to challenge a classification
The registry and index are public and versioned. They live in the project repository on GitHub at github.com/garretblanchette/Cerulean-search. Anyone can read them, audit them, and propose corrections.
To challenge a label, open a GitHub issue against the repository describing the domain, the current label, the proposed label, and the reasoning. Corrections are reviewed, batched, and folded into the next version. Egregious misclassifications get fixed faster than minor disagreements. Disagreements about edge cases are documented in the issue and left visible, even if the label does not change, so future readers can see the reasoning.
Rate-limiting and trust weighting will be added before any in-app flag UI ships, to prevent mass-flagging abuse. This page will be updated when that happens.
What a research librarian would do, in a search box.