Why annotate source roles?
A raw citation count cannot explain whether an answer is grounded in a company profile, primary documentation, a comparison page, a tutorial, or independent corroboration. Two answers can cite three pages each while supporting very different reader needs. Role annotation makes that difference explicit without pretending that the annotation measures authority.
Use the taxonomy only after preserving the prompt, answer surface, run timestamp, visible answer text, citation order, exact observed URL, and enough surrounding text to understand why the source appears. Apply roles to the source as used in that specific answer, not to the publisher in general.
Operational roles
entityIdentity and attributes
Names the organization, product, person, place, or service and supplies attributable facts such as official name, location, ownership, features, or availability.
definitionMeaning and category
Explains a term, concept, category, or problem so a reader can understand what it is and where its boundaries sit.
evidenceTraceable factual support
Supports a claim through data, primary documentation, a disclosed method, a dated observation, or another falsifiable record.
comparisonAlternatives and trade-offs
Supplies dimensions, criteria, limitations, pricing context, or competing approaches needed to make a fair comparison.
implementationReproducible action
Shows how to perform, configure, test, measure, or troubleshoot a workflow with enough detail to attempt it.
validationIndependent corroboration
Corroborates or challenges a claim through a separate publisher, audit, review, replication, standard, or other independent check.
actionNext-step destination
Enables a concrete next step such as visiting an official tool, booking, contacting, downloading, applying, or purchasing. It is not automatically evidence.
Inclusion and exclusion tests
| Role | Include when | Do not include merely because |
|---|---|---|
entity | The answer uses the source for attributable entity facts. | The entity name appears in a page title or navigation. |
definition | The source materially explains meaning or category boundaries. | The page repeats a short marketing tagline. |
evidence | A claim can be traced to data, documentation, method, or dated record. | The page contains numbers without provenance. |
comparison | The answer relies on explicit criteria or trade-offs across options. | Several brand names appear together. |
implementation | The source gives steps, inputs, expected outputs, or test conditions. | A call to action says “get started.” |
validation | The publisher is meaningfully independent of the claim owner and provides a check. | A company repeats its own claim on another owned page. |
action | The source is used as the destination for a user action. | The URL is clickable but not presented as a next step. |
Multi-label policy
A source may perform more than one role in the same answer. Assign every role supported by the captured context, but require a short evidence note for each label. Do not infer a role from the page type alone. An official documentation page may provide entity facts, evidence, and implementation; a review may provide comparison without meaningful validation.
- Annotate independently at the citation occurrence level.
- Record one primary role: the source's main job in the answer passage.
- Add secondary roles only when the answer visibly uses them.
- Store the supporting answer span and page span when review access permits.
- Mark ambiguous cases as
needs_review; do not force a label.
Minimum annotation record
| Field | Purpose |
|---|---|
observation_id, prompt_id, prompt_version | Connect the role to a frozen, repeatable observation. |
answer_system, surface, run_timestamp, market | State where and when the variable answer was observed. |
citation_id, citation_position, observed_url | Preserve the citation occurrence and exact raw URL. |
primary_role, secondary_roles | Store the role decision using taxonomy version 1.0. |
answer_evidence_span, page_evidence_note | Explain why each role was assigned without copying unnecessary content. |
publisher_relationship | Record owned, partner, independent, unknown, or not applicable. |
reviewer_id, review_status, taxonomy_version | Make review and later reclassification auditable. |
Synthetic annotation example
The following is fictional and demonstrates structure only. It is not a measured AI answer or a claim about any real publisher.
| Answer use | Primary | Secondary | Reason |
|---|---|---|---|
| Defines an imaginary category from a standards-style page | definition | evidence | The passage uses both the definition and a disclosed specification. |
| Lists setup steps from fictional product documentation | implementation | entity | The page supports steps and product-specific configuration facts. |
| Links to the fictional product signup page | action | None | The destination enables a next step but does not substantiate the answer claim. |
Reliability and responsible reporting
For benchmark work, have two reviewers independently annotate a stratified sample before scaling. Reconcile disagreements using the captured answer context and the inclusion rules above. Report the sample, systems, prompt-set version, collection window, reviewer process, taxonomy version, and unresolved cases. A simple agreement percentage is understandable; a chance-adjusted statistic can supplement it when the sample and label distribution justify one.
Never turn role coverage into a quality score by default
An answer can cover all seven roles and still be inaccurate, stale, biased, unsafe, or poorly sourced. Role coverage is a description of evidence responsibilities in a bounded sample. Assess factual support, independence, freshness, accessibility, and risk separately.
Version 1.0 · Published 2026-08-12 · Changes should create a new version rather than silently redefining historical labels.