Prerequisites
- An existing agent with at least one connector that produces data parts (for example, PubMed or Medical Coding)
- Understanding of the task lifecycle, artifacts, and parts
- Understanding of streaming responses
How inline citations work
- The orchestrator receives the user message and decides which connectors to call.
- Each connector call produces one or more data parts stored on the task. Every data part gets a stable data part ID (GID) such as
tool_data_01. - The orchestrator writes its answer text and embeds citation markers that reference the GIDs of the data parts it used.
- The server parses the markers, strips them from the text, and attaches a
citationsarray to the answer artifact’s text part metadata. - The cited data parts are emitted alongside the answer text in the same artifact, each carrying its GID in its
metadata.dataPartId.
citations metadata to link each claim back to its source data part.
Inline citations are produced by the orchestrator only, the agent runtime that serves your top-level task. Connector (expert) sub-answers do not embed citation markers themselves; their outputs are data parts that the orchestrator then cites or passes through. You only need to handle citations on the top-level task answer.
Marker format
The orchestrator uses the OpenAI citation formatting guide, with three private-use Unicode characters as delimiters:
A marker has the form:
cite: backs a claim with a source. The marker is stripped from the text and an entry is added to thecitationsmetadata.<source_id>: the GID of the cited data part, for exampletool_data_02.<locator>: an optional sub-reference into the cited part: a JSON Pointer path such as/data/5/population, or a line range such asL8-L13. Empty means the whole part is the source.
Citation metadata
Citations ride on the text part metadata of the answer artifact, not on the data parts. The text partmetadata contains a citations key: an array with one entry per citation marker, in text order.
Each entry in the citations array has the following fields:
The array is sorted by
offset.
Example: answer with citations
A PubMed-backed answer whose text part metadata looks like this:data_part_id (tool_data_02), find the sibling part whose metadata.dataPartId equals tool_data_02, then apply the locator (/results/0/findings) as a JSON Pointer to that part’s data.value.
Streaming and citations
When the answer is streamed, the orchestrator emits one text part per token delta, then a closing event that replaces the deltas with the canonical consolidated text. Thecitations metadata is attached to that final consolidated text part.
During streaming, the server also strips citation markers from each delta as it arrives, so partial markers never reach your client even when a delta splits a marker across a chunk boundary.
Render citations
The text you receive is already clean (markers stripped). Render it verbatim and use thecitations metadata to add your own citation UI:
- Read
metadata.citations. - For each entry, insert your rendered citation marker (superscript, footnote, or popover) at
offsetcharacters into the joined artifact text. - Resolve
data_part_idto the sibling data part whosemetadata.dataPartIdmatches. - If
locatoris non-empty and starts with/, apply it as a JSON Pointer to the part’sdata.valueto highlight the specific field. If it matchesL<digits>orL<digits>-L<digits>, treat it as a line or line-range reference. If empty, cite the whole part.
Best practices
Do
- Render the artifact text verbatim. The server has already stripped markers; do not attempt to find or strip
U+E200/U+E202/U+E201yourself. - Join all text parts of the artifact in order before applying offsets. Offsets index the concatenated text, not any single part.
- Treat
offsetas a character (rune) offset, not a UTF-8 byte offset. - Resolve citations by
data_part_id, joining against the cited part’smetadata.dataPartId. - Apply
locatoras a JSON Pointer when it starts with/, or as a line range when it matchesL<digits>. Fall back to citing the whole part when the locator is empty or does not resolve.
Don’t
- Don’t display raw
U+E200/U+E202/U+E201markers. They are private-use codepoints that render as nothing or boxes. If you see them, treat the input as malformed. - Don’t treat offsets as byte offsets, or apply them to a single text part.
- Don’t fail the whole response if a
locatordoesn’t resolve. The server keeps the citation with its locator intact; mirror that by falling back to citing the whole data part. - Don’t treat
user_data_*oruser_text_*GIDs as resolvable citation targets. The server rejects user-origin parts as citation sources; guard against them on the client as well.
Graceful error handling
Citations never fail a task. The server handles malformed citations gracefully, and your client should mirror that behavior:
If every cited part was user-origin and no answer text was produced, the server retries rather than completing with an empty artifact.
Orchestration and persistence
The orchestrator’s system prompt instructs it to emit citation markers only when it has backing source data parts. The raw text with markers intact is persisted internally for cache reuse across turns, but the artifact you receive always contains the clean text and resolved citation metadata. Connector outputs reach the user only when the orchestrator cites them in its final answer. Calling a connector produces a data part, but only a citation marker in the orchestrator’s text passes that data part through to the client artifact.Next steps
- Learn how to stream responses and consume token deltas
- Read about the task lifecycle and artifacts
- Explore registry connectors that produce citable data parts