> ## Documentation Index
> Fetch the complete documentation index at: https://docs.corti.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Render inline citations

> Learn how the orchestrator links claims in its answer to source data parts with inline citations, how to read citation metadata, and how to render citations correctly.

When an agent answers using data it retrieved from connectors, the orchestrator embeds **inline citations** that link each claim in its answer text to the specific source data part that backs it. The server strips the citation markers from all client-visible text and delivers the citations as structured metadata on the artifact, so you can render clean prose while still linking each claim to its source.

This guide explains the citation metadata, how to resolve a citation to its source data part, and how to render citations. It also covers the orchestrator-only restriction and the graceful error handling the server applies.

## Prerequisites

* An existing agent with at least one connector that produces data parts (for example, [PubMed](/agentic/registry/pubmed) or [Medical Coding](/agentic/registry/medical-coding))
* Understanding of the [task lifecycle](/agentic/task-lifecycle), [artifacts](/agentic/core-concepts#artifacts), and [parts](/agentic/core-concepts#part-types)
* Understanding of [streaming responses](/agentic/guides/stream-responses)

## How inline citations work

1. The orchestrator receives the user message and decides which connectors to call.
2. Each connector call produces one or more **data parts** stored on the task. Every data part gets a stable **data part ID** (GID) such as `tool_data_01`.
3. The orchestrator writes its answer text and embeds citation markers that reference the GIDs of the data parts it used.
4. The server parses the markers, strips them from the text, and attaches a `citations` array to the answer artifact's text part metadata.
5. The cited data parts are emitted alongside the answer text in the same artifact, each carrying its GID in its `metadata.dataPartId`.

You render the answer text verbatim and use the `citations` metadata to link each claim back to its source data part.

<Note>Inline citations are produced by the **orchestrator only**, the agent runtime that serves your top-level task. Connector (expert) sub-answers do not embed citation markers themselves; their outputs are data parts that the orchestrator then cites or passes through. You only need to handle citations on the top-level task answer.</Note>

## Marker format

The orchestrator uses the [OpenAI citation formatting guide](https://platform.openai.com/docs/guides/text-generation#citations), with three private-use Unicode characters as delimiters:

| Codepoint | Name | Role |
| - | - | - |
| `U+E200` | citation start | Opens a marker |
| `U+E202` | citation delimiter | Separates fields inside a marker |
| `U+E201` | citation stop | Closes a marker |

A marker has the form:

```
\ue200cite\ue202<source_id>\ue201                 whole-part citation
\ue200cite\ue202<source_id>\ue202<locator>\ue201   citation with a locator
```

* `cite`: backs a claim with a source. The marker is stripped from the text and an entry is added to the `citations` metadata.
* `<source_id>`: the GID of the cited data part, for example `tool_data_02`.
* `<locator>`: an optional sub-reference into the cited part: a [JSON Pointer](https://datatracker.ietf.org/doc/html/rfc6901) path such as `/data/5/population`, or a line range such as `L8-L13`. Empty means the whole part is the source.

<Warning>You should never observe raw `U+E200` / `U+E202` / `U+E201` characters in text you receive from the API. The server strips all valid markers before the artifact reaches you. If you do see these codepoints, treat the text as malformed. Do not display the markers to users.</Warning>

## Citation metadata

Citations ride on the **text part** metadata of the answer artifact, not on the data parts. The text part `metadata` contains a `citations` key: an array with one entry per citation marker, in text order.

Each entry in the `citations` array has the following fields:

| Field | Type | Description |
| - | - | - |
| `data_part_id` | string | GID of the cited data part (for example `tool_data_02`); join this against the cited part's `metadata.dataPartId` |
| `offset` | int | Character (rune) offset into the joined artifact text where the marker stood, snapped past any trailing punctuation so the citation lands on a clause boundary |
| `locator` | string | JSON Pointer (for example `/data/0/name`) or line range (for example `L8-L13`); empty means the whole part is the source |

The array is sorted by `offset`.

### Example: answer with citations

A PubMed-backed answer whose text part metadata looks like this:

```json theme={null}
{
  "text": "Doxorubicin is associated with dose-dependent cardiotoxicity. The recommended cumulative dose limit is 450 mg/m².",
  "metadata": {
    "citations": [
      { "data_part_id": "tool_data_02", "offset": 53, "locator": "/results/0/findings" },
      { "data_part_id": "tool_data_03", "offset": 107, "locator": "" }
    ]
  }
}
```

The cited data parts appear as sibling parts in the same artifact, each carrying its GID:

```json theme={null}
{
  "data": {
    "value": {
      "results": [
        { "findings": "Doxorubicin carries a dose-dependent risk of cardiomyopathy." }
      ]
    }
  },
  "metadata": {
    "dataPartId": "tool_data_02",
    "toolName": "pubmed_search"
  }
}
```

To resolve the first citation: take `data_part_id` (`tool_data_02`), find the sibling part whose `metadata.dataPartId` equals `tool_data_02`, then apply the `locator` (`/results/0/findings`) as a JSON Pointer to that part's `data.value`.

## Streaming and citations

When the answer is [streamed](/agentic/guides/stream-responses), the orchestrator emits one text part per token delta, then a closing event that replaces the deltas with the canonical consolidated text. The `citations` metadata is attached to that final consolidated text part.

<Tip>Offsets index the **joined** artifact text, the concatenation of all text parts in order. If you process the artifact as multiple parts, join them before applying offsets. Do not apply an offset to a single part.</Tip>

During streaming, the server also strips citation markers from each delta as it arrives, so partial markers never reach your client even when a delta splits a marker across a chunk boundary.

## Render citations

The text you receive is already clean (markers stripped). Render it verbatim and use the `citations` metadata to add your own citation UI:

1. Read `metadata.citations`.
2. For each entry, insert your rendered citation marker (superscript, footnote, or popover) at `offset` characters into the joined artifact text.
3. Resolve `data_part_id` to the sibling data part whose `metadata.dataPartId` matches.
4. If `locator` is non-empty and starts with `/`, apply it as a JSON Pointer to the part's `data.value` to highlight the specific field. If it matches `L<digits>` or `L<digits>-L<digits>`, treat it as a line or line-range reference. If empty, cite the whole part.

## Best practices

### Do

* **Render the artifact text verbatim.** The server has already stripped markers; do not attempt to find or strip `U+E200` / `U+E202` / `U+E201` yourself.
* **Join all text parts of the artifact in order before applying offsets.** Offsets index the concatenated text, not any single part.
* **Treat `offset` as a character (rune) offset**, not a UTF-8 byte offset.
* **Resolve citations by `data_part_id`**, joining against the cited part's `metadata.dataPartId`.
* **Apply `locator` as a JSON Pointer when it starts with `/`**, or as a line range when it matches `L<digits>`. Fall back to citing the whole part when the locator is empty or does not resolve.

### Don't

* **Don't display raw `U+E200` / `U+E202` / `U+E201` markers.** They are private-use codepoints that render as nothing or boxes. If you see them, treat the input as malformed.
* **Don't treat offsets as byte offsets**, or apply them to a single text part.
* **Don't fail the whole response if a `locator` doesn't resolve.** The server keeps the citation with its locator intact; mirror that by falling back to citing the whole data part.
* **Don't treat `user_data_*` or `user_text_*` GIDs as resolvable citation targets.** The server rejects user-origin parts as citation sources; guard against them on the client as well.

## Graceful error handling

Citations never fail a task. The server handles malformed citations gracefully, and your client should mirror that behavior:

| Condition | Server behavior | Client handling |
| - | - | - |
| Invalid marker (does not match the grammar) | Kept as literal text; logged | Render the text as normal prose |
| Unknown GID (the cited data part is missing) | Citation dropped; logged | You will not receive an entry for it |
| Unresolvable JSON Pointer locator | Citation kept with the locator intact | Fall back to citing the whole data part |
| User-origin GID (`user_data_*`, `user_text_*`) | Rejected as a citation target | Guard against these GIDs |

If every cited part was user-origin and no answer text was produced, the server retries rather than completing with an empty artifact.

## Orchestration and persistence

The orchestrator's system prompt instructs it to emit citation markers only when it has backing source data parts. The raw text with markers intact is persisted internally for cache reuse across turns, but the artifact you receive always contains the clean text and resolved citation metadata.

Connector outputs reach the user only when the orchestrator cites them in its final answer. Calling a connector produces a data part, but only a citation marker in the orchestrator's text passes that data part through to the client artifact.

## Next steps

* Learn how to [stream responses](/agentic/guides/stream-responses) and consume token deltas
* Read about the [task lifecycle](/agentic/task-lifecycle) and artifacts
* Explore [registry connectors](/agentic/registry/overview) that produce citable data parts
