The assistant (chat)
The Futuros assistant is not a chatbot that "knows things". It is a reasoning layer over the verified corpus: every figure it states comes from a tool result, tied to its citation, never from the model's memory. That discipline is what sets it apart from a generic assistant — and what makes it usable by a regulator or an investor. It is available across the entire platform (the launcher lives in the root layout, on any page); surfaces like /sala pass it questions with context via their "go deeper in chat" cards. It embodies Pillar 2 — Sovereign Model: the moat is not the base weights, it is the corpus with provenance and the grounding discipline.
The reasoning pattern
The assistant does not fire tools at its first reading of the question. The order is explicit: think about the question → look at the corpus → research online if needed → ask if the question is ambiguous → reason → then answer. The sentence that drove this design, in the call it came from, was "it has to stop for a minute and reason": the failure it corrects is not answering wrongly, it is answering fast with the first figure retrieval returned.
That order is also what makes the sections below cohere: the specifying question arrives before the tool budget is spent, and strict relevance is applied when choosing which corpus figure deserves to enter the answer, not when laying it out.
Answer depth: thorough on the first response
There are three tiers, and the one served by default is the complete answer:
mode | What it is | Tool rounds | Output tokens |
|---|---|---|---|
full | The default. The complete first-shot answer: verdict, indicators, regulatory frame, financing read, precedent, who moves the lever, and the improvement close. This used to be what the Expand control bought. | 6 | 8,192 |
brief | The short executive brief. Still available on the API if a client sends mode:"brief"; the product UI no longer serves it or offers a control to expand it. | 6 | 4,096 |
briefing | Deep Briefing: a long, board-ready deliverable with seven fixed sections. | 9 | 16,384 |
full is the default mode. expand and the retired normal value are accepted as aliases of full. The UI does not offer an Expand / Ampliar button: the first response is already the complete version, in all three product languages (ES / EN / PT). An explicit, substantive brief may still carry expandable: true on the wire so an API client can re-send the same turn with mode:"full". A brief that was cut short is never offered as expandable: it is declared truncated first.
What changes between brief and full is not the grounding rules, which are identical across all three modes. What changes is how many ingredients are enforced. The completeness verifier (server/chat/completeness.ts) applies ENFORCED_BY_MODE: a brief enforces only the improvement close, while full and briefing additionally enforce the regulatory frame and the financing read. Grading is unaffected by mode — the report still states the true status of every ingredient, so telemetry stays comparable across modes; only the enforceable list narrows. A brief is shorter, never less cited: the internal rule is that a brief with no numbers is a failure, and the "how to improve" close is never dropped, because it is the point. Thorough does not mean inventing rates, tenors, tickets or counts.
When the brief serves a financing decision
The brief block explicitly told the model to leave out the financing read. For a minister skimming a pillar that is right. For an investment officer at IDB or CAF, the money IS the question, and dropping it produced a 1,357-character answer with zero indicator values, zero years and zero sources against a prompt asking for exactly those three.
detectDfiIntent (server/chat/dfi-intent.ts) settles this deterministically: no model call, no latency, same question and same verdict every time. It fires on a single signal — a development-bank counterpart (BID/IDB/CAF), capex or opex, ticket, tenor, bankability, investment committee, financial close, or a currency amount — and returns which signals matched, so a failure is diagnosable rather than opaque. The acronyms are matched case-sensitively on purpose: a case-insensitive \bbid\b would fire on the English verb bid and on bidding in any procurement question.
On the product default (full) the financing read is already a required ingredient, so a DFI question gets it on the first turn. When the detector fires, and only in brief (full and briefing already enforce the financing read), one appended block inverts exactly one line: the money read stays in. Everything else on the exclusion list stays out. The target rises to 180-260 words, and the block insists that a ticket range is what the issuer can lend, never what was committed to this pilot. It costs no extra retrieval round: recommend_action already returns capex, opex, ticket-fitted instruments and the ranked decision-makers.
Up to three specifying questions
Faced with a genuinely ambiguous or complex question, the assistant may return up to three questions before answering (MAX_QUESTIONS = 3, with at most 4 options each). Never on a simple lookup: asking back for a direct figure is friction, not rigor.
They are not only clarifying but specifying questions: they let the asker narrow what they are after. That is why a question may legitimately carry no options — it is free text. When none of the questions carries options, the futuros:clarify block is replaced by prose, so a legacy client shows the questions as text rather than an empty bubble. The questions field always carries the full ordered list; question/options mirror the first option-bearing question, so a client that only understands the old shape still renders its chips.
Grounding rules
The rules are strict and explicit:
- Every figure comes from a tool result, marked with a
[[cite:id]]token. No tool, no number. - No answers from memory. The model does not fill in figures from its training; if it did not get them from a tool, it says it does not have them.
- Contradictions are shown. If two sources disagree, the assistant flags it instead of silently averaging.
- Vintage is marked. An old figure is presented with its year, not as if it were current.
- Answer first (BLUF). The first sentence carries the verdict and the most important cited figure — not a preamble.
- Ask before guessing. Faced with a genuinely scope-ambiguous question, the assistant emits a
futuros:clarifyblock and stops, instead of silently assuming (up to three questions — see above). - Strict relevance. A corpus figure only enters if it answers this question. The case that set the rule: asked about compute sovereignty in Mexico, "active GitHub developers" is the largest number in the innovation cell and is a proxy for nothing — the relevant constraint was the electricity grid. Evidence is ranked by pertinence to the question, not by size or by the cell's own order.
- People are cited. When the answer touches who decides something, the key-people corpus is cited like any other source, with its
persona-…-src-…id. - It answers in the language of the question — any language. Not just ES/EN/PT: Dutch for Suriname, French or Kreyòl for Haiti, Caribbean English.
- It does not invent coverage. The live product is 25 countries + the LATAM row, 10 pillars, 260 cited cells. GDB (Oxford Insights) and GIRAI are not ingested; ILIA is not a live corpus index. It never claims ~9,000 sources. Chile INE.Stat is live only for unemployment and informality. A thorough answer names the hole when the source is not in the catalog.
Live coverage versus what is promised
Thorough does not mean a larger product. If a room asks for an index or an office the corpus did not ingest, the first sentence says so, then the nearest cited measure is given and labelled adjacent. Live paths for the recurring questions:
- CEPAL versus INEC / INDEC / DANE: both series from the cell, cited, never averaged;
get_discoveries/get_regional_pulsefor documented contradictions (curated list, not exhaustive). - Ecuador malnutrition:
cepal_desnutricionin health and cohesion; vintage cited. - Ecuador HE / research owners: personas + regulation + pilots (CEDIA; SENESCYT absorbed into the viceministry, Decree 100 / 2025). INEC is the statistics office, not the HE owner.
- AI readiness versus ILIA/GDB, and who has a strategy:
ai-governance/national_ai_strategy.status. GDB, GIRAI and ILIA are not cited as Futuros data. - Training-corpus energy (Ember/IRENA): live Ember and IRENA series in environment +
frontier-country. - SDG coverage and civic-space gaps:
get_data_health(coverage_gaps); civic-pulse (GDELT, directional signal); civic-attention (21/25); democratic resilience; V-Dem in the cells. Comparable civic space for all 25 countries, when ingested, is cited ascivicus_monitor_score(CIVICUS Monitor, country page) — a second index is not invented. - Argentina fiscal / BOOST:
boost_spending_*in ARG cells (2022 vintage, aging) +markets-screener. - Chile INE.Stat health/innovation: not ingested; the health or innovation cell is cited from its real sources (WHO, World Bank, ECLAC).
The tool suite
The assistant acts through deterministic tools, with a budget of 6 tool rounds per answer (9 in briefing mode). In plain language:
| Tool | What it does |
|---|---|
resolve_metric | Translates what the user asks for into a concrete indicator from the corpus. |
get_parameter_data | Retrieves an indicator's value by country, with its citation. |
search_corpus / get_dataset | Semantic search over the corpus and direct access to a cataloged dataset. |
compare_countries | Compares countries on an indicator. |
get_news / get_regional_pulse | Cited news and the regional pulse. |
compute | Deterministic computation. Includes forecasts (marked as projection, not fact) and explain_change (which declares association, not causation). |
get_data_health | Health and freshness status of the data. |
get_discoveries | Baked findings (movements, anomalies). |
find_pilots / get_pilot | Bankable pilots by criteria, and a pilot's profile. |
recommend_action | Chains pilot → financing instrument → decision-maker: from data to action. Includes policy_screen (Bhutan/NZ/Humphrey; rubric in docs/plataforma/cribado-politicas.md; human override to export). |
web_search / fetch_url / perplexity_search | External search, last resort and visibly badged. |
The design key: adding a dataset (server/chat/datasets.ts) automatically makes it cataloged, retrievable and searchable — without writing new tool code. Retrieval combines Voyage embeddings (voyage-3-lite, 512-dim) over search-index.json with a BM25/keyword fallback and reranking.
Briefing mode. With mode: "briefing" the assistant produces a long, board-ready deliverable with seven fixed sections (executive summary, current situation, trend, peer comparison, data quality and currency, contradictions and risks, "what to watch") and the larger budget from the table above. The same grounding rules apply inside the briefing.
Blocks an answer may emit
Beyond cited prose, an answer may include blocks the UI renders as components. They all share one rule: they only show values a tool returned, and every datapoint keeps its citation.
| Block | What it renders |
|---|---|
futuros:viz | Chart or matrix (5 kinds). At most one per normal answer; a briefing may use one per section. A brief prefers none. |
futuros:map | Map of per-country datapoints. This is the surface where competitors were beating us, and the one that makes a multi-country pilot legible: each mapped datapoint carries its ISO3, its value and its citation, and the citation chips render beneath the map. |
futuros:pilots | Pilots aligned to the asker's real problem (see strict relevance). |
futuros:personas | Corpus people cited as the decision counterpart. |
futuros:news | News and social pulse, explicitly marked as "latest news, not 100% verified data" — never with an indicator's blue citation. |
futuros:laws | The law-and-articles snapshot (see below). |
Verified corpus versus external source: blue and gold
The provenance split is visual and carries a legend:
- Blue — inside the Futuros corpus: verified, with provenance and a deep link.
- Gold — outside the corpus: an external source a web search returned in that turn. It is not verified by us, and it says so.
Gold additionally carries a ⌖ glyph next to the citation, and the palette keeps three separate tones (--gold for text, --gold-line for the stroke, --gold-soft for the fill) precisely because on a light surface warning amber and provenance gold are neighbours and must not be confused. The wording is softened ("unverified figure") but the distinction is never removed: it is what underwrites the provenance promise.
Regulation: the law and its articles, not "51 instruments"
Faced with a regulatory question, the answer names the law and the articles that govern the topic and offers a pop-up snapshot of the relevant passage, instead of reporting an aggregate ("51 instruments") and sending the reader to an official portal raw. The futuros:laws block is built from the instrument's own citation record — a regulation citation id is the instrument id and always starts with the country's ISO3 in caps, which makes the partition exact rather than heuristic.
The regulation-instruments and regulation-provisions datasets count as regulatory evidence but not as core data: they are lookups, so a purely legal answer never grades as "substantive" on the strength of having consulted them.
When the model names a law in prose or in a table but omits the [[cite:id]] token, the renderer underlines the name if it matches an alias of an instrument already present in that turn's citation payload — acronyms (LGPD), official references (Ley N° 26.743), full titles. It never invents links to instruments the answer did not ground: if the law is not in the citation payload, the name stays plain text. System prompt rule 9 requires [[cite:id]] on every named law; this fallback is interface honesty for the case where the model grounded the instrument but forgot the inline token.
Automatic data-gap ticket
When a corpus tool comes back empty, the server files a structured ticket rather than letting the signal die with the request, and publishes it at /vacios via GET /api/backlog. This is the mechanism that turns an unknown unknown into a known unknown, ranked by real demand.
Two things that matter:
- A network failure or a tool error never files a ticket. The signal is
returnedEmpty, which records only true corpus emptiness; an outage is never mistaken for a hole.search_corpusdoes not count either: poor phrasing looks identical to a gap. - The ticket does not keep the question. It carries the country, the axis and the tools that came back empty — all platform vocabulary. The endpoint is public and unauthenticated, so anything the ticket carried would be something the asker published without meaning to. Each row's headline on
/vaciosis composed from the country and the axis. (The schema also allowsmissing_indicators, but the handler never passes it, so the field is always absent: documented as what it is rather than what it promises.) The full rule lives at the top ofserver/backlog/gaps.ts; the endpoint is documented in the public API.
The gap reaches the model, not only the backlog
Filing the hole is useless if the answer papers over it. Emptiness reached three places — telemetry, the public backlog and the client chip — and never the one consumer that decides what gets written. On the wire, a zero-result query was 68 bytes of {"dataset":"financing","total_matches":0,"records":[]}: no prose, no prohibition. An empty array reads as "nothing to add here", not as "you have nothing to write this from". One financing question logged financing → 0 records twice, filed its ticket, and printed a six-row table of tenors, rates and concessionality grades anyway.
Every empty result now travels to the model with empty: true and an explicit corpus_empty_note, in two variants, because "empty" means two different things: a real absence (a corpus tool asked about a specific cell that does not exist) bans producing any table, row, figure, name, date, rate, tenor, grade or amount for that dataset; an unproven absence (search_corpus, resolve_metric, the web tools, which can come back empty on phrasing alone) says exactly that instead. Both share ONE definition of empty (isEmptyResult) and ONE list of which tools mean absence (NOT_A_GAP, now in server/chat/empty-note.ts and imported by gaps.ts) — that they could disagree about the meaning of "empty" was the underlying defect.
And a rule that runs against intuition: the model may no longer claim the gap was registered. It cannot observe that — the ticket is written after the answer ends, and the client renders the backlog link itself — so asserting it was a false claim about the platform's own behaviour. A production answer promised that Futuros "automatically registers these gaps in the public backlog" while telemetry showed zero tickets. Describe the hole, yes; promise its filing, no.
A cargo survives a cabinet change; a name does not
The corpus's differentiating asset is knowing who to call, and it was where the product failed hardest: across ten development-finance questions, wrong office-holders appeared in six, load-bearing every time. A minister who had left ten days earlier still carried "readiness 90/100, sentiment favorable" — a confidence number that converts a dated record into an assertion about the present.
Three changes, none of which requires re-baking the 1,346 profiles. former is honoured: 47 index entries mark an ex-holder and no consumer read them; filtering them stops 41 country×pillar cells from offering someone who has left (first-ranked in 18 of them), and it is free — of 278 non-empty cells, none empties and none drops below three candidates. Every row carries its date: there is no verified_as_of field in the corpus, so the stamp derives from last_updated (present on 1,346 of 1,346 profiles, already loaded by the ranking) as verified_as_of, verified_days_ago and currency (recent ≤30 days, aging ≤90, unverified beyond). The median profile is ~66 days old, which is why recent cannot mean 90. The confidence numbers are gated on currency: when currency is not recent, readiness and relationship_status come back null with a readiness_withheld reason — justified by the data, since readiness is exactly 55 on 1,281 of 1,346 rows because that is the stub default, not a measurement.
The prompt closes the loop: when a row is not recent, name the cargo and the institution, and only then, optionally, the last recorded holder with its date. Pilot-embedded contacts get the same treatment — only 155 of 336 carry a persona_id, and 172 of the 297 distinct names are absent from the directory, so a name is emitted only when its id resolves; otherwise the office survives and the person does not.
The corpus has no prices
The financing catalog holds issuer, type, ticket range, eligibility text and application URL. It holds no rate, spread, tenor, grace period, concessionality grade or credit rating — verified field by field across all 64 instruments. Those were exactly the figures one answer invented: IDB sovereign money at "1.5-2.5%", CAF described as "AAA-rated", and a mechanism whereby a MIGA wrap compressed 40-60 bp off a board-set spread. They are the figures most likely to end up in a term sheet.
Two defences. A prompt rule: name the instrument and its issuer, give its ticket range with a citation, point at application_url, and say plainly that Futuros carries no pricing or terms; concessional-loan is a catalog label, not a verified concessionality grade, and a pilot's financing[] is a design split, never committed co-financing. And a deterministic detector: detectFinancialTerms flags the co-occurrence, within one claim unit, of lending vocabulary and a term token (a percentage, basis points, benchmark + margin, a duration, a rating grade). It is advisory, not a rewrite: it caps the confidence chip and rides the verify event as fabricated_terms. Tables are read as column header + row label + cell, because in a terms table the word lives in the header and the number in the cell. What it does not flag matters as much: the discount rate (0.07-0.14 on all 67 pilots), the horizon in years, counterpart share, disbursement lag, TRL and the whole affordability panel are real corpus data, and a false positive there would discredit a correct answer.
Private question log
The ticket describes the shape of the gap, not the question — the right constraint for a public surface, but it left the single most useful signal for improving retrieval unrecorded: what people actually ask. Hence a second store, private and separated by construction:
- A different key.
chat:questions:v1, neverbacklog:gaps:v1. Nothing that reads the backlog can reach it, even by accident, and the two are flushed independently. - The public read path is untouched.
sanitizeTicketstill rebuilds every published ticket field by field, so a question cannot ride out throughGET /api/backlogor/vacios. This adds a store; it removes no guarantee. - The private read is fail-closed.
GET /api/questionsrequiresAuthorization: Bearer $QUESTION_LOG_TOKENand refuses every request when the token is unset (theapi/alerts-digest.tsposture). A deployment that forgets to set it serves nothing, rather than publishing the log. - The client stays shape-only; the server carries the text. PostHog's
chat_query(and the funnel eventchat_submit) carry length, language, route, tier and ageneration_id— never the text. That id joins those halves to the private question log and to the server eventschat_generation/$ai_generation, which do include the question (and a clipped answer) for LLM traces. The words never travel on the client analytics pipeline.
Each row keeps question, timestamp, language, country, axis, route, tier and generation_id. The write is best-effort and fires before the stream opens: it never blocks or delays the answer, and a store failure costs a datum, never a turn.
Anatomy of an answer (the SSE protocol)
An answer is an SSE stream of named events (server/chat/handler.ts produces; src/lib/chat-client.ts hand-parses the wire format, because EventSource only does GET). The loop: stream one model turn → if it requested tools, execute them against the static corpus, append results, stream the next turn → repeat until end_turn. Volatile context (country, parameter, reply_language) is appended to the last user turn, never to the system prompt — so the cached prefix (cache_control: ephemeral) stays byte-stable across requests. The incoming request is trimmed before touching the model: last 12 turns, 2,000 characters per message, and the last turn must be from the user.
Event catalog, with what the interface does with each one (src/store/chatStore.ts):
| Event | Payload | What the interface does |
|---|---|---|
provider | {id, sovereign, inRegion?, model} | Sets the turn's provider / sovereign-mode badge. First thing emitted; re-emitted on every failover. |
text | {d} — incremental delta | Appends the delta. The first token after a wipe discards the "shadow" and starts the fresh text. |
reset_text | {} | Never wipes to black: it moves the already-emitted text into a dimmed shadow layer — the user never sees answer → dots → answer. |
tool | {name, detail, found?} | Adds a step to the "show the work" trace. found is derived from the actual result, after execution. |
citations | {items} | Registers the new chips (deduped by id) before the next text turn, so [[cite:id]] resolves while the prose streams. |
revising | {unverified_ids} | Shadow + "verifying…" chip + marks the turn as self-corrected; clears any previous verify/followups. |
verify | {claims, traceable, sources, unverified_ids, uncited_figures, contradictions, values_confirmed, values_checked} | Feeds the trust badge. Emitted only if there is something to report. |
followups | {items} | Shows the 3–4 suggested follow-up questions. |
error | {code, message, reason?, status?} | Attached inline without wiping the partial. reason (auth/credits/unavailable, derived from the HTTP status only, no secrets) decides whether "retry" makes sense. |
done | {usage, truncated?} | Closes the turn. Final floor: if it ended empty but a shadow exists, it is restored — a turn never ends blank if prior text existed. |
The client adds its own sentinel: a stream EOF without a terminal done/error event is a severed stream — the partial answer is marked as truncated, never as complete, and never routed to error (that would wipe the partial).
Faithfulness verification
Before delivering an answer, the assistant audits itself with a deterministic grounding verification — with no extra model call whatsoever (server/chat/faithfulness.ts):
- Claim extraction — deterministic, over the
[[cite:id]]tokens, including the data points of visualization blocks and table cells. - Grounding verification — every cited id must be an id a tool actually returned in this conversation (
verifyGrounding). The judge is thereturnedIdsset — which includes ids embedded in the payloads even if they were never emitted as chips — not the citation registry: judging against the registry would trigger corrections on honest answers. An id that came from no tool is a fabricated citation, and is treated as such. - Value confirmation — it is not enough for the citation to resolve: the number the model wrote must match the source's value (
confirmValues), harvested from the tool results themselves. The mechanics are conservative: take the figure closest to the citation (skipping years 1900–2099), compare at the precision it was written at (75.9 → one decimal), and ×/÷100 scale jumps are only accepted if the figure was written as a percentage or the source is a fraction (|source| < 1) — an unconditional wildcard would wave through a two-orders-of-magnitude error. The signal is positive-only: an unparseable figure neither confirms nor accuses. - One-pass self-correction — if the final answer cites ungrounded ids, the handler emits a
revisingevent and spends one extra round correcting. A salvage guarantees the answer never ends up blank or worse than the draft: the revision is discarded as degenerate if it comes out empty, if it is an acknowledgment (the "You're right…" family, or an apology opener combined with a redo verb — the apology alone never triggers, because it also opens honest abstentions), or if it does not shrink the set of ungrounded ids (the semantic signal outranks the textual one). In that case the previous draft is restored with the invalid citation tokens removed — exactly the rewrite that was requested. The salvage runs after the loop, not only on a cleanend_turn: the clock can expire before or mid-revision, and the full draft must still be restored.
An LLM judge that classifies claims as SUPPORTED / PARTIAL / UNSUPPORTED also exists, but it is offline faithfulness evaluation (scripts/chat-faithfulness.ts), not a runtime step — the two are never conflated.
The result travels to the interface in a verify event with counts and id lists — total claims, how many are traceable, distinct verified sources, unverified ids, figures without a citation — plus the detected contradictions and the value confirmation. It is emitted only when there is something to report. From it comes the trust badge the user sees: an honest reading of how grounded the answer ended up, not an ornament.
There is also defense against prompt injection: quarantine.ts isolates untrusted content brought in by tools so it cannot rewrite the assistant's instructions.
Grounded visualizations
The assistant can emit at most one futuros:viz JSON block per normal answer (a briefing may use one per section) — chart or matrix — which the interface renders interactively. Its data points carry citation ids and go through the same grounding verification as the prose: a chart that cites an id no tool returned triggers the same correction pass.
Honest limits
The answer runs under explicit budgets: ~100 s of wall clock (within the function's 120 s maxDuration, with ~20 s reserved for verification and follow-ups), 8,192 output tokens (16,384 in briefing), 12 turns of history and 2,000 characters per message. The clock is shared across rounds: each model turn runs with an abort signal set to the remaining budget (and to client disconnect — closing the tab aborts the in-flight turn, the tools and the follow-ups, without billing tokens nobody reads). It governs the tools too, not just the model: external ones carry fixed 12–20 s timeouts in series, so with the budget exhausted the remaining ones degrade to an is_error result from which the model lands with what it already has — the protocol requires one result per tool_use, and this avoids a silent platform kill with no terminal done. If the budget runs out, the cut is declared: the final event carries done.truncated, the interface marks the answer as "truncated" instead of presenting a fragment as complete, and a guard (stripDanglingFence) removes any half-emitted visualization block: an odd number of fence markers means the last block never closed — it is trimmed from there, so the truncated JSON neither renders broken nor inflates the count of uncited figures.
The provider chain and sovereign mode
server/chat/provider.ts builds a provider chain with transparent failover (providerChain()). The primary precedence remains sovereign → OpenRouter → Anthropic — but every configured provider is also appended as a fallback: if the primary fails at the transport level (connection timeout, auth or credit 4xx, 5xx after the SDK's own retries), the round is retried against the next one without consuming recovery budget (round -= 1), re-resolving the model id per provider namespace (anthropic/claude-sonnet-4.6 on OpenRouter vs claude-sonnet-4-6 direct). Failover has a strict gate: it only fires if the round has not yet emitted visible text — with tokens already on screen, a second provider re-streaming would duplicate the answer, so the host is never switched mid-text. Follow-ups also run on the provider that actually answered — after a failover the primary is dead. The default is Sonnet for answers; follow-up questions run on Haiku (claude-haiku-4-5), which is cheaper.
The usingSovereign() function checks whether SOVEREIGN_INFERENCE_URL / _KEY are configured; if they are, inference runs on an in-region gateway (LiteLLM/vLLM) over open weights, and the corpus and the query never leave the host. That is the seam of Pillar 2: sovereignty is a provider toggle, not a rewrite, and frontier quality remains the default until parity is reached in the evaluations.
The active provider is streamed as a provider SSE event on every answer — the first thing emitted, and re-emitted on every failover, so the badge names the host that actually answered. The "sovereign mode" badge in the interface appears only when the inference was sovereign: the user knows, inside the conversation itself, where the inference they are reading lives.
External search as a last resort
The open web is only used when the corpus is not enough, and always badged as external (web- citations). A fact from the web is never presented at the same level as a datum from the verified corpus — the separation is visible, as explained in Provenance and citations.
And it is optional: with the corpus-only scope (corpusOnly) the external tools (web_search, fetch_url, perplexity_search) are withdrawn entirely, and the model answers solely from the verified corpus — or declines.
Handoffs
The assistant is not a dead end. It can hand control to analytical surfaces: to /explorar to query the data with SQL in the browser (DuckDB-WASM), and to an embed board (/embed/board) to pin and share a set of findings with their provenance intact. The same corpus is exposed to external agents via the MCP server — the chat's corpus tools plus two MCP-only ones (get_series, resolve_citations).
The close is also action, not just data: every substantive answer ends with a "how to improve" lever and a futuros:pilots block of 1–3 validated pilot cards — with explicit rules for omitting it (clarifications, refusals, no data, questions that are already about pilots). And after each answer come 3–4 suggested follow-up questions, generated best-effort by the cheap model (Haiku) — even when the answer was truncated. The full envelope of an answer: cited prose (with its table, if it compares) → visualization → workbench/board link → pilots → follow-ups.
In one sentence: the assistant turns a corpus with provenance into answers that declare their own confidence, correct what they cannot support, and say where their inference runs. The honesty discipline that governs it is detailed in Methodology and honesty.