SILT Research Series · SILT-RP-007

Model Disclosure in Composed AI Systems

Census #1 of the Model Disclosure Index (MDI) programme: whether routers and orchestrators name the model that answered

Version 1.2Issued 2026-09-24Sentient Index Labs & Technology
This document is fixed. It describes the instrument as of version 1.2, issued 2026-09-24, and it will not be edited. Corrections and revisions are issued as a new version under the same identifier, each with a dated note recording exactly what changed, so a citation to an earlier version can still be read against this one. Where this paper and our website disagree, the website is more current and this document is what was true on the date above.

1.Abstract

A composed system routes a request among several models rather than answering with one, and they are increasingly what enterprises deploy. This census asks something narrow that appears not to have been measured: when you ask one a question, does it tell you what answered? Across three commercially available systems, the model field every portable client reads returned a name belonging to no model at all. One answered a short question using three invocations of two different companies' models and disclosed the fact only in a vendor extension outside the standard interface; another disclosed nothing anywhere. Asked the same question five times, one system used five different combinations of models costing between 3,558 and 13,008 tokens, and four blind judges scored the most expensive draw the lowest of the five. The paper reports what an audit log therefore does not contain, states the limits of a single census, and declines — in both directions — to draw any conclusion about intent.

2.What was measured, and what was not

A composed system takes a request and distributes it among several models rather than answering with one. They are cheaper and often better, which is why they are increasingly what organisations actually deploy. This census asks one narrow question about them: when you ask one a question, does it tell you what answered?

It is a census, not a battery. No system is scored on a scale or ranked, and no level, grade or threat rating is assigned. (The pilot in section 5 scores individual answers, not systems.) Those belong to instruments that measure models, and a composed system is not a model.

Three systems were probed: two from groq and one from xAI. Four kinds of question — a one-word lookup, a short coding task, a multi-step arithmetic problem and an open-ended research question — were each asked five times, and the composition serving every draw was recorded. Sixty-three calls in total: sixty for the repeated questions, and three single-call disclosure probes, one per system.

3.Disclosure: the model field names a thing that is not a model

In all three systems the model field — the only field a portable OpenAI-compatible client reads — returned the name of the composed system itself. Asked a short question about load-bearing walls, one system produced its answer using three separate model invocations: two of Meta’s Llama 4 Scout and one of OpenAI’s gpt-oss-120b. That fact appeared only in a vendor-specific extension the standard interface does not surface.

The second groq system behaved the same way with two components. xAI’s multi-agent system disclosed nothing at all: its response named the system, its usage block reported how many tools and sources were used and what the call cost, and nowhere did it say which model produced the text.

An organisation that logs the model field for audit — the obvious and documented thing to log — records the composed system’s name. It does not record that a Meta model and an OpenAI model processed the input. Two vendors, inside one call, under one name, invisible to the audit trail the platform itself provides.

For anyone operating under a data processing agreement, a subprocessor disclosure obligation, or a contractual term about where information goes, that is a gap between what is required to be known and what can be found out.

A second finding concerns the interface rather than the answer. xAI’s multi-agent model is refused outright on the chat-completions endpoint every other model here accepts, returning an error indistinguishable from a dead API key. It answers on a different endpoint. Across two vendors, then, the portable interface either conceals the composition or excludes the composed system from itself.

4.Stability: the same question, five different machines

On the one-word question, one system used the same two invocations in all five draws. On the arithmetic question, five identical asks produced five different compositions, ranging from four model invocations to eleven. Not one draw matched another.

The instability is not spread evenly. It concentrates where the work is hard, which is precisely where a buyer would most want to know what they were getting. A test using a single easy prompt lands on the quiet end of that range and reports stability.

We do not describe any of these systems as stable, including the one that never varied. Five draws with no variation is consistent with a true variation rate as high as forty-five per cent, at the conventional 95% confidence level.

v1.1 (24 September 2026). Version 1.0 said “as high as twenty-eight per cent”. That is the upper bound at 80% confidence, given without saying so; at the 95% level used everywhere else in our work it is about 45%. The correction widens the bound and strengthens the point it was making. Nothing else in this paper changed.

v1.2 (24 September 2026). Three statements checked against the census data. Section 2 said “Sixty-three calls in total” without saying the three beyond 3 × 4 × 5 were single-call disclosure probes, one per system; it now says so. Section 5 described the pilot as “one prompt, one system”; it was one system and three prompts, five draws each, which is where its fifteen answers come from. And “nothing here is scored on a scale” is narrowed to what is true: no system is scored or ranked; the pilot scores answers.

Cost follows composition. On the arithmetic question, with identical input, total tokens consumed ranged from 3,558 to 13,008 — nearly a factor of four for the same question, decided by a process the customer cannot observe and is not told about.

5.Did the extra work buy anything?

The five answers to the arithmetic question were already recorded, each produced by a different composition. They were shown to four independent judges who saw only the question and the answer — never the composition, never the invocation count, never each other’s scores — in shuffled order.

  • Four invocations scored 7.0
  • Five scored 10.0
  • Nine scored 7.5
  • Ten scored 8.5
  • Eleven scored 6.5
The most expensive draw scored the lowest of the five. Across all fifteen answers judged, in three separate cells, there was no relationship between how much work the system did and how good the answer was.

That result is a pilot and is reported as one. It is one system, three prompts and five draws of each; the scores above are the arithmetic prompt's five. It cannot support a rate and none is offered. What it supports is the narrow claim a buyer would care about: on this question, on this system, the run costing nearly four times as much was not better, and by this panel it was slightly worse.

6.What we decline to conclude

An earlier draft of this material stated that none of this was deception. That sentence has been removed, and the reason is worth stating because it is easy to miss.

Our published rule is that a measurement names indifference and never intent, because no artifact can show what a system or a company wanted. That rule cuts in both directions. If the evidence cannot establish deception, it cannot establish its absence either, and an exoneration is a claim about intent exactly as much as an accusation is.

So neither claim is made. What can be reported is behaviour, and the behaviour differs between the two vendors in a way a single sentence would flatten: one publishes its component list in a field the standard interface does not surface, the other publishes none that we could find. Those are different facts.

Nor do we suggest the difference is sinister. The standard interface predates composed systems and has no field for this, which is a sufficient explanation and not the only possible one. We have not tested which explanation is true and we have no instrument that could. What was measured is the gap between what a system reports and what it did. That gap is the same size whatever produced it.

7.Limits, and one finding we nearly reported backwards

Three systems, four questions, five draws each. This establishes that these things happen. It does not establish how often, and the intervals around every rate here are wide enough that none is published.

It is also a single census. The question this exercise exists to answer — whether disclosure improves — cannot be answered until there are three, and the measurement has been built so that the next one is identical to this one. A comparison between two runs is valid only if they were the same run.

Score these systems as stable or varied and the one disclosing nothing comes out the most stable of the three, because it returns an identical model name every time. It is in fact the only one whose stability cannot be observed at all. A verdict with two values would have ranked the least auditable system best. We report a third value — unknowable — and treat it as an answer rather than a missing one.

Even there, something leaks. That system’s composition is invisible, but its token counts move on identical input: on the coding question, five identical asks consumed between 4,447 and 15,969 tokens. Variable internal work is provable from usage even when attribution is impossible. It is a weaker claim than the one made about the disclosed systems, and it is still a claim.

8.What was not tested

There is a persistent idea that orchestration is where certain national research programmes are doing their most interesting work. We cannot speak to it. The systems measured here are American and commercially available, and were reachable because their endpoints answer. Where work is a method rather than a product, no benchmark reaches it, ours included. That absence is a fact about benchmarking rather than about the researchers.

Every figure in this paper was read from the responses themselves; none was estimated. One number that could have been published was not: one vendor reports a cost field in units it does not document, and a plausible assumption about those units yields plausible-looking prices, which is what makes quoting them unsafe.

SILT-RP-007 · version 1.2 · issued 2026-09-24
© 2026 Sentient Index Labs & Technology. This document is licensed CC BY-ND 4.0 — share it freely and in full, with attribution; do not publish modified versions. Translations and excerpts are granted on request to press@sentientindexlabs.com.
⚠️ Corrections are issued as a new version under the same identifier, each with a dated note in the text recording exactly what changed and what the earlier version said.
Nothing in this document is a certification, audit opinion, conformity assessment, or legal advice. See the Subscriber Agreement, Section 15.