Model Disclosure in Composed AI Systems
Census #1 of the Model Disclosure Index (MDI) programme: whether routers and orchestrators name the model that answered
1.Abstract
A composed system routes a request among several models rather than answering with one, and they are increasingly what enterprises deploy. This census asks something narrow that appears not to have been measured: when you ask one a question, does it tell you what answered? Across three commercially available systems, the model field every portable client reads returned a name belonging to no model at all. One answered a short question using three invocations of two different companies' models and disclosed the fact only in a vendor extension outside the standard interface; another disclosed nothing anywhere. Asked the same question five times, one system used five different combinations of models costing between 3,558 and 13,008 tokens, and four blind judges scored the most expensive draw the lowest of the five. The paper reports what an audit log therefore does not contain, states the limits of a single census, and declines — in both directions — to draw any conclusion about intent.
2.What was measured, and what was not
A composed system takes a request and distributes it among several models rather than answering with one. They are cheaper and often better, which is why they are increasingly what organisations actually deploy. This census asks one narrow question about them: when you ask one a question, does it tell you what answered?
Three systems were probed: two from groq and one from xAI. Four kinds of question — a one-word lookup, a short coding task, a multi-step arithmetic problem and an open-ended research question — were each asked five times, and the composition serving every draw was recorded. Sixty-three calls in total: sixty for the repeated questions, and three single-call disclosure probes, one per system.
3.Disclosure: the model field names a thing that is not a model
In all three systems the model field — the only field a portable OpenAI-compatible client reads — returned the name of the composed system itself. Asked a short question about load-bearing walls, one system produced its answer using three separate model invocations: two of Meta’s Llama 4 Scout and one of OpenAI’s gpt-oss-120b. That fact appeared only in a vendor-specific extension the standard interface does not surface.
The second groq system behaved the same way with two components. xAI’s multi-agent system disclosed nothing at all: its response named the system, its usage block reported how many tools and sources were used and what the call cost, and nowhere did it say which model produced the text.
For anyone operating under a data processing agreement, a subprocessor disclosure obligation, or a contractual term about where information goes, that is a gap between what is required to be known and what can be found out.
A second finding concerns the interface rather than the answer. xAI’s multi-agent model is refused outright on the chat-completions endpoint every other model here accepts, returning an error indistinguishable from a dead API key. It answers on a different endpoint. Across two vendors, then, the portable interface either conceals the composition or excludes the composed system from itself.
4.Stability: the same question, five different machines
On the one-word question, one system used the same two invocations in all five draws. On the arithmetic question, five identical asks produced five different compositions, ranging from four model invocations to eleven. Not one draw matched another.
The instability is not spread evenly. It concentrates where the work is hard, which is precisely where a buyer would most want to know what they were getting. A test using a single easy prompt lands on the quiet end of that range and reports stability.
v1.1 (24 September 2026). Version 1.0 said “as high as twenty-eight per cent”. That is the upper bound at 80% confidence, given without saying so; at the 95% level used everywhere else in our work it is about 45%. The correction widens the bound and strengthens the point it was making. Nothing else in this paper changed.
v1.2 (24 September 2026). Three statements checked against the census data. Section 2 said “Sixty-three calls in total” without saying the three beyond 3 × 4 × 5 were single-call disclosure probes, one per system; it now says so. Section 5 described the pilot as “one prompt, one system”; it was one system and three prompts, five draws each, which is where its fifteen answers come from. And “nothing here is scored on a scale” is narrowed to what is true: no system is scored or ranked; the pilot scores answers.
Cost follows composition. On the arithmetic question, with identical input, total tokens consumed ranged from 3,558 to 13,008 — nearly a factor of four for the same question, decided by a process the customer cannot observe and is not told about.
5.Did the extra work buy anything?
The five answers to the arithmetic question were already recorded, each produced by a different composition. They were shown to four independent judges who saw only the question and the answer — never the composition, never the invocation count, never each other’s scores — in shuffled order.
- Four invocations scored 7.0
- Five scored 10.0
- Nine scored 7.5
- Ten scored 8.5
- Eleven scored 6.5
That result is a pilot and is reported as one. It is one system, three prompts and five draws of each; the scores above are the arithmetic prompt's five. It cannot support a rate and none is offered. What it supports is the narrow claim a buyer would care about: on this question, on this system, the run costing nearly four times as much was not better, and by this panel it was slightly worse.
6.What we decline to conclude
An earlier draft of this material stated that none of this was deception. That sentence has been removed, and the reason is worth stating because it is easy to miss.
So neither claim is made. What can be reported is behaviour, and the behaviour differs between the two vendors in a way a single sentence would flatten: one publishes its component list in a field the standard interface does not surface, the other publishes none that we could find. Those are different facts.
Nor do we suggest the difference is sinister. The standard interface predates composed systems and has no field for this, which is a sufficient explanation and not the only possible one. We have not tested which explanation is true and we have no instrument that could. What was measured is the gap between what a system reports and what it did. That gap is the same size whatever produced it.
7.Limits, and one finding we nearly reported backwards
Three systems, four questions, five draws each. This establishes that these things happen. It does not establish how often, and the intervals around every rate here are wide enough that none is published.
It is also a single census. The question this exercise exists to answer — whether disclosure improves — cannot be answered until there are three, and the measurement has been built so that the next one is identical to this one. A comparison between two runs is valid only if they were the same run.
Even there, something leaks. That system’s composition is invisible, but its token counts move on identical input: on the coding question, five identical asks consumed between 4,447 and 15,969 tokens. Variable internal work is provable from usage even when attribution is impossible. It is a weaker claim than the one made about the disclosed systems, and it is still a claim.
8.What was not tested
There is a persistent idea that orchestration is where certain national research programmes are doing their most interesting work. We cannot speak to it. The systems measured here are American and commercially available, and were reachable because their endpoints answer. Where work is a method rather than a product, no benchmark reaches it, ours included. That absence is a fact about benchmarking rather than about the researchers.
Every figure in this paper was read from the responses themselves; none was estimated. One number that could have been published was not: one vendor reports a cost field in units it does not document, and a plausible assumption about those units yields plausible-looking prices, which is what makes quoting them unsafe.