Splitting research from writing
A cheaper model runs the search tools and Gemini 3.7 Flash writes the answer. Three research models against today’s single-model agent (new tools, prompt v29). 100 real questions, 112 turns, 11 Oct 2026.
No research model is worth shipping as the default today. The two cheap ones cut cost by about half but lose clearly on answer quality: GPT-6 Luna won 14 of 83 decided cases and GLM 5.3 won 25 of 84. DeepSeek V4.1 Flash keeps quality (46–32, not significant, and partly a matter of its longer answers) but saves only 16% while adding 13 s before the first word.
The split itself works. With Gemini doing both jobs there is no significant overall quality difference (31–36), and citations stay valid, so the code can ship behind its flag, which is off by default. The cheap models fall short in the research step, not in the handoff: they write weaker queries and find fewer of the right passages.
How the split works
Off by default. CHAT_RESEARCH_MODEL picks the research model, and a per-request option lets the benchmark choose. Prompts are new Langfuse entries, platform/chat-research and platform/chat-writer, labelled latest only; reviewed copies live in the repo.
Results per arm
| Arm | Cost per turn research + writer + rerank | First word p50 / p95 | Done p50 / p95 | Win rate vs baseline decided, 95% CI | Supported claims | Serious errors | Invented citations |
|---|---|---|---|---|---|---|---|
| BaselineGemini runs tools and writes | $0.053— + 0.045 + 0.008 | 14.9 / 32.3 s | 29.0 / 57.6 s | — | 72.5% | 6 | 7 |
| GPT-6 LunaAzure, reasoning low | $0.024 −54% · 0.004 + 0.013 + 0.007 | 15.8 / 24.6 s | 21.8 / 34.8 s | 17% 14–69 · CI 10–26% | 74.2% | 8 | 0 |
| DeepSeek V4.1 FlashOpenRouter → DeepSeek, low | $0.045 −16% · 0.014 + 0.019 + 0.011 | 28.2 / 54.7 s | 36.2 / 67.6 s | 59% 46–32 · CI 48–69% | 73.5% | 1 | 1 |
| GLM 5.3 FlashOpenRouter → Fireworks, low | $0.028 −47% · 0.008 + 0.015 + 0.005 | 24.8 / 41.8 s | 32.2 / 54.7 s | 30% 25–59 · CI 21–40% | 72.4% | 9 | 1 |
| Gemini → Geminicontrol: same split, no cheap model | $0.063 +18% · 0.040 + 0.015 + 0.008 | 22.1 / 43.7 s | 29.1 / 51.2 s | 46% 31–36 · CI 35–58% | 74.8% | 4 | 0 |
Per turn, means for cost, 112 turns per arm; every turn completed in every arm. Win rate counts a case only when the blinded judge picks the same winner in both answer orders (100 cases). Sign-test p: Luna and GLM below 0.001; DeepSeek 0.14; control 0.63. “Invented citations” is the number of answers with a citation that does not resolve to a retrieved passage (for GLM, refs left unmapped). Supported claims are within ±2.4 points of the baseline for every arm; only the control’s +2.3 points has a CI that excludes zero (+0.3 to +4.2).
Does more reasoning help?
| Research model | Effort | Cost per turn | First word p50 / p95 | Win rate vs baseline | At similar length |
|---|---|---|---|---|---|
| GPT-6 Luna | low | $0.024 | 15.8 / 24.6 s | 17% 14–69 | 8–25 |
| medium | $0.035 | 24.4 / 46.6 s | 25% 18–55 · CI 16–36% | 9–17 | |
| xhigh | $0.039 | 33.3 / 70.5 s | 37% 29–50 · CI 27–48% | 15–16 | |
| DeepSeek V4.1 Flash | low | $0.045 | 28.2 / 54.7 s | 59% 46–32 · p 0.14 | 24–18 |
| medium | $0.044 | 29.1 / 50.0 s | 62% 46–28 · CI 51–72% · p 0.047 | 19–15 | |
| xhigh | $0.052 | 41.5 / 79.2 s | 65% 48–26 · CI 54–75% · p 0.014 | 22–17 | |
| GLM 5.3 Flash | low | $0.028 | 24.8 / 41.8 s | 30% 25–59 | 12–28 |
| medium | $0.030 | 24.7 / 44.2 s | 18% 14–63 · CI 11–28% | 9–37 | |
| xhigh | $0.035 | 33.3 / 60.6 s | 34% 24–47 · CI 24–45% | 12–26 |
Follow-up run on the same 100 cases: research effort medium and xhigh, writer unchanged, research timeout raised to 120 s, judged against the same baseline answers ($0.053, first word 14.9 s; a fresh baseline run alongside measured the same 14.9 s). More effort helps Luna and GLM a little but neither catches up; most of Luna’s remaining gap is shorter answers (15–16 at similar length on xhigh). DeepSeek gets significantly better than the baseline at medium and xhigh, but its wins are mostly longer answers (at similar length 19–15 and 22–17, not significant), xhigh costs as much as the baseline, and every setting adds 13–27 s before the first word. Claim audit was not repeated for these arms.
What goes wrong
- Cheap models find less of the right evidence. Luna’s research found 44% of the passages the baseline went on to cite, GLM 35%, DeepSeek 60%. Gemini researching for itself in the split found 61%. Luna writes queries as paraphrased questions where Gemini writes the wording a source would use; a prompt rule asking for source wording did not close the gap on the dev set. Answers lose mainly on completeness: Luna 11–74, GLM 18–56.
- The handoff is not the main problem. On the dev set, giving the writer every passage Luna retrieved instead of its selection still lost (3–11). The same split with Gemini researching ties the baseline. Luna also loses on follow-ups (0–10), questions that name a source (9–43) and questions that need the full text (6–23).
- DeepSeek’s edge is partly length. It wins 20–5 when its answer is longer than the baseline’s and 24–18 (p = 0.44) when lengths are similar. Its research takes a median 19 s: each model round on DeepSeek’s endpoint takes 4–5 s with reasoning on low, and it runs 4.4 searches per turn against 3.1. In 19 of 112 turns it ended without submitting, usually by calling search again in the final “submit only” round, so the server picked its passages instead.
- Provider routing traps. OpenRouter drops endpoints that cannot force a tool call, including DeepSeek’s and Z.ai’s own, and silently sends the request to slow third parties instead. Z.ai’s GLM endpoint took 12 s per round even with reasoning on low, and with default thinking all 6 test runs hit our 150 s timeout. GLM only became usable on Fireworks.
- Copying ids. With full UUIDs, 6 of the 8 ids Luna selected in a smoke test did not exist. Short 8-character ids cut unresolved selections to about 8% for Luna and 1–3% for DeepSeek and GLM. The server drops the rest and tops the selection up to 10 passages.
- Splitting does not save money by itself. With Gemini on both sides the turn costs 18% more and the first word comes 7 s later. Savings come only from a cheaper model reading the long tool results.
Recommendation
- Keep the single-model agent as the default. Merge the split behind its flag (off) so it can be tested in production, without rolling it out.
- Don’t use Luna or GLM for research. They save $0.025–0.029 per turn at low effort but lose about 70–83% of decided comparisons, and medium or xhigh effort does not close the gap (best: Luna xhigh, 37%).
- DeepSeek is the only candidate, and not yet. Quality holds, and at medium effort it beats the baseline (62%, p 0.047, mostly longer answers); the audit at low effort counted 1 serious error against 6 for the baseline (small numbers, p ≈ 0.12). But the saving is at most $0.009 per turn for 13–14 s more wait. Try it only if latency can be cut, for example with fewer rounds, a faster host, or by streaming research progress.
- Cheap win regardless: the citation guard that maps refs and drops undelivered ids kept invented citations to 2 answers across all four split arms, against 7 for the baseline. One was a fallback turn that bypasses the guard, and one GLM answer left bare S-refs outside citation tags. The guard could also protect the single-model path.
Method and caveats
- Cases. The same 100 hand-picked production questions as the v29 benchmark: 50 Arabic and 50 English, 62 naming a source, 34 needing the full text, 11 multi-turn follow-ups, for 112 turns. Prompts were tuned on a separate dev set of 22 other production questions plus a made-up greeting and app question, never on these.
- Arms. All arms ran at the same time through the real engine path on this branch, with Gemini 3.7 Flash on Vertex at the same thinking level and the same tools and reranker. The split arms used
platform/chat-researchv1 andplatform/chat-writerv1. The baseline is the single-model path withplatform/main-geminiv29. Each case ran once per arm. - Judging. Claude Opus compared each arm with the baseline without labels, in both orders: 800 verdicts, with 71–89% order agreement. A separate Opus auditor checked the claims in all five answers to each case against the passages each answer cited: 7,769 claims. This audit used shorter excerpts than the v29 benchmark, so its supported-claim rates are not comparable with that report.
- Cost. Measured tokens times list prices. Gemini: $0.75/M input, $0.075/M cached, $3.75/M output. Luna: $0.10/$0.01/$0.50 per M. DeepSeek and GLM: the cost OpenRouter actually billed. Cohere rerank: $0.0025 per search. Turbopuffer and embeddings are not priced. Latency excludes the Cloudflare gateway hop.
- Noise. Intervals are Wilson for win rates and paired bootstrap for differences. Luna loses in every slice; GLM in every slice except hadith takhrīj (4–3, n 9) and quotations (3–3, n 8). DeepSeek’s lead is not significant.