QafAgent benchmark

Splitting research from writing

A cheaper model runs the search tools and Gemini 3.7 Flash writes the answer. Three research models against today’s single-model agent (new tools, prompt v29). 100 real questions, 112 turns, 11 Oct 2026.

Don’t switch yet

No research model is worth shipping as the default today. The two cheap ones cut cost by about half but lose clearly on answer quality: GPT-6 Luna won 14 of 83 decided cases and GLM 5.3 won 25 of 84. DeepSeek V4.1 Flash keeps quality (46–32, not significant, and partly a matter of its longer answers) but saves only 16% while adding 13 s before the first word.

The split itself works. With Gemini doing both jobs there is no significant overall quality difference (31–36), and citations stay valid, so the code can ship behind its flag, which is off by default. The cheap models fall short in the research step, not in the handoff: they write weaker queries and find fewer of the right passages.

How the split works

User turn + history Research model lookup_catalog · search · expand sees 8-character passage ids 3–6 rounds max, 60 s timeout → submit_research Server handoff resolve ids to retrieved passages verbatim text as S1…Sn 10–20 when available + notes, gaps, searches, status Gemini writer no tools, writer prompt cites S refs only S refs mapped back to ids, anything else dropped One streamed message: research tool parts (search labels, passages) then the writer’s text same shape as today, so citations, history and follow-ups need no client change no submission → server picks top retrieved passages error with no evidence → old Gemini tool loop greeting or app question: 1 call

Off by default. CHAT_RESEARCH_MODEL picks the research model, and a per-request option lets the benchmark choose. Prompts are new Langfuse entries, platform/chat-research and platform/chat-writer, labelled latest only; reviewed copies live in the repo.

Results per arm

ArmCost per turn
research + writer + rerank
First word
p50 / p95
Done
p50 / p95
Win rate vs baseline
decided, 95% CI
Supported claimsSerious errorsInvented citations
BaselineGemini runs tools and writes$0.053— + 0.045 + 0.00814.9 / 32.3 s29.0 / 57.6 s—72.5%67
GPT-6 LunaAzure, reasoning low$0.024 −54% · 0.004 + 0.013 + 0.00715.8 / 24.6 s21.8 / 34.8 s17% 14–69 · CI 10–26%74.2%80
DeepSeek V4.1 FlashOpenRouter → DeepSeek, low$0.045 −16% · 0.014 + 0.019 + 0.01128.2 / 54.7 s36.2 / 67.6 s59% 46–32 · CI 48–69%73.5%11
GLM 5.3 FlashOpenRouter → Fireworks, low$0.028 −47% · 0.008 + 0.015 + 0.00524.8 / 41.8 s32.2 / 54.7 s30% 25–59 · CI 21–40%72.4%91
Gemini → Geminicontrol: same split, no cheap model$0.063 +18% · 0.040 + 0.015 + 0.00822.1 / 43.7 s29.1 / 51.2 s46% 31–36 · CI 35–58%74.8%40

Per turn, means for cost, 112 turns per arm; every turn completed in every arm. Win rate counts a case only when the blinded judge picks the same winner in both answer orders (100 cases). Sign-test p: Luna and GLM below 0.001; DeepSeek 0.14; control 0.63. “Invented citations” is the number of answers with a citation that does not resolve to a retrieved passage (for GLM, refs left unmapped). Supported claims are within ±2.4 points of the baseline for every arm; only the control’s +2.3 points has a CI that excludes zero (+0.3 to +4.2).

Does more reasoning help?

Research modelEffortCost per turnFirst word p50 / p95Win rate vs baselineAt similar length
GPT-6 Lunalow$0.02415.8 / 24.6 s17% 14–698–25
medium$0.03524.4 / 46.6 s25% 18–55 · CI 16–36%9–17
xhigh$0.03933.3 / 70.5 s37% 29–50 · CI 27–48%15–16
DeepSeek V4.1 Flashlow$0.04528.2 / 54.7 s59% 46–32 · p 0.1424–18
medium$0.04429.1 / 50.0 s62% 46–28 · CI 51–72% · p 0.04719–15
xhigh$0.05241.5 / 79.2 s65% 48–26 · CI 54–75% · p 0.01422–17
GLM 5.3 Flashlow$0.02824.8 / 41.8 s30% 25–5912–28
medium$0.03024.7 / 44.2 s18% 14–63 · CI 11–28%9–37
xhigh$0.03533.3 / 60.6 s34% 24–47 · CI 24–45%12–26

Follow-up run on the same 100 cases: research effort medium and xhigh, writer unchanged, research timeout raised to 120 s, judged against the same baseline answers ($0.053, first word 14.9 s; a fresh baseline run alongside measured the same 14.9 s). More effort helps Luna and GLM a little but neither catches up; most of Luna’s remaining gap is shorter answers (15–16 at similar length on xhigh). DeepSeek gets significantly better than the baseline at medium and xhigh, but its wins are mostly longer answers (at similar length 19–15 and 22–17, not significant), xhigh costs as much as the baseline, and every setting adds 13–27 s before the first word. Claim audit was not repeated for these arms.

What goes wrong

Recommendation

Method and caveats