Zhengkid/AutoTTS: Agentic Discovery for Test-Time Scaling
AutoTTS introduces an agentic approach to test-time scaling in large language models by automating the discovery of inference controllers through a replay-based environment. The method eliminates the need for hand-crafted heuristics and gradient updates, relying instead on a coding agent that iteratively refines code-defined controllers. Evaluated on AIME and HMMT benchmarks, the discovered Confidence Momentum Controller achieves competitive accuracy with significant token savings.
- ▪AutoTTS uses a coding agent to automatically discover inference controllers in a replay environment without LLM calls during evaluation.
- ▪The discovered Confidence Momentum Controller reduces token usage by ~69.5% compared to SC@64 while maintaining similar accuracy across multiple model scales.
- ▪The entire discovery process costs an estimated $39.9 and takes 160 minutes, with zero LLM calls during evaluation due to cached replays.
- ▪Controllers are trained on AIME24 data and generalize well to held-out AIME25 and HMMT25 benchmarks.
- ▪The system enables fine-grained policy improvements by recording full execution traces and scaling curves for iterative refinement.
Hacker News (Newest) files mainly under programming. We currently carry 5,306 of its stories.
Story provenance
Source · retrieval · rights · ranking — open for full record
inspect →
Story provenance
Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.
Record
| Original publisher | GitHub |
| Canonical URL | https://github.com/zhengkid/AutoTTS |
| Publication time | Sun, 17 May 2026 02:01:16 +0000 |
| Retrieval time | 2026-05-17T02:10:19.091Z |
| Last seen | 2026-05-17T02:10:19.091Z |
| Headline source | Publisher (no WeSearch rewrite) |
| Excerpt source | publisher body |
| Excerpt method | First ~120 words (~800 chars) of extracted publisher body, fair-use limited. |
| Summary | WeSearch · cerebras-chat (WeSearch summarizer) |
| Summary source text | contentText |
| Citation coverage | Summary is a WeSearch-generated derivative; primary citation is the original publisher URL. |
| Cluster | LGHGMT3SFNIQ |
| Cluster logic | Grouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison. |
| Ranking reason | Story pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking. |
| Publisher visit | Yes — open original |
| Substitutes article? | No — link-out required for full text |
Rights status (four layers)
WeSearch handling by dimension
| Indexing | May the item be indexed (stored, ranked, made findable)? | Allowed |
| Snippet | May a short excerpt of the publisher's text be shown? | Allowed |
| AI summary | May WeSearch generate its own short summary of the article? | Limited |
| Retrieval / RAG | May the content be exposed for third-party retrieval-augmented generation? | Not asserted |
| Model training | May the content be used to train AI models? | Not asserted |
| Commercial reuse | May the content be reused commercially? | Not permitted |
Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.
Opening excerpt (first ~120 words) tap to expand
AutoTTS LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling Tong Zheng, Haolin Liu, Chengsong Huang, Huiwen Bao, Sheng Zhang, Rui Liu, Runpeng Dai, Ruibo Chen, Chenxi Liu, Tianyi Xiong, Xidong Wu, Hongming Zhang, Heng Huang UMD · UVA · WUSTL · UNC · Google · Meta Project page AutoTTS reframes TTS strategy design from hand-crafting heuristics to environment-driven automatic search: humans only construct an offline replay environment (states, actions, feedback, objectives), and a coding agent iteratively proposes and refines code-defined controllers within it — code edits, no gradient updates. Cheap: 0 LLM calls, fully replay.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at GitHub.