Reducing Hallucinations in LLM-Gen. Code via Semantic Triangulation (OOPSLA 26)
The paper introduces semantic triangulation as a method to reduce hallucinations in code generated by large language models. It involves transforming the original problem into a dissociative version and cross‑examining solutions using a bijection‑inducing relation. Experiments on CodeElo and LiveCodeBench show that this approach outperforms simple plurality voting in identifying correct programs.
- ▪LLM‑generated code often contains hallucinated bugs that are hard to detect because expected behavior is rarely formally specified.
- ▪Semantic triangulation uses a dissociative problem transformation and a hyperproperty relation to create an independent witness that reveals inconsistencies in hallucinated solutions.
- ▪The method ensures that correct solutions map to correct solutions across the original and transformed problems, providing higher confidence than majority voting.
- ▪The authors implement transformations such as partial inversion, answer enumeration, and problem decomposition, and evaluate the approach on benchmark datasets.
- ▪Theoretical analysis proves that agreement with a triangulated witness yields strictly higher correctness confidence than plurality voting.
Hacker News (AI / LLM) files mainly under ai. We currently carry 4,153 of its stories.
Story provenance
Source · retrieval · rights · ranking — open for full record
inspect →
Story provenance
Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.
Record
| Original publisher | GitHub |
| Canonical URL | https://github.com/msv-lab/just-tri-it |
| Publication time | Sun, 09 Aug 2026 10:42:39 +0000 |
| Retrieval time | 2026-08-09T10:50:42.314Z |
| Last seen | 2026-08-09T10:50:42.314Z |
| Headline source | Publisher (no WeSearch rewrite) |
| Excerpt source | publisher body |
| Excerpt method | First ~120 words (~800 chars) of extracted publisher body, fair-use limited. |
| Summary | WeSearch · cerebras-chat (WeSearch summarizer) |
| Summary source text | contentText |
| Citation coverage | Summary is a WeSearch-generated derivative; primary citation is the original publisher URL. |
| Cluster | _J5saXZcTFRp · 1 stories |
| Cluster logic | Grouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison. |
| Ranking reason | Story pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking. |
| Publisher visit | Yes — open original |
| Substitutes article? | No — link-out required for full text |
Rights status (four layers)
WeSearch handling by dimension
| Indexing | May the item be indexed (stored, ranked, made findable)? | Allowed |
| Snippet | May a short excerpt of the publisher's text be shown? | Allowed |
| AI summary | May WeSearch generate its own short summary of the article? | Limited |
| Retrieval / RAG | May the content be exposed for third-party retrieval-augmented generation? | Not asserted |
| Model training | May the content be used to train AI models? | Not asserted |
| Commercial reuse | May the content be reused commercially? | Not permitted |
Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.
Opening excerpt (first ~120 words) tap to expand
Reducing Hallucinations in LLM-Generated Codevia Semantic Triangulation Yihan Dai, Sijie Liang, Haotian Xu, Peichu Xie, Sergey Mechtaev arXiv:2511.12288 LLM-generated code often contains hallucinated bugs, and since expected behavior is rarely formally specified, they are hard to detect automatically. Identifying which, if any, of the sampled programs are correct is akin to a police detective questioning suspects. Because LLMs make correlated errors, most suspects have colluded on the same fake alibi — so plurality (majority) voting does not identify the truth; it merely amplifies their shared deception. Previous methods bring in extra witnesses: LLM-generated tests, or specifications auto-formalized from the problem description (e.g., Hoare-style postconditions).
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at GitHub.