WeSearch

Show HN: WorldBuild Bench repo: testing LLM world coherence with 3D games

·7 min read · 0 reactions · 0 comments · 34 views
#show#worldbuild#bench#repo#testing
Show HN: WorldBuild Bench repo: testing LLM world coherence with 3D games
TL;DR · WeSearch summary

WorldBuild Bench Same brief, same tools, same harness. Build a 3D game someone would actually want to keep playing. One brief: Sunset Apex, a 3-lap circuit racer.

Key facts
Original article
GitHub
Read full at GitHub →
Opening excerpt (first ~120 words) tap to expand

WorldBuild Bench Same brief, same tools, same harness. Build a 3D game someone would actually want to keep playing. One brief: Sunset Apex, a 3-lap circuit racer. Three models. Three completely different games. Real, unedited gameplay. All 27 builds from the round are playable in your browser right now. ▶ Play the builds · Vote in the Arena · Methodology · Round data Why this exists Coding benchmarks usually stop at "does it run." As models move toward "world models," a harder question is spatial, temporal, and causal coherence in a 3D space: does the model understand where things are, stay consistent over time, and when something happens, do the consequences make sense? Those qualities are hard to capture with static benchmark questions.

Excerpt limited to ~120 words for fair-use compliance. The full article is at GitHub.

Anonymous · no account needed
Share 𝕏 Facebook Reddit LinkedIn Threads WhatsApp Bluesky Mastodon Email

Discussion

0 comments

More from GitHub