10 stories tagged with #inference, in publish-time order across the WeSearch catalog. Tag pages update as new stories ingest.
⌘ RSS feed for this tag → or search "Inference"
LLM Inference vs. the OOM Killer
Mobile memory warnings and handling them in Rust.…
Bypassing inference bottlenecks: Accelerating complex AI search
The KV Cache Tax: Why Inference Servers Run Out of Memory Before Compute
A VRAM budget formula for LLM serving, and three optimization strategies mapped to the traffic patterns that trigger the OOM.…
The Inference Hardware Revolution of 2026
Today’s tidal wave of queries is forcing hardware makers to pivot…
Sovereign: A Unified GPU Inference Substrate (Fractal Memory, Manifold Routing)
All of my Whitepapers. Contribute to CuppaTea1983/Sovereign development by creating an account on GitHub.…
Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs (Bryan Shan/SemiAnalysis)
By Bryan Shan / SemiAnalysis. View the full context on Techmeme.…
RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems
AI is beginning to make substantive contributions to LLM inference optimization. Existing AI optimizations are predominantly profiling-based. Profiling-bound feedback confines the …
Categorizing AI inference: Initialization, Reasoning, Orchestration, Synthesis
A framework for understanding where LLM inference goes during agentic tasks: Initialization, Reasoning, Orchestration, and Synthesis. Consider "Reasoning Yield" as a way to measure…
NanoGPT-inference LLM inference from scratch
LLM inference from scratch…
Show HN: Determinstic LLM inference for lowest price Gemma 4, with Windows XP
Open-weight models, fully deterministic. The same answer every time, byte for byte, at the floor price.…