r/LocalLLaMA
social · source
r/LocalLLaMA on WeSearch
Recent social headlines from r/LocalLLaMA.
r/LocalLLaMA
Gemma 4 2B handling structured JSON output + tool calling + reasoning traces correctly via Spring AI / LM Studio — including identifying a real Java bug in code review
r/LocalLLaMA
Qwen3.6-35B-A3B vs Gemma4-26B-A4B
r/LocalLLaMA
Measuring AI intelligence vs Human intelligence
r/LocalLLaMA
gemma 4 e2b quality degrades after ~30-40 continuous inferences on 4gb vram?
r/LocalLLaMA
Qwen Plays ̶p̶̶o̶̶k̶̶e̶̶m̶̶o̶̶n̶ ? / QWEN PLAYS DCSS! - qwen3.6-35b-a3b@q4_k_xl plays open source roguelike adventure DCSS (and does a decent job)
r/LocalLLaMA
Frustrating results with product searching
r/LocalLLaMA
Why not dynamic active parameters (and other questions for the knowledgeable)
r/LocalLLaMA
Choosing an abliterated version of Gemma 4 31B and 26B-A4B
r/LocalLLaMA
Qwen3.6-35B-A3B-Uncensored-Genesis-APEX-MTP
r/LocalLLaMA
I built a local GUI for the TradingAgents framework — works with Ollama
r/LocalLLaMA
Anyone down to test this? Just uploaded a model using rys
r/LocalLLaMA
TTS Benchmark Comparison (all known TTS up until May 2026)
r/LocalLLaMA
Performance When Offloading Large Models to System RAM?
r/LocalLLaMA
How are you all handling agents and sub agents?
r/LocalLLaMA
Is there any reason for an uncensored model if you have no interest in roleplaying?
r/LocalLLaMA
Vision-capable LLMs vs. OCR for long-document (including charts, images, tables, etc.) QA
r/LocalLLaMA
minor speed bump for MTP with Qwen3.6-27B-MTP Q6_K_XL
r/LocalLLaMA
llampart 1.0.0 - I released a standalone local web UI for llama-server with translations, extended settings and a polished conversation sidebar
r/LocalLLaMA
How to keep up to date on latest models?
r/LocalLLaMA
llama.cpp server have built-in native tools (exec_shell, edit_file, etc.)
r/LocalLLaMA
Local model doing accounting tasks
r/LocalLLaMA
MLID claims nova lake-ax not cancelled just renamed razor lake-ax
r/LocalLLaMA
For users have have both 6000 PRO MaxQ and Workstation Edition (or Server Edition), how much slower is the MaxQ vs the WS/SV on compute? (Prompt processing, Diffusion, etc)
r/LocalLLaMA
Command A+ (218B MoE) running on Apple Silicon — MLX port, PR open
r/LocalLLaMA
Inference provider tiers by Cache-hit rates, using openrouter data
r/LocalLLaMA
Any reason to run dense over MOE for RAGs?
r/LocalLLaMA
$16 refactor, 400 steps, 95% routed to open MoE
r/LocalLLaMA
7900XTX idle power draw when running headless?
r/LocalLLaMA
Local, low code, node based agentic development workspace... that actually works?
r/LocalLLaMA
Qwen3.6 35B-A3B MTP hits 249 t/s on a 24GB consumer GPU (RTX 5090M) — 3.4× the dense 27B variant on the same image
r/LocalLLaMA
found this little known channel with some really good content
r/LocalLLaMA
First AI to Beat Every Human in a Programming Competition - Agentic GRPO Explained
r/LocalLLaMA
Have we passed the peak of inflated expectations?
r/LocalLLaMA
DGX Spark agentic usage numbers
r/LocalLLaMA
Best open-source & proprietary options for Indic language ASR
r/LocalLLaMA
LLaMa.cpp basic question
r/LocalLLaMA
Gemma4 26b a4b Apex quant is quite good
r/LocalLLaMA
Gemma is so much better than Qwen, prove me wrong
r/LocalLLaMA
G4-MeroMero-26B-A4B-it-uncensored-heretic Is Out Now, a Finetune of gemma-4-26B-A4B-it, With KLD of 0.0152 and 12/100 Refusals!
r/LocalLLaMA
Qwen3.6-35B-A3B Q4 262k context on 8GB 3070 Ti = +30tps
r/LocalLLaMA
NVIDIA Removes Gaming Revenue Category From Financial Reports
r/LocalLLaMA
How small can the orchestration model in an agent be? (separating it from code-gen — that obviously wants a big model)
r/LocalLLaMA
If one .gguf makes it past the great filter, humanity survives in some way.
r/LocalLLaMA
Seeking resources to read about llama.cpp server and how offloading works
r/LocalLLaMA
OpenBMB presents the model BitCPM-CANN 1.58 bit
r/LocalLLaMA
Holding machine upgrade waiting for a model?
r/LocalLLaMA
Quick note on sudden performance loss when running GGUFs
r/LocalLLaMA
ztok — a fast multithreaded tokenizer in Zig that loads tiktoken / HF / SentencePiece and is 2–5× faster
r/LocalLLaMA
New Release of ROCm based MLX LLM Engine - lemon-mlx-engine
How WeSearch handles this source
WeSearch's declared handling of r/LocalLLaMA's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.
Indexing: Allowed
Snippet: Allowed
AI summary: Limited
Retrieval / RAG: Not asserted
Model training: Not asserted
Commercial reuse: Not permitted