r/LocalLLaMA
social · source
r/LocalLLaMA on WeSearch
Recent social headlines from r/LocalLLaMA.
r/LocalLLaMA
Low-level coding dataset
r/LocalLLaMA
Anyone evaluated the difference between Qwen Code for the local qwen models vs another harness? CC, OC, LC, Aider etc..
r/LocalLLaMA
What model weights (quantized included) under 150GB have the best general knowledge depth?
r/LocalLLaMA
When your LLM treats data center GPUs like an optional DLC
r/LocalLLaMA
Latest b9274 Addresses MTP VRAM leak
r/LocalLLaMA
Waiting for Qwen 3.7 open weight... The new King has arrived...
r/LocalLLaMA
Gorgon Halo is 6.7% faster than predecessor Strix Halo
r/LocalLLaMA
Strix Halo 128GB vs M5 pro 64GB
r/LocalLLaMA
Honesty in a small model drops from 35% to 0% by changing the tone of the prompt. Sharing the findings.
r/LocalLLaMA
110 tok/s with 12GB VRAM on Qwen3.6 35B A3B and ik_llama.cpp
r/LocalLLaMA
Open-source LLMs are still weak against long reasoning jailbreaks, even with lightweight defenses
r/LocalLLaMA
Model Golf for some Runpod Credits!
r/LocalLLaMA
Back again, many changes have taken place.
r/LocalLLaMA
How can you stop your model from looping
r/LocalLLaMA
"AWS secures rare Mac Studios while ordinary Apple customers remain completely locked out"
r/LocalLLaMA
Guide to building smoltorrent | A Distributed ML Checkpoint Storage System
r/LocalLLaMA
What small speech to text (STT) model is best at recognizing whispered speech?
r/LocalLLaMA
Gemma 4 MTP with LlamaCPP
r/LocalLLaMA
Impulse Purchase.
r/LocalLLaMA
Qwen3.7 Max scored by Artificial Analysis, 27B/35B waiting room
r/LocalLLaMA
Guardrails take an 8B model from 53% to 99% on agentic tasks [ACM CAIS '26 preprint]
r/LocalLLaMA
Show HN: Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks
Reddit
LM Studio finally added support for MTP Speculative Decoding
r/LocalLLaMA
Claude Code plugins a risk to local ecosystem?
r/LocalLLaMA
anyone else spending more time managing ai markdown files than actually coding?
r/LocalLLaMA
Carbon: Decoding the Language of Life
r/LocalLLaMA
Llama-server and MTP
r/LocalLLaMA
Qwen is cooking hard
r/LocalLLaMA
We have sub-agents at home
r/LocalLLaMA
Why might MTP be net negative for tool heavy agentic flows?
r/LocalLLaMA
Is there any <3B model with usable 200k+ context window?
r/LocalLLaMA
How many GPUs do you have on your local system/server/AI PC?
r/LocalLLaMA
favorite Agentic Coding Harness
r/LocalLLaMA
Still happy for yall
r/LocalLLaMA
Is the llama.cpp nixos flake just broken?
r/LocalLLaMA
MTP (Multi-Token Prediction): 2x Faster Token Generation on AMD Strix Halo & Radeon 9700 AI Pro
r/LocalLLaMA
Qwen cant wait to release 3.7 models
r/LocalLLaMA
Qwen 35b a3b surprises me
r/LocalLLaMA
Hopes and dreams for Google IO tomorrow? 👀
Reddit
What happens to local LLM if/when LLMs are no longer released for free?
r/LocalLLaMA
I tested 42 LLMs on their willingness to build the apocalypse. The "safest" closed-source models are lying to you.
r/LocalLLaMA
Quantizing MTP KV Cache = free lunch?
r/LocalLLaMA
GGUF with MTP vs MLX without. Is mlx still the way to go for mac users?
r/LocalLLaMA
New models when? Forecasting release date.
r/LocalLLaMA
The Lurk Report - The last 30 days of r/LocalLLaMA
r/LocalLLaMA
Is anyone prioritizing code quality checks via a small local model?
r/LocalLLaMA
Big new memory tool with local benchmarks
r/LocalLLaMA
I built a coding agent that gets 87% on benchmarks with a 4B parameter model, here's how
r/LocalLLaMA
May 2026 updated chart of strix halo mini pc size chart
How WeSearch handles this source
WeSearch's declared handling of r/LocalLLaMA's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.
Indexing: Allowed
Snippet: Allowed
AI summary: Limited
Retrieval / RAG: Not asserted
Model training: Not asserted
Commercial reuse: Not permitted