r/LocalLLaMA
social · source
r/LocalLLaMA on WeSearch
Recent social headlines from r/LocalLLaMA.
r/LocalLLaMA
PSA
r/LocalLLaMA
Step 3.7 Flash passes the car wash test
r/LocalLLaMA
Llama.cpp B9406 MTP mmproj fix
r/LocalLLaMA
FP16 on Qwen 3.6 27B
Reddit
Comparing Vector search libraries
r/LocalLLaMA
OAM waterblocks
r/LocalLLaMA
A moment of thanks for DeepSeek
r/LocalLLaMA
How do I make MTP work in llama-server?
r/LocalLLaMA
StepFun 3.7 Flash - Speed Benchmark in M5 Max
r/LocalLLaMA
Step 3.7 Flash Config + Early Data on 2x RTX 6000's
r/LocalLLaMA
Liquid AI releases LFM2.5-8B-A1B
r/LocalLLaMA
Which Coding Agent Features Are Useful For Local LLMs
r/LocalLLaMA
Beware!! Users trying to fork and steal your projects
r/LocalLLaMA
StepFun 3.7 Flash
r/LocalLLaMA
UPDATE: "Gentle Coding" is mathematically proven. 1,500+ test runs show major gain for Kimi K2.6 and even more for GLM-5.1! GPT 5.4/5.5 and Claude Sonnet 3.5/Opus 4.6 also better, with ZERO REGRESSION ACROSS THE BOARD.
r/LocalLLaMA
Ubuntu 26.04 on DGX Spark
r/LocalLLaMA
Upgrade path from 4x 3090s
r/LocalLLaMA
Mimo 2.5 Pro - 40t/s on 8x Nvidia Spark/GB10 cluster
r/LocalLLaMA
CrankGPT by Squeez Labs - hand-cranked edge AI - talk about local AI!!!
Reddit
I built a 103B-token Usenet corpus (1980–2013) — pre-web, human-only, zero AI contamination. Got strong traction on r/ML, thought this community would find it useful.
r/LocalLLaMA
Inferencing at 10.33 t/s on Qwen 3.5 35B on a $300 laptop
r/LocalLLaMA
Qwen3.6 huge quality gain from Q4 to Q6 for coding agent
r/LocalLLaMA
Looking for a working Deepseek-v4-Flash quant
r/LocalLLaMA
Why are the AI Companies spreading F.U.D. about AI?
r/LocalLLaMA
Is a 128 GB MacBook Pro M5 Max actually too slow for large-context local LLM coding workflows?
r/LocalLLaMA
Hugging Face Dataset Lineage Explorer
r/LocalLLaMA
Finally pioneering beyond the local 256k context window frontier!
r/LocalLLaMA
Found a Rust TUI coding agent that aggressively trims context with AST-level chunking. Cut my token bleed sharply with DeepSeek V4 Flash.
r/LocalLLaMA
Hyvemind OSS - Looking for some testers
r/LocalLLaMA
I made a small tool to inspect retrieval results before feeding them into RAG
r/LocalLLaMA
New DeepSWE benchmark finds Claude Opus cheats
r/LocalLLaMA
Turning every "no thats not what i meant" in chat into actual LoRA training data
r/LocalLLaMA
Does Engram Do Memory Retrieval in Autoregressive Image Generation?
r/LocalLLaMA
Stop traumatizing AI into loops and turn hallucinations into an honest "I don't know!" by being NICE to them (Proof of Concept, Research, I don't want to sell anything)
r/LocalLLaMA
How Qwen3.6-35B-A3B fails differently as a sub agent compared to solo
r/LocalLLaMA
I made a Windows app for managing llama.cpp in WSL/Ubuntu
r/LocalLLaMA
Long-context performance at lower quants
r/LocalLLaMA
OpenMOSS-Team/MOSS-TTS-v1.5 · Hugging Face
r/LocalLLaMA
Built a local-first AI memory system that indexes screen activity, meetings, and voice notes ( MCP + automations)
r/LocalLLaMA
SkillOpt treats markdown skill files as trainable parameters with proper optimization machinery
r/LocalLLaMA
Strix Halo users, a rejected PR can give you up to 30% faster PP for MOEs.
r/LocalLLaMA
Stop pretending self-hosting is cheaper. It's not. We do it for different reasons and we should say so.
r/LocalLLaMA
Qwen3.5 35B A3B uncensored heretic Native MTP Preserved is Out Now With the Full 785 MTPs Preserved and Retained, Available in Safetensors, GGUFs. NVFP4, NVFP4 GGUFs and GPTQ-Int4 Formats
r/LocalLLaMA
Running on a macbook, and having issues with crashing? Maybe this will help...
r/LocalLLaMA
I finally put my NPU (Intel Arrow Lake) to use doing ASR for my smart home
r/LocalLLaMA
CXMT started selling ram to corsair
r/LocalLLaMA
Is something went wrong with those online free model, why I feel they worse than Gemma 4 26B A4B Q4_KM ??
r/LocalLLaMA
One letter to appease them all
r/LocalLLaMA
Shard - getting to 10× KV cache compression
How WeSearch handles this source
WeSearch's declared handling of r/LocalLLaMA's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.
Indexing: Allowed
Snippet: Allowed
AI summary: Limited
Retrieval / RAG: Not asserted
Model training: Not asserted
Commercial reuse: Not permitted