r/LocalLLaMA
social · source
r/LocalLLaMA on WeSearch
Recent social headlines from r/LocalLLaMA.
r/LocalLLaMA
MTP experiences on 7900xtx?
r/LocalLLaMA
Grafting vision onto text models for fun and profit.
r/LocalLLaMA
Are local models good enough yet for AI meeting memory?
r/LocalLLaMA
llama: avoid copying logits during prompt decode in MTP by am17an · Pull Request #23198 · ggml-org/llama.cpp
r/LocalLLaMA
The power of structured workflows and small local models
r/LocalLLaMA
Developers who use local AI - Q4_0 vs Q8_0 KV quant?
r/LocalLLaMA
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
r/LocalLLaMA
How do I get the superfast DFlash / MTP tokens per second that I'm seeing on here? Dual 3090s
r/LocalLLaMA
Dual GPU llama.cpp speedup
r/LocalLLaMA
Convert With MPT Support?
r/LocalLLaMA
Good candidate model to act as a PA
r/LocalLLaMA
Is that was a right purchase for Qwen3.6 27/35
r/LocalLLaMA
Llama.cpp MTP with Qwen3.6 27B on Headless RTX 3090
r/LocalLLaMA
Jackrong/Qwopus3.5-9B-Coder-GGUF · Hugging Face
r/LocalLLaMA
Very happy with Qwen 3.5 122B output. But is slowness expected?
r/LocalLLaMA
LeanLoop, the Tool Claude Leans on
r/LocalLLaMA
"Elias Thorne" is what eight different LLMs name a lighthouse keeper. He's also selling cancer treatment advice on Amazon
r/LocalLLaMA
Looking to migrate off of Ollama and LMStudio
Reddit
Hardware Recommendations for realtime voice and a simple personal assistant/organisation agent.
r/LocalLLaMA
Meet Ronald
r/LocalLLaMA
webui: support video files as input by foldl · Pull Request #22830 · ggml-org/llama.cpp
r/LocalLLaMA
How do I correct a memory that was retrieved without asking for any help from the backend team? (personal experience)
r/LocalLLaMA
G4-Meromero-31B-Uncensored-Heretic Is Out Now, a Finetune of Gemma 4 31B It Designed for Creative Tasks, With Kld of 0.0100 and 15/100 Refusals!
r/LocalLLaMA
Ran the same models across Strix Halo, RTX 3090, and RTX 5070 because I wanted my own numbers
r/LocalLLaMA
an alternative = similar experience to using windsurf but on local?
r/LocalLLaMA
Now that MTP is merged... What's the best outputs you're getting on Qwen 3.6 35B on 2x3090s?
r/LocalLLaMA
WSL can't reach Kobold.cpp running on Windows, even though the API works fine in PowerShell, SillyTavern & a Kenshi SentientSands Mod. Does anyone know the solution?
r/LocalLLaMA
I fitted the new δ-mem research for apple silicon using mlx and openclaw integration! My findings
r/LocalLLaMA
Qwen3.5-122B-Q5-MTP - Qwen3.5-122B-Q6-MTP
r/LocalLLaMA
Best llama.cpp launch config for Qwen3.6 27B on RX 7800 XT (16 GB VRAM) for OpenClaw?
r/LocalLLaMA
gemma-4-Ortenzya-The-Creative-Wordsmith-31B-it-uncensored-heretic is Out Now, A Writing Finetune that Aims to Improve Gemma 4 31B it Writing Quality with More Natural English and Better Prose, Good for Creative Writings, Translations and RPs!
r/LocalLLaMA
Local Qwen 3.6 vs frontier models on a coding primitive: single-file HTML canvas driving animation - results and GIFs
r/LocalLLaMA
How I started programming differently over the last year. What about you?
r/LocalLLaMA
Corsair desktop PC with Ryzen 395 and 128GB of unified RAM, has anyone tested it for LLM? Seems "a good" price
r/LocalLLaMA
Qwen 27b MTP Config, Llama.cpp Single 3090
r/LocalLLaMA
Using Intel Arc Pro series, any thoughts ?
r/LocalLLaMA
b9180 llama.ccp MTP landed
r/LocalLLaMA
LLM Phone Home: Reliable Apps that can deliver inference from local backend
r/LocalLLaMA
Extension idea: llama-server with custom samplers
r/LocalLLaMA
Local speech to text for iOS using Apple Watch
r/LocalLLaMA
I've updated my glorified Llama fork (LLM Inference Server) for P40's to utilise MTP + TurboQuant + DFlash
r/LocalLLaMA
A very important milestone for me in the AI field.
r/LocalLLaMA
Built a 6x cheaper CodeRabbit alternative using open source models
r/LocalLLaMA
Reduce your GPU power limit
r/LocalLLaMA
When you run small LMM on RAM, dont use all Theards.
r/LocalLLaMA
What's in a GGUF, besides the weights - and what's still missing?
How WeSearch handles this source
WeSearch's declared handling of r/LocalLLaMA's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.
Indexing: Allowed
Snippet: Allowed
AI summary: Limited
Retrieval / RAG: Not asserted
Model training: Not asserted
Commercial reuse: Not permitted