47 stories tagged with #distill, in publish-time order across the WeSearch catalog. Tag pages update as new stories ingest.
⌘ RSS feed for this tag → or search "Distill"
Show HN: Distill and serve small models with frontier quality for half the cost
Hi HN, we built world-model-optimizer, an open-source tool for continually improve models specialized to agents. Today we are launching `wmo serve`, a tool to route repetitive task…
From Silicon Valley to DC, the tech world is suddenly obsessed with one concept in AI: Distillation
Distillation has long been a topic for AI wonks, but it's become a hot-button issue of late as techies and lawmakers debate how it should be regulated.…
Ryan Reynolds-backed distillery closes as share of Americans drinking alcohol hits new low
Ryan Reynolds’ Portland gin distillery – once advertised as “Disneyland for adults” – just poured its last drink. It comes as booze brands struggle to adapt with barely half of Ame…
Trump under pressure over China’s AI surge, amid ‘distillation’ accusations
China has been accused of "distilling" US AI models to produce cheaper versions. Read more at straitstimes.com. Read more at straitstimes.com.…
OpenAI's Brockman says distillation is a technical problem - Yahoo
Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…
Old Forester® Marks a New Chapter for its Main Street Distillery with newest 117 Series expression: Triple Char - Morningstar
Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…
Updating IP Regulations for AI Distillation
A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions
Despite the success of knowledge distillation (KD) in Large Language Models (LLMs), the underlying mechanism behind its efficacy remains unclear. In this paper, we propose a unifie…
Open weight AI models are facing an existential policy test in the US, with Anthropic leading a campaign against Chinese models over distillation concerns (Nathan Lambert/Interconnects AI)
Nathan Lambert / Interconnects AI : Open weight AI models are facing an existential policy test in the US, with Anthropic leading a campaign against Chinese models over distillatio…
ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents
arXiv:2606.27814v1 Announce Type: new Abstract: Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. …
Anthropic Alleges Largest-Ever Claude Distillation Attack by Alibaba
SITUATION DETECTED: Anthropic has disclosed to the U.S. Government that Alibaba executed the largest known distillation attack on Claude to date, generating 28.8 million exchanges …
Anthropic Accuses Alibaba of Largest AI Distillation Attack: 28.8M Fraudulent
Anthropic sends letter to U.S. Senate accusing Alibaba of conducting the largest AI distillation attack using 25,000 fraudulent accounts and 28.8M exchanges with Claude AI. Stock d…
Distilling Stale Gasoline to Make it Usable Again
The propensity of gasoline to ‘go stale’ through the process of oxidation is the reason why gasoline that has been stored for a long period of time is considered to be unusable, as…
Show HN: The Logos Machine – AI that distill knowledge into weights during Sleep
DocSend helps you communicate more effectively by telling you what happens to content after you send them and letting you keep control in real time.…
Show HN: Terraform RAG - index modules, distill conventions, compose via MCP
Distilling Answer-Set Programming Rules from LLMs for Neurosymbolic Visual Question Answering
Visual Question Answering (VQA) is the task of answering questions about images, requiring the integration of multimodal input and reasoning. Modular approaches that incorporate lo…
Skill Distillation
How a personal AI agent built on markdown skills lets a frontier model teach smaller, local models to do real work, without retraining.…
8-step FLUX.2-dev DMD2 distillation
Whisky Weekend: Maker’s Mark Turns Wheat Whisky Into Something Worth Thinking About
Author’s note: Whiskey Wednesday is usually a midweek reward, but every now and then, the calendar hands you an excuse to bend the rules. When the opportunity for a new experience …
How Model Distillation Actually Works (and What the 'China Distilled Our Model' Headlines Really Mean)
A practical, no-hype explainer of knowledge distillation in LLMs — the actual mechanics, why distilling from a closed API is different, and what the OpenAI/Anthropic vs DeepSeek al…
Sequoia's 'This is AGI' talk, distilled — what it means if you build on the models
Sequoia's AI Ascent 2026 keynote ("This is AGI") is worth 32 minutes of your time. I distilled it...…
High Liquor Taxes and a Home Distillation Ban Guarantee a Thriving Booze Black Market
Between a home distillation ban and high liquor taxes, government officials have created the perfect conditions for a black market in distilled spirits.…
When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability
Guided Soft Actor-Critic (GSAC) distills knowledge from a privileged full-state teacher to a partial-observation student for autonomous driving, but uses a fixed distillation coeff…
StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning
Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hinges on a few local decisions…
Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation
Domain specialization can improve LLM behavior in vertical domains, but often weakens the general capabilities inherited from the original model. Recent Multi-Teacher On-Policy Dis…
Maryland governor signs bill making cocktails to-go permanent in Baltimore County
Maryland Gov. Wes Moore has signed legislation permanently restoring cocktails to-go in Baltimore County, reviving a pandemic-era convenience that lapsed nearly three years ago.…
PANDO: Efficient Multimodal AI Agents via Online Skill Distillation
Recent advances in multimodal web agents often rely on increased inference-time computation, including rollout search, verifier passes, offline skill discovery, and specialist mode…
Distilling Game Code World Model Generation into Lightweight Large Language Models
Large Language Models (LLMs) have shown great ability in generating executable code from natural language, opening the possibility of automatically constructing environments for AI…
Ask HN: Local model experiences with 'high-reasoning distill' finetunes
Is this real enough for that baseball trend going on =P - WAN2GP LTX2.3 distilled 1.1
96% Correct Next Token Prediction, with No DNN, No Training, Autodistilled Model
Over the last 12 months, I’ve built a model to predict the next token and to suggest synonyms or related queries to a user prompt, with 100% correct predictions on the training set…
EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation
On-Policy Distillation (OPD) has gained wide attraction as an LLM post-training paradigm due to its effectiveness in improving capabilities without introducing model distribution d…
Journalists Distill News on Ebola, Licensing Midwives, and California’s Budget
KFF Health News journalists made the rounds on national or local media recently to discuss topical stories. Here’s a collection of their appearances.…
PACD-Net: Pseudo-Augmented Contrastive Distillation for Glycemic Control Estimation from SMBG
Effective diabetes management requires continuous monitoring of glycemic levels. Clinically, glycemic control is assessed using metrics such as Time in Range (TIR), Time Below Rang…
AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals
Self-distillation enables language models to learn on-policy from their own trajectories by using the same model as both student and teacher, with the teacher being conditioned on …
Consistently Informative Soft-Label Temperature for Knowledge Distillation
Knowledge distillation (KD) transfers knowledge from a high-capacity teacher to a compact student by matching their predictive distributions, with temperature scaling serving as a …
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
Contextual Integrity (CI) defines privacy not merely as keeping information hidden, but as governing information flows according to the norms of a given context. As large language …
From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation
Self-attention serves as the core foundation of large-scale transformer pretraining, but its quadratic token interaction cost makes inference expensive. Replacing attention with si…
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
Reinforcement learning can train LLM agents from sparse task rewards, but long-horizon credit assignment remains challenging: a single success-or-failure signal must be distributed…
AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment
The alignment of Large Language Models (LLMs) for complex reasoning heavily relies on Reinforcement Learning with Verifiable Rewards (RLVR). However, standard algorithms like GRPO …
SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented Reasoning
Search-augmented reasoning agents interleave internal reasoning with calls to an external retriever, and their performance relies on the quality of each issued query. However, unde…
Trying to distill the soon-to-be-sunset Imagen 4 to a LoRA for Illustrious 2.0 but the result is a bit wonky, would appreciate some pointers
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
Block attention, which processes the input as separate blocks that cannot attend to one another, offers significant potential to improve KV cache reuse in long-context scenarios su…
DeltaPrompts: Escaping the Zero-Delta Trap in Multimodal Distillation
Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process are typically chosen via simple heuristics …
Gold medal-winning vodka distillery files for Chapter 11 bankruptcy
Self-Distillation Enables Continual Learning [PDF]
Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models. While on-pol…
I built an open-source tool to distill books into knowledge graphs
I have a bad habit: I buy books faster than I read them. Not because I'm lazy — I start most of...…