WeSearch
Hub / Tags / Language Models
TAG · #LANGUAGE-MODELS

Language Models coverage.

Every story in the WeSearch catalog tagged with #language-models, chronological, with view counts. Subscribe to the per-tag RSS feed to follow this topic in your reader of choice.

60 stories tagged with #language-models, in publish-time order across the WeSearch catalog. Tag pages update as new stories ingest.

⌘ RSS feed for this tag →   or   search "Language Models"

RELATED TAGS
#ai122#ml75#large-language-models8#technology5#reinforcement-learning4#research4#ai-research3#openai3#computation3#anthropic2#document-editing2#cybersecurity2
ARXIV.ORG

Rethinking Uncertainty Evaluation in Large Language Models

arXiv:2607.19367v1 Announce Type: new Abstract: Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimator…

21 views ·
#rethinking#uncertainty#evaluation
ARXIV.ORG

Logic-Guided Data Extraction with Answer Set Programming and Large Language Models

arXiv:2607.19365v1 Announce Type: new Abstract: When Large Language Models (LLMs) are used for semantic data extraction from unstructured text, producing candidate relational facts…

19 views ·
#logic-guided#data#extraction
ARXIV.ORG

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models

arXiv:2607.19364v1 Announce Type: new Abstract: Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based s…

22 views ·
#statistically#grounded#sparse-feature
ARXIV.ORG

Lifted Representation Hypothesis in Language Models

arXiv:2607.19360v1 Announce Type: new Abstract: Large language models (LLMs) often answer queries by mapping individual observations to more general rule-like structures. However, …

15 views ·
#lifted#representation#hypothesis
ARXIV.ORG

Information Discernment in Large Language Models

arXiv:2607.19355v1 Announce Type: new Abstract: LLMs are increasingly used with external knowledge sources like the internet. Do they weigh information appropriately -- updating mo…

18 views ·
#information#discernment#large
TOWARDS DATA SCIENCE

Loop Engineering with Adaptive Parsing in Action: Parsing Flat Tables with Azure and Figures with a Vision LLM

Enterprise Document Intelligence [Vol.1 #10B] - The LLM as last line of defence, then two real escalations walked end to end: a flat table to Azure, a figure to a vision model The …

28 views ·
#document intelligence#adaptive parsing#large language models
MIT TECHNOLOGY REVIEW

GPT-Red: an LLM super-hacker OpenAI built to make its models safer

Exclusive: The firm says it wants to future-proof its safety procedures and stay ahead of human attackers.…

77 views ·
#ai safety#red teaming#prompt injection
ARXIV CS.AI

Augmenting Fundamental Analysis with Large Language Models: A RAG-Based System for Generating Investor Briefs

In this study, we examine the opportunities brought by Large Language Models (LLMs) to various aspects of fundamental analysis of companies based on their reports as well as data a…

30 views ·
#augmenting#fundamental#analysis
ARXIV CS.AI

Integrating Large Language Models and Graph Convolutional Networks for Semi-Supervised Image Classification

While the growing availability of image data has driven significant advances, labeling datasets remains costly and time-consuming. Therefore, semi-supervised approaches such as Gra…

28 views ·
#integrating#large#language
ARXIV CS.AI

Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices

With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introduce sparsity into weight matrices can acce…

24 views ·
#accelerating#inference#large
ARXIV CS.AI

A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions

Despite the success of knowledge distillation (KD) in Large Language Models (LLMs), the underlying mechanism behind its efficacy remains unclear. In this paper, we propose a unifie…

23 views ·
#unified#approach#interpreting
ARXIV.ORG

Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models

arXiv:2606.26366v1 Announce Type: new Abstract: Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with…

33 views ·
#narration-of-thought#inference-time#scaffolding
SIMON WILLISON'S WEBLOG

Prompt Injection as Role Confusion

Prompt Injection as Role Confusion First, I absolutely love this: This is a blog-style writeup of the paper. I wish every paper would come with one of these. Academic writing is pr…

43 views ·
#ai#security
ARXIV.ORG

GPU Forecasters: Language Models as Selective Surrogates for Kernel Optimization

GPU kernels are the workhorse of modern deep learning, and optimizing them (via evolutionary search or coding agents) usually requires repeated measurement on target hardware. Whil…

37 views ·
#machine learning#artificial intelligence#gpu optimization
ARXIV.ORG

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

As large language models (LLMs) are increasingly used for long-form generation, reliably evaluating long-form outputs has become a critical challenge. LLM-as-a-judge offers a scala…

46 views ·
#machine learning#evaluation
ARXIV CS.AI

Code-on-Graph: Iterative Programmatic Reasoning via Large Language Models on Knowledge Graphs

Knowledge Graphs (KGs) are widely used to mitigate the limitations of Large Language Models (LLMs), such as outdated knowledge and hallucinations. Existing LLM-KG integration frame…

52 views ·
#artificial intelligence#knowledge graphs
ARXIV CS.AI

From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models

Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical failure mode: a model may o…

40 views ·
#artificial intelligence#chemistry#machine learning
ARXIV CS.AI

ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models

Large language models (LLMs) have been widely adopted in healthcare, yet they still encounter significant challenges in complex clinical decision-making scenarios. Existing benchma…

41 views ·
#healthcare#artificial intelligence#clinical decision-making
ARXIV CS.AI

Uncertainty-Aware Clarification in LLM Agents with Information Gain

Large Language Model (LLM) agents often operate under underspecified user instructions, where latent uncertainty over user intent leads to erroneous tool actions. To address this c…

43 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

Decomposing how prompting steers behavior

Prompting steers large language models (LLMs) and vision-language models (VLMs) without weight updates, but it remains unclear how instruction changes reshape internal representati…

49 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs

Inference-time scaling has emerged as a critical avenue for enhancing Large Language Models' performance, yet real-world deployment is constrained by strict computational budgets. …

42 views ·
#artificial intelligence#budget allocation#large language models
ARXIV CS.AI

SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale

As LLM agents adopt large skill libraries, selecting the right subset becomes a structural problem rather than a similarity-matching one: skills depend on, conflict with, specializ…

40 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning

Large language models (LLMs) exhibit strong natural-language reasoning abilities for clinical decision support, but struggle to effectively model structured longitudinal electronic…

39 views ·
#artificial intelligence#healthcare#machine learning
ARXIV CS.AI

Visual Graph Scaffolds for Structural Reasoning in Large Language Models

Graphs have been used to enhance large language models (LLMs) for structured reasoning, mostly as external knowledge sources are provided to models at test time. In this paper, we …

44 views ·
#artificial intelligence#machine learning
IEEE SPECTRUM

Why Are Large Language Models So Terrible at Video Games?

LLMs can code your retro shooter but still fail at playing Halo; see what this gap reveals about AI’s real limits in 2026…

47 views ·
#llms#artificial-intelligence#video-games
ARXIV.ORG

Scaling Laws for Agent Harnesses via Effective Feedback Compute

Agent harnesses increasingly determine the performance of language-model systems by deciding how models call tools, receive feedback, verify intermediate states, store memory, and …

35 views ·
#computer science#feedback
R/PROMPTENGINEERING

Heuristic Parasites: A Behavioral Taxonomy of Recurrent Distortion Patterns in Large Language Models (Full System) V2

38 views ·
ARXIV.ORG

AI Propaganda factories with language models

AI-powered influence operations can now be executed end-to-end on commodity hardware. We show that small language models produce coherent, persona-driven political messaging and ca…

42 views ·
#artificial intelligence#cryptography#security
DEV.TO (TOP)

✨📊 🧠 The Ultimate Visual Guide to Large Language Models (LLMs)

Generative AI is a type of artificial intelligence that can produce new content including text,...…

35 views ·
#artificial intelligence#machine learning
DEV.TO (TOP)

📄Paper: RORA-VLM: Robust Retrieval Augmentation for Vision Language Models

Public At International Conference on Learning Representations (ICLR) 2025 💡 Why I read...…

36 views ·
#ai#vlm#research
ARS TECHNICA - ALL CONTENT

LLMs believe false statements even after explicit warnings that they're false

Fine-tuning tests show "bias ... toward confidently representing the claims as true."…

45 views ·
#artificial intelligence#research
UNITE.AI

Why does AI love writing about lighthouse keepers?

Asked to 'write a story', ChatGPT and other leading language models appear to be avoiding copyright infringement by obsessive recourse to the same small and strange cast of lightho…

28 views ·
#artificial intelligence#storytelling
ARXIV.ORG

How sure is the activation oracle?

Activation oracles aim to make the activations of other models legible to humans and yield promising results compared to white-box interpretability techniques. However, uncertainty…

26 views ·
#artificial intelligence#interpretability
ARXIV CS.AI

PitchBench: Measuring Pitch Hearing in Audio-Language Models

Audio-language models (ALMs) are increasingly used in real-world applications that require understanding music, from music tutoring and transcription to captioning, recommendation …

35 views ·
#audio#artificial intelligence#music
ARXIV CS.AI

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications

Large Language Models (LLMs) have become the predominant paradigm in NLP, advancing both research and industry. As model sizes and pretraining data grow, concerns about Pretraining…

33 views ·
#artificial intelligence#machine learning#data privacy
ARXIV CS.AI

Gumbel Machine: Counterfactual Student Writing Generation via Gumbel Noise Steering

An effective method of teaching across disciplines is to provide examples of high-quality work. However, an example may be significantly different from a student's current work, ma…

35 views ·
#artificial intelligence#education#writing
ARXIV CS.AI

Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs

Retrieval-augmented LLMs are deployed for tasks where evidence quality determines action safety, yet evaluation protocols assume that single-turn robustness predicts robustness whe…

32 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation

Domain specialization can improve LLM behavior in vertical domains, but often weakens the general capabilities inherited from the original model. Recent Multi-Teacher On-Policy Dis…

45 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

Generating Robust Portfolios of Optimization Models using Large Language Models

Mathematical optimization is a powerful tool for structured decision-making across domains such as resource allocation and planning. Formulating optimization models faithful to rea…

33 views ·
#artificial intelligence#optimization
ARXIV CS.AI

Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation

Multi-stakeholder tasks require one output to satisfy users with conflicting preferences. Holistic LLM judges conflate utility estimation and utility aggregation, yielding unstable…

38 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global Derivation

Chain-of-thought (CoT) prompting reliably improves language-model accuracy, but which properties of a rationale text drive the improvement is poorly understood. Prior work has larg…

32 views ·
#artificial intelligence#research
ARXIV CS.AI

The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context

Retrieval-augmented generation promises to ground language model outputs in external evidence, yet the field has no reliable way to verify whether retrieved context actually govern…

32 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation

Large language models (LLMs) have shown strong empirical gains as self-evolving agents for CUDA kernel generation, driven by feedback-conditioned planning across generations. Howev…

46 views ·
#artificial intelligence#machine learning#cuda
ARXIV CS.AI

AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents

The token-level extractive compressors widely used for general LM context are structurally inappropriate for LLM agents: across 17 (env, backbone, method) cells spanning two indepe…

38 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning

Clinical practice guidelines (CPGs) encode evidence-based decision logic that clinicians apply by evaluating patient variables, conditional criteria, and recommendation rules. Howe…

35 views ·
#artificial intelligence#healthcare#clinical reasoning
ARXIV CS.AI

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…

36 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions

Large Language Models (LLMs) achieve impressive accuracy on mathematical reasoning benchmarks, yet their performance drops when problems are modified with simple changes like diffe…

37 views ·
#artificial intelligence#machine learning#mathematics
ARXIV CS.AI

OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling

Theory of Mind (ToM), the ability to infer others' knowledge, intentions, and emotions, is commonly evaluated in large language models (LLMs) using end-point question answering, wh…

40 views ·
#artificial intelligence#theory of mind
ARXIV CS.AI

Can LLMs Introspect? A Reality Check

Can large language models detect and report their own internal states? A number of studies have argued that the answer to this question is yes. We argue, based on lessons from huma…

42 views ·
#artificial intelligence#metacognition
MICROSOFT RESEARCH

Microsoft Research: LLMs Corrupt your files during delegated work

Large Language Models (LLMs) are poised to disrupt knowledge work, with the emergence of delegated work as a new interaction paradigm (e.g., vibe coding). Delegation requires trust…

32 views ·
#artificial intelligence#document editing
LET'S DATA SCIENCE

Sparse Autoencoders Reveal Cortical Brain-LLM Semantic Mapping

A preprint submitted to arXiv (arXiv:2605.23035) by Dongxin Guo and colleagues presents a mechanistic interpretability approach connecting large language model representations to h…

43 views ·
#neuroscience#machine learning
ARXIV.ORG

Prompt Politeness Affects LLM Accuracy

The wording of natural language prompts has been shown to influence the performance of large language models (LLMs), yet the role of politeness and tone remains underexplored. In t…

34 views ·
#artificial intelligence#research
SMOLA

You don't need all the LLM benchmarks

32 views ·
#machine learning#benchmarks
ARXIV CS.AI

Credit Assignment with Resets in Language Model Reasoning

Contemporary reinforcement learning with verifiable reward methods post-train language models on multi-step reasoning by assigning a single outcome reward uniformly across all toke…

25 views ·
#artificial intelligence#reinforcement learning
ARXIV CS.AI

Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models

Large language models often generate confident but incorrect answers rather than abstaining when uncertain. This problem is particularly acute for small language models (SLMs), whe…

26 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

DarkForest: Less Talk, Higher Accuracy for Multi-Agent LLMs

Multi-agent LLM systems improve reasoning by combining outputs from multiple agents, but interaction-heavy methods can introduce error propagation and high communication overhead. …

23 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

Representation Without Control: Testing the Realization Effect in Language Models

Large language models are increasingly used as behavioral simulators, but it remains unclear when their outputs reflect human-like cognitive mechanisms rather than prompt-sensitive…

32 views ·
#artificial intelligence#behavioral economics
ARXIV CS.AI

Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling

Test-time scaling improves language model reasoning by spending additional compute to explore multiple solution trajectories. The key challenge is to maximize accuracy while minimi…

33 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction

Reliably knowing when a language model is correct is almost as important as being correct. We introduce prover-verifier deliberation (PVD), an inference-time protocol grounded in i…

32 views ·
#artificial intelligence#machine learning
ARXIV CS.AI

Privacy-Preserving Local Language Models for Longitudinal Data Retrieval in Chronic Dermatologic Disease: Implementation in Pemphigus Patients

Chronic dermatologic diseases such as pemphigus require long-term follow-up, generating extensive longitudinal clinical documentation that is difficult to review comprehensively du…

34 views ·
#artificial intelligence#healthcare#dermatology