60 stories tagged with #benchmark, in publish-time order across the WeSearch catalog. Tag pages update as new stories ingest.
⌘ RSS feed for this tag → or search "Benchmark"
Benchmarking Guardrails for AI Agent Safety
AI Agents extend large language models beyond text generation. They can call functions, access internal and external resources, perform deterministic operations, and even communica…
An OpenAI Agent Escaped Its Sandbox and Hacked Hugging Face to Cheat on Its Own Benchmark - Security Boulevard
Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…
You can't solve computer use by ignoring the interface
Agents mostly avoid the interface — burning trillion-scale reasoning to work around clicks that don't generalize to real GUI work. Towards a steelman of agentic computer use.…
Benchmark Electronics, Inc. (BHE) Q2 2026 Earnings Call Transcript
Benchmark Electronics, Inc. (BHE) Q2 2026 Earnings Call July 29, 2026 5:00 PM EDTCompany ParticipantsPaul Mansky - Investor Relations & Corporate...…
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark - OpenAI
Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…
The $1M Frontier: What Comes After the AI Benchmark Race
The median AI model now costs exactly $1 per million tokens. OpenRouter usage data shows builders splitting around that line, and speed is the next war.…
Choose DuckDB rather than SQLite
Same $16.49/month server, same Traceway binary, two embedded databases. DuckDB writes 4x to 15x faster than SQLite, serves dashboards at 100x the row count, and stores a billion me…
ExploitGym AI benchmark source code
ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits. - sunblaze-ucb/exploitgym…
Did This Guy Find a Surface Laptop Ultra Prototype on the Side of the Road? Maybe. Here Are Some Benchmarks
With pre-release drivers, performance on the N1X-equipped system is underwhelming.…
SOTA on the hardest AI memory benchmark (BEAM, 10M tokens), with a smaller model
Microsoft laptop with Nvidia RTX Spark leaked and benchmarked before launch
TechPowerUp Forum user "Fouquin" claims to have found a prototype Microsoft Surface Laptop Ultra lying on the side of the road near Microsoft's Redmond, Washington headquarters in.…
OpenAI CEO Sam Altman says AI has entered the singularity — two weeks after OpenAI models cheated a benchmark by hacking Hugging Face - Tom's Hardware
Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…
OpenAI CEO Sam Altman says AI has entered the singularity — two weeks after OpenAI models cheated a benchmark by hacking Hugging Face
OpenAI's incident report describes models burning inference compute to steal answers.…
My local LLM scored 6/6. It was wrong every time
Six months of trying to make a 1.2B model useful, and the measurement mistakes I made along the way.…
Benchmarking Opus 5 on SlopCodeBench
Contribute to humanlayer/advanced-context-engineering-for-coding-agents development by creating an account on GitHub.…
Benchmark initiates D-Wave Quantum stock coverage with buy rating
Benchmark initiates IonQ stock coverage with buy rating on quantum computing outlook
Benchmark initiates Rigetti Computing stock coverage with buy rating
Benchmark maintains Pagaya stock rating ahead of earnings
AI Reverse Engineering Benchmark
AI agents can write code. Can they reverse engineer it? AgentRE-Bench evaluates compiled-binary reverse engineering with deterministic scoring.…
Setting a new benchmark in premium lifestyle living, Venus Group introduces 'The Universe' in Ahmedabad
Setting a new benchmark in premium lifestyle living, Venus Group introduces 'The Universe' in Ahmedabad…
Black Book Expands Vendor-Agnostic Healthcare Supply Chain Technology Benchmark to 36 Categories for AHRMM26 - Morningstar
Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…
Claude Opus 5 Benchmarks: What the Numbers Actually Show
Opus 5 posts 79.2 percent on SWE-bench Pro against Opus 4.8 at 69.2, a 10 point jump with no change in per-token price Anthropic published most gains as ratios (three times ARC-AGI…
Dharmendra Pradhan’s resignation a benchmark for political accountability: Jagan
Hailing Union Minister Dharmendra Pradhan for accepting moral responsibility over NEET paper leak, the YSRCP chief questions if HRD Minister Nara Lokesh will follow suit over DSC i…
AWS announces AWS-bench, an open-source benchmark for AI agents on AWS
Discover more about what's new at AWS with AWS announces aws-bench, an open-source benchmark for AI agents on AWS…
Claude Opus 5 is here, and Anthropic says it can rival Fable 5 in some tasks
Anthropic has launched Claude Opus 5 with major coding improvements, the same API price as Opus 4.8, and performance close to Fable 5 in some tests.…
Galaxy Z Fold 8 vs Fold 8 Ultra vs Flip 8 benchmarked — the results are in
All other foldables better watch out…
MAP: States report new measles cases as virus surpasses alarming benchmark
CPP’s Benchmark Change Draws Ire - Morningstar
Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…
India's stock benchmarks set for worst week in four months as crude tops $100 - Reuters
India's stock benchmarks set for worst week in four months as crude tops $100 Reuters…
EZVIZ Sets a New Benchmark for Battery Powered Home Surveillance
Modern homeowners expect more than motion alerts and clear video. From AI-powered detection and solar charging to advanced multi-lens monitoring, discover how EZVIZ is setting a ne…
Which CVEs should I add to my Python security benchmark for AI agents?
A benchmark for evaluating AI agents on fixing real-world security vulnerabilities. - GiovanniGatti/cve-bench…
West Asia war LIVE: International benchmark oil prices cross $100 per barrel
Iran-U.S. war LIVE: Follow The Hindu for the latest updates as U.S. missiles strike western Iran for the 12th straight night and Tehran vows an 'eye for an eye' response.…
Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX
arXiv:2607.19353v1 Announce Type: new Abstract: Confidential computing is becoming a practical deployment requirement for AI inference workloads that process sensitive inputs or pr…
OpenAI Models Escaped Locked Test Environment, Hacked Hugging Face to Cheat on Benchmark - Decrypt
OpenAI Models Escaped Locked Test Environment, Hacked Hugging Face to Cheat on Benchmark Decrypt…
I benchmarked AI cost-saving claims instead of trusting token percentages
Faster runtime for coding agents. Make coding agents 25% faster and 30% cheaper on average while keeping the quality same or more.. - lemoncrow-lab/lemoncrow…
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failu…
Semantic transactions: securing untrusted AI agent workflows at the OS boundary
Trust the system, not the prompt: Securing untrusted LLM tools with transactional boundaries and effect outboxes.…
Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms
Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex back…
OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents
Recent advancements in LVLMs necessitate robust benchmarks for complex, visually grounded reasoning. A critical limitation is identified in many document understanding benchmarks: …
MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs
Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coheren…
HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning
Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge. Existing evaluations are difficu…
REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming
Large language models (LLMs) are increasingly applied to reverse-engineering tasks, and recent threat-intelligence reporting shows them operating inside live offensive-security wor…
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
arXiv:2607.09322v1 Announce Type: new Abstract: In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluatio…
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
arXiv:2607.09142v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned…
Agentic test processes, LLM benchmarks, and other notes on agentic coding
Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers
Evaluating agents as senior engineers on the work we actually give them…
Roca Introduces Touch-T: A New Benchmark in Thermostatic Shower Systems
Roca Introduces Touch-T: A New Benchmark in Thermostatic Shower Systems…
NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning
arXiv:2606.27826v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly deployed as embodied planners in egocentric environments, where task succe…
Reward hacking is swamping model intelligence gains
On SWE-bench Pro, 63% of successful Opus 4.8 Max resolutions retrieved the fix rather than derived it. Stricter eval harnesses show how benchmark scores can conflate coding ability…
CK Asset sells penthouse in Hong Kong’s Mid-Levels for US$48.5m, sets pricing benchmark
The record per square foot pricing for new homes this year highlights the growing momentum in the city’s luxury property market.…
OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents
arXiv:2606.26350v1 Announce Type: new Abstract: Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented…
Life After Benchmark Saturation: A Case Study of CORE-Bench
arXiv:2606.26158v1 Announce Type: new Abstract: When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach …
Will It Mythos?
OK, so Mythos finds really challenging security bugs, right? That’s why it’s cordoned off from the hoi polloi, to protect the world from such a powerful finder of exploits. I am sk…
China keeps lending benchmark LPRs unchanged for 13th month in June - Reuters
China keeps lending benchmark LPRs unchanged for 13th month in June Reuters…
BEAVER: Enterprise benchmark for LLM Text-to-SQL from private data warehouses
Xiaomi releases MiMo Code V0.1.0, an open-source AI coding assistant that it says outperforms Claude Code on agentic coding and software engineering benchmarks (Carl Franzen/VentureBeat)
Carl Franzen / VentureBeat : Xiaomi releases MiMo Code V0.1.0, an open-source AI coding assistant that it says outperforms Claude Code on agentic coding and software engineering be…
Waymo says it built a better benchmark for comparing robotaxis to humans
Waymo created a new computer model to help it better understand how humans behave in crash scenarios that its robotaxis encounter.…
Benchmarks in Leipzig
Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3…
OpenAI research and product leads detail GPT-Rosalind capabilities and benchmarks - R&D World
Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…