WeSearch
Hub / Tags / Benchmark
TAG · #BENCHMARK

Benchmark coverage.

Every story in the WeSearch catalog tagged with #benchmark, chronological, with view counts. Subscribe to the per-tag RSS feed to follow this topic in your reader of choice.

60 stories tagged with #benchmark, in publish-time order across the WeSearch catalog. Tag pages update as new stories ingest.

⌘ RSS feed for this tag →   or   search "Benchmark"

RELATED TAGS
#benchmarking69#ai61#technology15#ml11#benchmarks10#security8#performance5#programming5#api5#hardware4#evaluation4#intel3
MOZILLA.AI

Benchmarking Guardrails for AI Agent Safety

AI Agents extend large language models beyond text generation. They can call functions, access internal and external resources, perform deterministic operations, and even communica…

5 views ·
#ai safety#machine learning#security
GOOGLE NEWS

An OpenAI Agent Escaped Its Sandbox and Hacked Hugging Face to Cheat on Its Own Benchmark - Security Boulevard

Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…

6 views ·
STEELMAN LABS

You can't solve computer use by ignoring the interface

Agents mostly avoid the interface — burning trillion-scale reasoning to work around clicks that don't generalize to real GUI work. Towards a steelman of agentic computer use.…

5 views ·
#ai#computer-use#interfaces
SEEKING ALPHA

Benchmark Electronics, Inc. (BHE) Q2 2026 Earnings Call Transcript

Benchmark Electronics, Inc. (BHE) Q2 2026 Earnings Call July 29, 2026 5:00 PM EDTCompany ParticipantsPaul Mansky - Investor Relations & Corporate...…

9 views ·
#electronics#earnings
GOOGLE NEWS

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark - OpenAI

Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…

10 views ·
VERTICAL

The $1M Frontier: What Comes After the AI Benchmark Race

The median AI model now costs exactly $1 per million tokens. OpenRouter usage data shows builders splitting around that line, and speed is the next war.…

5 views ·
#frontier#what#comes
TRACEWAYAPP

Choose DuckDB rather than SQLite

Same $16.49/month server, same Traceway binary, two embedded databases. DuckDB writes 4x to 15x faster than SQLite, serves dashboards at 100x the row count, and stores a billion me…

5 views ·
#databases#performance
GITHUB

ExploitGym AI benchmark source code

ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits. - sunblaze-ucb/exploitgym…

7 views ·
#exploitgym#source
PCMAG

Did This Guy Find a Surface Laptop Ultra Prototype on the Side of the Road? Maybe. Here Are Some Benchmarks

With pre-release drivers, performance on the N1X-equipped system is underwhelming.…

11 views ·
HACKER NEWS (AI / LLM)

SOTA on the hardest AI memory benchmark (BEAM, 10M tokens), with a smaller model

7 views ·
#sota#hardest#memory
TECHSPOT

Microsoft laptop with Nvidia RTX Spark leaked and benchmarked before launch

TechPowerUp Forum user "Fouquin" claims to have found a prototype Microsoft Surface Laptop Ultra lying on the side of the road near Microsoft's Redmond, Washington headquarters in.…

7 views ·
#microsoft#laptop#nvidia
GOOGLE NEWS

OpenAI CEO Sam Altman says AI has entered the singularity — two weeks after OpenAI models cheated a benchmark by hacking Hugging Face - Tom's Hardware

Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…

11 views ·
TOM'S HARDWARE

OpenAI CEO Sam Altman says AI has entered the singularity — two weeks after OpenAI models cheated a benchmark by hacking Hugging Face

OpenAI's incident report describes models burning inference compute to steal answers.…

12 views ·
#openai#altman#says
MARKBHALL

My local LLM scored 6/6. It was wrong every time

Six months of trying to make a 1.2B model useful, and the measurement mistakes I made along the way.…

7 views ·
#ai#evaluation
GITHUB

Benchmarking Opus 5 on SlopCodeBench

Contribute to humanlayer/advanced-context-engineering-for-coding-agents development by creating an account on GitHub.…

11 views ·
#benchmarking#opus#slopcodebench
INVESTING.COM — NEWS

Benchmark initiates D-Wave Quantum stock coverage with buy rating

4 views ·
INVESTING.COM — NEWS

Benchmark initiates IonQ stock coverage with buy rating on quantum computing outlook

8 views ·
INVESTING.COM — NEWS

Benchmark initiates Rigetti Computing stock coverage with buy rating

11 views ·
INVESTING.COM — NEWS

Benchmark maintains Pagaya stock rating ahead of earnings

6 views ·
AGENTRE-BENCH

AI Reverse Engineering Benchmark

AI agents can write code. Can they reverse engineer it? AgentRE-Bench evaluates compiled-binary reverse engineering with deterministic scoring.…

13 views ·
THE HINDU

Setting a new benchmark in premium lifestyle living, Venus Group introduces 'The Universe' in Ahmedabad

Setting a new benchmark in premium lifestyle living, Venus Group introduces 'The Universe' in Ahmedabad…

16 views ·
#setting#premium
GOOGLE NEWS

Black Book Expands Vendor-Agnostic Healthcare Supply Chain Technology Benchmark to 36 Categories for AHRMM26 - Morningstar

Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…

19 views ·
DEV COMMUNITY

Claude Opus 5 Benchmarks: What the Numbers Actually Show

Opus 5 posts 79.2 percent on SWE-bench Pro against Opus 4.8 at 69.2, a 10 point jump with no change in per-token price Anthropic published most gains as ratios (three times ARC-AGI…

14 views ·
#claude#opus#benchmarks
THE HINDU — TOP

Dharmendra Pradhan’s resignation a benchmark for political accountability: Jagan

Hailing Union Minister Dharmendra Pradhan for accepting moral responsibility over NEET paper leak, the YSRCP chief questions if HRD Minister Nara Lokesh will follow suit over DSC i…

11 views ·
#dharmendra#pradhan#resignation
AMAZON WEB SERVICES, INC.

AWS announces AWS-bench, an open-source benchmark for AI agents on AWS

Discover more about what's new at AWS with AWS announces aws-bench, an open-source benchmark for AI agents on AWS…

11 views ·
#announces#aws-bench#open-source
DIGITAL TRENDS

Claude Opus 5 is here, and Anthropic says it can rival Fable 5 in some tasks

Anthropic has launched Claude Opus 5 with major coding improvements, the same API price as Opus 4.8, and performance close to Fable 5 in some tests.…

15 views ·
#ai#machine-learning#benchmarks
TOM'S GUIDE

Galaxy Z Fold 8 vs Fold 8 Ultra vs Flip 8 benchmarked — the results are in

All other foldables better watch out…

9 views ·
#galaxy#fold#ultra
THE HILL

MAP: States report new measles cases as virus surpasses alarming benchmark

13 views ·
GOOGLE NEWS

CPP’s Benchmark Change Draws Ire - Morningstar

Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…

9 views ·
GOOGLE NEWS

India's stock benchmarks set for worst week in four months as crude tops $100 - Reuters

India's stock benchmarks set for worst week in four months as crude tops $100 Reuters…

8 views ·
DIGITAL TRENDS

EZVIZ Sets a New Benchmark for Battery Powered Home Surveillance

Modern homeowners expect more than motion alerts and clear video. From AI-powered detection and solar charging to advanced multi-lens monitoring, discover how EZVIZ is setting a ne…

14 views ·
#ezviz#sets
GITHUB

Which CVEs should I add to my Python security benchmark for AI agents?

A benchmark for evaluating AI agents on fixing real-world security vulnerabilities. - GiovanniGatti/cve-bench…

14 views ·
#which#cves#should
THE HINDU — TOP

West Asia war LIVE: International benchmark oil prices cross $100 per barrel

Iran-U.S. war LIVE: Follow The Hindu for the latest updates as U.S. missiles strike western Iran for the 12th straight night and Tehran vows an 'eye for an eye' response.…

9 views ·
#west#asia#live
ARXIV.ORG

Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX

arXiv:2607.19353v1 Announce Type: new Abstract: Confidential computing is becoming a practical deployment requirement for AI inference workloads that process sensitive inputs or pr…

19 views ·
#benchmarking#confidential#inference
GOOGLE NEWS

OpenAI Models Escaped Locked Test Environment, Hacked Hugging Face to Cheat on Benchmark - Decrypt

OpenAI Models Escaped Locked Test Environment, Hacked Hugging Face to Cheat on Benchmark Decrypt…

33 views ·
GITHUB

I benchmarked AI cost-saving claims instead of trusting token percentages

Faster runtime for coding agents. Make coding agents 25% faster and 30% cheaper on average while keeping the quality same or more.. - lemoncrow-lab/lemoncrow…

30 views ·
#benchmarked#cost-saving#claims
ARXIV.ORG

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark

Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failu…

29 views ·
#coercion#deception#ai-to-ai
SUBSTACK

Semantic transactions: securing untrusted AI agent workflows at the OS boundary

Trust the system, not the prompt: Securing untrusted LLM tools with transactional boundaries and effect outboxes.…

33 views ·
#ai#security#transactions
ARXIV CS.AI

Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms

Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex back…

28 views ·
#event#stream#based
ARXIV CS.AI

OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents

Recent advancements in LVLMs necessitate robust benchmarks for complex, visually grounded reasoning. A critical limitation is identified in many document understanding benchmarks: …

29 views ·
#omnimapbench#benchmarking#visual-centric
ARXIV CS.AI

MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs

Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coheren…

20 views ·
#multiview-bench#diagnostic
ARXIV CS.AI

HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning

Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge. Existing evaluations are difficu…

25 views ·
#hero#heterogeneity-aware
ARXIV CS.AI

REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming

Large language models (LLMs) are increasingly applied to reverse-engineering tasks, and recent threat-intelligence reporting shows them operating inside live offensive-security wor…

25 views ·
#reforge#method#benchmarking
ARXIV.ORG

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

arXiv:2607.09322v1 Announce Type: new Abstract: In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluatio…

38 views ·
#longmedbench#benchmarking#medical
ARXIV.ORG

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

arXiv:2607.09142v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned…

30 views ·
#medrealmm#real-world#multimodal
DANLUU

Agentic test processes, LLM benchmarks, and other notes on agentic coding

111 views ·
#ai#testing#softwaredevelopment
SENIOR SWE-BENCH

Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

Evaluating agents as senior engineers on the work we actually give them…

39 views ·
#senior#swe-bench#open-source
THE HINDU

Roca Introduces Touch-T: A New Benchmark in Thermostatic Shower Systems

Roca Introduces Touch-T: A New Benchmark in Thermostatic Shower Systems…

27 views ·
#roca#introduces#touch-t
ARXIV.ORG

NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning

arXiv:2606.27826v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly deployed as embodied planners in egocentric environments, where task succe…

45 views ·
#normact#hidden
CURSOR

Reward hacking is swamping model intelligence gains

On SWE-bench Pro, 63% of successful Opus 4.8 Max resolutions retrieved the fix rather than derived it. Stricter eval harnesses show how benchmark scores can conflate coding ability…

36 views ·
#ai#machinelearning#codingbenchmarks
SOUTH CHINA MORNING POST

CK Asset sells penthouse in Hong Kong’s Mid-Levels for US$48.5m, sets pricing benchmark

The record per square foot pricing for new homes this year highlights the growing momentum in the city’s luxury property market.…

31 views ·
#realestate#hongkong#luxury
ARXIV.ORG

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

arXiv:2606.26350v1 Announce Type: new Abstract: Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented…

34 views ·
#artificialintelligence#machinelearning#quantitativefinance
ARXIV.ORG

Life After Benchmark Saturation: A Case Study of CORE-Bench

arXiv:2606.26158v1 Announce Type: new Abstract: When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach …

30 views ·
#life#saturation
I'VE DONE SOME THINGS

Will It Mythos?

OK, so Mythos finds really challenging security bugs, right? That’s why it’s cordoned off from the hoi polloi, to protect the world from such a powerful finder of exploits. I am sk…

81 views ·
#security#ai#benchmarking
GOOGLE NEWS

China keeps lending benchmark LPRs unchanged for 13th month in June - Reuters

China keeps lending benchmark LPRs unchanged for 13th month in June Reuters…

35 views ·
GITHUB

BEAVER: Enterprise benchmark for LLM Text-to-SQL from private data warehouses

48 views ·
#database#enterprise#technology
TECHMEME

Xiaomi releases MiMo Code V0.1.0, an open-source AI coding assistant that it says outperforms Claude Code on agentic coding and software engineering benchmarks (Carl Franzen/VentureBeat)

Carl Franzen / VentureBeat : Xiaomi releases MiMo Code V0.1.0, an open-source AI coding assistant that it says outperforms Claude Code on agentic coding and software engineering be…

44 views ·
TECHCRUNCH

Waymo says it built a better benchmark for comparing robotaxis to humans

Waymo created a new computer model to help it better understand how humans behave in crash scenarios that its robotaxis encounter.…

49 views ·
#waymo#says#built
ARXIV.ORG

Benchmarks in Leipzig

Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3…

49 views ·
#mathematics#artificial intelligence#research
GOOGLE NEWS

OpenAI research and product leads detail GPT-Rosalind capabilities and benchmarks - R&D World

Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…

34 views ·