Sens.aiAI signal desk
TodayRadarBriefing
Admin sign in
Sens.aiSensif.aiAI signal desk
ModelsResearchCountriesOpen vs ClosedComputeBetaCommentary
Loading…
TodayRadarBriefingAdmin

Radar

Research Radar

Which research topics are accelerating, and where earlier papers have already landed in shipped models.

Filtered to Evaluation & Benchmarks — 100 papersClear filter

More papers match this topic — see “Show more” below.

Papers this week

50

Trending topic

Agents

▲ 24%

papers this window vs prior window

Papers linked to models

148

linked by the desk, past 90 days

Median days paper → model

—

lower is faster

Topic velocity

papers this window by topic · Δ vs prior window · a paper can carry more than one topic

AgentsAgents
135▲24%
Evaluation & BenchmarksEvaluation & Benchmarks
38▲41%
MultimodalMultimodal
23▼30%
Safety & Alignment

Topic momentum

prior window → this window

Agents

135 this window ▲24%

Evaluation & Benchmarks

38 this window ▲41%

Multimodal

23 this window ▼30%

Where research lands: paper → model linkage

Links are AI-inferred by the desk's LLM from tech reports, system cards and citations — each carries a confidence level and is not a claim by the authors.

ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment

arXiv:2608.06110AI-inferred · high

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

arXiv:2608.05212AI-inferred · high

StepReflect: Structured UI Transition Reflection for Mobile GUI Agents

arXiv:2608.05587AI-inferred · high

Runtime Observability for Heterogeneous Attention Memory

Filtered to Evaluation & Benchmarks — 100 papersClear filter

More papers match this topic — see “Show more” below.

Papers on Evaluation & Benchmarks

Clear filter
mathematical reasoninglocal inferencepublished Aug 25, 2026 · arXiv:2608.22048

More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning

LinkedQwen 3.8 Max
language model evaluationstructured perturbationspublished Aug 25, 2026 · arXiv:2608.22138
Safety & Alignment
23▲53%
Training & OptimizationTraining & Optimization
21▲50%
Efficiency & InferenceEfficiency & Inference
20▼5%
ReasoningReasoning
15▲36%
InterpretabilityInterpretability
12▼20%
Reinforcement LearningReinforcement Learning
11▼21%
Retrieval & RAGRetrieval & RAG
9▼18%
Privacy & SecurityPrivacy & Security
8▲300%
Code GenerationCode Generation
8▼27%
Generative ModelsGenerative Models
7▼53%
Robotics & Embodied AIRobotics & Embodied AI
6▼14%
Language UnderstandingLanguage Understanding
5▲25%
Speech & AudioSpeech & Audio
5▼17%
Mixture of ExpertsMixture of Experts
3▲50%
Long ContextLong Context
3▼40%
Computer VisionComputer Vision
3▼63%
Data & Synthetic DataData & Synthetic Data
1▼75%
OtherOther
163▲11%
View as table
TopicPapers this windowPrior windowΔ
Agents135109+24%
Evaluation & Benchmarks3827+41%
Multimodal2333-30%
Safety & Alignment2315+53%
Training & Optimization2114+50%
Efficiency & Inference2021-5%
Reasoning1511+36%
Interpretability1215-20%
Reinforcement Learning1114-21%
Retrieval & RAG911-18%
Privacy & Security82+300%
Code Generation811-27%
Generative Models715-53%
Robotics & Embodied AI67-14%
Language Understanding54+25%
Speech & Audio56-17%
Mixture of Experts32+50%
Long Context35-40%
Computer Vision38-62%
Data & Synthetic Data14-75%
Other163147+11%

Safety & Alignment

23 this window ▲53%

Training & Optimization

21 this window ▲50%

Efficiency & Inference

20 this window ▼5%
arXiv:2608.05863
AI-inferred · high

Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

arXiv:2608.05411AI-inferred · high

C^3PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

arXiv:2608.05381AI-inferred · high

SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation

arXiv:2608.05628AI-inferred · high

Counterfactual Analysis via Large Language Models

arXiv:2608.05367AI-inferred · high

EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents

arXiv:2608.05519AI-inferred · high

Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints

arXiv:2608.06949AI-inferred · high

Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays

arXiv:2608.22068AI-inferred · high
GPT-5.68 links
DeepSeek V4-Flash1 link
Gemini 3.7 Flash2 links
claude-fable-5, claude-sonnet-5, claude-opus-52 links
Grok1 link
View as table
PaperLinked modelRelationConfidence
ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment (2608.06110)GPT-5.6evaluates85%
SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents (2608.05212)GPT-5.6evaluates98%
StepReflect: Structured UI Transition Reflection for Mobile GUI Agents (2608.05587)GPT-5.6evaluates95%
Runtime Observability for Heterogeneous Attention Memory (2608.05863)DeepSeek V4-Flashrelated90%
Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index (2608.05411)Gemini 3.7 Flashevaluates99%
Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index (2608.05411)GPT-5.6evaluates99%
C^3PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models (2608.05381)Gemini 3.7 Flashevaluates98%
SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation (2608.05628)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation (2608.05628)GPT-5.6evaluates98%
Counterfactual Analysis via Large Language Models (2608.05367)GPT-5.6evaluates98%
EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents (2608.05519)GPT-5.6evaluates95%
Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints (2608.06949)GPT-5.6evaluates98%
Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068)Grokevaluates96%
Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates96%
Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068)Gemini 3.7 Flashevaluates96%
Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068)GPT-5.6evaluates96%
MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning (2608.22167)GPT-5.6technique used95%
Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems (2608.22143)Qwen 3.8 Maxevaluates98%
Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems (2608.22143)Gemma 4evaluates98%
GenCoord: Skill-Path Commitments under Private Information (2608.22055)Qwen 3.8 Maxtechnique used98%
More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning (2608.22048)Qwen 3.8 Maxevaluates100%
Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents (2608.22191)Qwen 3.8 Maxevaluates98%
Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents (2608.22191)GPT-5.6evaluates98%
Redteaming Leading Arabic LLMs with ASAS (2608.21985)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Redteaming Leading Arabic LLMs with ASAS (2608.21985)GPT-5.6evaluates98%
Task-Driven 3D Printability Assistance via Geometry- and Knowledge-Grounded LLM Reasoning (2608.22128)Gemini 3.7 Flashevaluates95%
MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds (2608.22061)GPT-5.6evaluates95%
HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews (2608.21868)Qwen 3.8 Maxtechnique used99%
Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning (2608.21811)Qwen 3.8 Maxevaluates98%
HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries (2608.21792)Qwen 3.8 Maxtechnique used99%
ATHENA: Knowledge-guided agentic neural architecture search for AutoFormer-based electronic health record modeling (2608.21712)Athena-Brain-8Btechnique used99%
Context as an Environment: Programmatic Context Management for Long-Horizon Agents (2608.21690)Qwen 3.8 Maxevaluates95%
From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668)GPT-5.6evaluates98%
From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668)Gemini 3.7 Flashevaluates98%
Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning (2608.21501)Qwen 3.8 Maxevaluates99%
KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference (2608.21362)Qwen 3.8 Maxevaluates98%
K-Bench: measuring model performance on real scientific agent requests (2608.21601)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates100%
K-Bench: measuring model performance on real scientific agent requests (2608.21601)GPT-5.6evaluates100%
IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents (2608.06735)DeepSeek-V4-Flash-0731evaluates95%
NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs (2608.07167)GPT-5.6evaluates98%
WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader (2608.06474)GPT-5.6evaluates99%
Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578)GPT-5.6evaluates99%
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory (2608.07169)GPT-5.6 Lunatechnique used98%
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills (2608.07885)GPT-5.6 Lunaevaluates98%
SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents (2608.08055)DeepSeek-V4-Flash-0731technique used98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)DeepSeek-V4-Flash-0731evaluates98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)GLM-5.3evaluates98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)GPT-5.6 Lunaevaluates98%
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651)GPT-5.6 Lunaevaluates99%
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651)Gemini 3.7 Flashevaluates99%
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines (2608.07813)DeepSeek-V4-Flash-0731evaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)Grok 4.6evaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)GPT-5.6 Lunaevaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)Gemini 3.7 Flashevaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Can Legal AI Know When It Is Wrong? And Do Students Know When It Is? (2608.21089)GPT-5.6 Lunaevaluates99%
Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol (2608.20729)Qwen 3.8 Maxevaluates98%
Why2Speak: Faithful Reasoning for Abstaining Action Policies (2608.20670)Qwen 3.8 Maxevaluates98%
No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators (2608.20938)Qwen 3.8 Maxevaluates99%
Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation (2608.20797)Qwen 3.8 Maxtechnique used98%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)GPT-5.6 Lunaevaluates99%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)Qwen 3.8 Maxevaluates99%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)Gemma 4evaluates99%
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design (2608.20755)GPT-5.6 Lunatechnique used95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Gemma 4evaluates95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Mistral OCR 4evaluates95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Qwen 3.8 Maxevaluates95%
Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory (2608.20397)Qwen 3.8 Maxevaluates99%
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378)Gemma 4evaluates98%
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378)Qwen 3.8 Maxevaluates98%
PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure (2608.20342)claude-fable-5, claude-sonnet-5, claude-opus-5technique used98%
Personalized Privacy Control in LLMs via Attention Head Intervention (2608.21209)Qwen 3.8 Maxevaluates99%
TreeWY: Speculative Verification for Gated DeltaNet Hybrids (2608.20961)Qwen 3.8 Maxevaluates98%
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631)Gemma 4evaluates99%
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631)Qwen 3.8 Maxevaluates99%
FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth (2608.20574)Grok 4.6evaluates99%
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414)GPT-5.6 Lunaevaluates99%
SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation (2608.10775)GPT-5.6 Lunaevaluates98%
CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation (2608.10090)DeepSeek-V4-Flash-0731evaluates98%
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? (2608.10366)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence (2608.11341)GPT-5.6 Lunaevaluates98%
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research (2608.11216)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation (2608.11616)GPT-5.6 Lunatechnique used98%
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs (2608.11403)Qwen 3.8 27Bevaluates99%
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication (2608.11676)Qwen 3.8 27Bevaluates95%
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges (2608.12097)GPT-5.6 Lunaevaluates95%
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting (2608.11692)Qwen 3.8 27Bevaluates99%
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning (2608.11994)GPT-5.6 Lunaevaluates95%
From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate (2608.11381)Qwen 3.8 27Btechnique used98%
Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets (2608.11233)Qwen 3.8 27Btechnique used99%
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232)GPT-5.6 Lunaevaluates95%
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents (2608.11683)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (2608.12036)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates86%
Harnessing agent memory to build lifelong AI partners for materials scientists (2608.11224)GPT-5.6 Lunaevaluates95%
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343)GPT-5.6 Lunaevaluates99%
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS (2608.12249)claude-fable-5, claude-sonnet-5, claude-opus-5technique used95%
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration (2608.11210)Qwen 3.8 27Bevaluates95%
From Monolithic to Modular: Segment-level Automatic Prompt Optimization (2608.11219)GPT-5.6 Lunaevaluates98%
RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle (2608.11241)claude-fable-5, claude-sonnet-5, claude-opus-5technique used90%
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585)GPT-5.6 Lunaevaluates95%
Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes (2608.13420)Qwen 3.8 27Bevaluates98%
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing (2608.13156)Qwen 3.8 27Bevaluates95%
Jointly Predicting Courses and Grades Using a Transformer-Based Model (2608.13409)trace-supervised symbolic neural CPUtechnique used99%
Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) (2608.13063)Qwen 3.8 27Bevaluates99%
ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522)AlphaEvolverelated94%
ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522)GPT-5.6 Lunaevaluates98%
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents (2608.12476)Qwen 3.8 27Bevaluates99%
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs (2608.12675)Qwen 3.8 27Bevaluates95%
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence (2608.12928)GPT-5.6 Lunaevaluates99%
From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL (2608.13787)GPT-5.6 Lunaevaluates95%
Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation (2608.13754)GPT-5.6 Lunaevaluates99%
Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375)GPT-5.6 Lunaevaluates95%
MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends (2608.13883)GPT-5.6 Lunaevaluates95%
Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence (2608.13958)GPT-5.6 Lunaevaluates100%
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (2608.14290)Qwen 3.8 Maxrelated95%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)DeepSeek-V4-Flash-0731evaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)GPT-5.6 Lunaevaluates95%
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning (2608.14552)GPT-5.6 Lunaevaluates99%
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry (2608.14680)DeepSeek-V4-Flash-0731evaluates98%
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579)GPT-5.6 Lunatechnique used98%
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145)GPT-5.6 Lunaevaluates95%
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558)GPT-5.6 Lunaevaluates98%
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588)GPT-5.6 Lunaevaluates98%
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks (2608.14927)GPT-5.6 Lunaevaluates95%
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338)Qwen 3.8 Maxevaluates98%
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338)GPT-5.6 Lunaevaluates98%
When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models (2608.19230)GPT-5.6 Lunaevaluates98%
MidTool: Mid-training Data Synthesis for Agentic Tool Use (2608.20314)Qwen 3.8 Maxtechnique used98%
Automatic bioinformatic software named entity recognition from literature (2608.19201)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)Grokevaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)Gemini 3.7 Flashevaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)GPT-5.6 Lunaevaluates90%
Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping (2608.19220)GPT-5.6 Lunatechnique used99%
A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to Deployment (2608.19199)Athena-Brain-8Bevaluates98%
Phantom Gains: Auditing Self-Improvement Against a Measured Null (2608.20290)Qwen 3.8 Maxevaluates98%
ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974)Gemini 3.7 Flashevaluates95%
ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974)DeepSeek-V4-Flash-0731evaluates95%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)Gemini 3.7 Flashevaluates99%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)GPT-5.6 Lunaevaluates99%
SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning (2608.19842)Qwen 3.8 Maxevaluates98%
Can Agent Memory Systems Track Evolving State? (2608.19652)Qwen 3.8 Maxevaluates98%
Can Agent Memory Systems Track Evolving State? (2608.19652)DeepSeek-V4-Flash-0731evaluates98%
From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG (2608.19535)Qwen 3.8 Maxevaluates99%
Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluation (2608.19299)GPT-5.6 Lunaevaluates80%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Grokevaluates98%
Different Facets of Verbalised Overconfidence: an Interpretability Study (2608.18106)Qwen 3.8 Maxevaluates99%
DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models (2608.18103)DeepSeek V4-Flashtechnique used99%
Abliteration Mitigation via Refusal Aliases (2608.18093)Gemma 4evaluates99%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Gemma 4evaluates85%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Qwen 3.8 Maxevaluates85%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Mistral OCR 4evaluates85%
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089)Qwen 3.8 Maxevaluates99%
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089)Mistral OCR 4evaluates99%
Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication (2608.19161)Qwen 3.8 Maxevaluates98%
Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference (2608.18591)Qwen 3.8 Maxtechnique used100%
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389)Qwen 3.8 Maxevaluates98%
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair (2608.18324)Qwen 3.8 Maxevaluates98%
Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu (2608.18142)Mistral OCR 4evaluates99%
Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions (2608.18078)DeepSeek V4-Flashevaluates95%
Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS) (2608.18100)GPT-5.6 Lunaevaluates99%
Breaking the weakest link to evade vision language models (2608.18938)Qwen 3.8 Maxevaluates99%
Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models (2608.18884)Qwen 3.8 Maxevaluates98%
CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence (2608.18613)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates80%
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents (2608.18423)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates100%
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates90%
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336)Qwen 3.8 Maxevaluates98%
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates80%
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111)GPT-5.6 Lunaevaluates80%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)Qwen 3.8 Maxevaluates99%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)Gemini 3.7 Flashevaluates99%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)GPT-5.6 Lunaevaluates99%
Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis (2311.06273)GPT-5.6 Lunaevaluates95%
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050)GPT-5.6 Lunaevaluates98%
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050)Gemini 3.7 Flashevaluates98%
Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch (2608.17684)Qwen 3.8 Maxevaluates98%
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract (2608.16956)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing (2608.17638)GPT-5.6 Lunaevaluates95%
TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation (2608.17588)GPT-5.6 Lunaevaluates95%
Agent Lightning v1.0: Towards Harnessed Agentic RL (2608.17528)Qwen 3.8 Maxevaluates99%
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (2608.17393)Qwen 3.8 Maxevaluates99%
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration (2608.17336)Qwen 3.8 Maxevaluates98%
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning (2608.17301)Qwen 3.8 Maxevaluates98%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Grok 4.6evaluates99%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Gemini 3.7 Flashevaluates98%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)GPT-5.6 Lunaevaluates98%
FedPref: Federated Preference Learning for Structured Radiology Report Extraction (2608.16971)Qwen 3.8 Maxtechnique used99%
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification (2608.17247)GPT-5.6 Lunaevaluates98%
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents (2608.16890)GPT-5.6 Lunaevaluates99%
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents (2608.16890)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)DeepSeek V4-Flashevaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)GPT-5.6 Solevaluates99%
Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5 (2608.14992)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates100%
Small Models Scout Bottleneck Order for Large-Model Data Control (2608.14936)Qwen 3.8 Maxevaluates95%

Measuring Stability and Failure Behavior in Language Models Under Structured Perturbations

literature review agentsexpert evaluationpublished Aug 25, 2026 · arXiv:2608.21374

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

multiple-choice benchmarksleaderboard reliabilitypublished Aug 25, 2026 · arXiv:2608.21382

There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items

LLM-as-Judgetrust scoringpublished Aug 24, 2026 · arXiv:2608.21097

When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge

misinformation interventionscontent moderationpublished Aug 24, 2026 · arXiv:2608.20649

Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions

AI evaluator auditingcounterfactual reasoningpublished Aug 24, 2026 · arXiv:2608.20938

No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

LinkedQwen 3.8 Max
pathology foundation modelswhole-slide imagingpublished Aug 24, 2026 · arXiv:2608.21060

CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

automated evaluationlanguage-model rankingpublished Aug 24, 2026 · arXiv:2608.20574

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

LinkedGrok 4.6
large language modelslegal advicepublished Aug 21, 2026 · arXiv:2608.20220

InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries

legal contractscontract scrubbingpublished Aug 21, 2026 · arXiv:2608.20204

ContractScrub: A benchmark for final review of legal contracts

LLM memorymemory benchmarkspublished Aug 21, 2026 · arXiv:2608.20202

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

large language modelssocial simulationpublished Aug 21, 2026 · arXiv:2608.19689

Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

large language modelsair traffic controlpublished Aug 21, 2026 · arXiv:2608.19299

Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluation

LinkedGPT-5.6 Luna
LLM evaluationself-preferencepublished Aug 20, 2026 · arXiv:2608.18091

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

text-space skill optimizationreference-free LLM judgespublished Aug 20, 2026 · arXiv:2608.18719

Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization

structured decompositionlanguage model evaluationpublished Aug 20, 2026 · arXiv:2608.18303

SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition

LLM-as-a-Judgerecommendation explanationspublished Aug 20, 2026 · arXiv:2608.18300

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

AI benchmarksleaderboard governancepublished Aug 20, 2026 · arXiv:2608.18117

Position: AI Leaderboards Are Underserving the Global South: A Case Study from India

language model benchmarkingpartial-credit evaluationpublished Aug 20, 2026 · arXiv:2608.18336

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

Linkedclaude-fable-5, claude-sonnet-5, claude-opus-5Qwen 3.8 Max
medical consultationlarge language modelspublished Aug 19, 2026 · arXiv:2608.17330

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

time series foundation modelszero-shot forecastingpublished Aug 19, 2026 · arXiv:2608.17299

LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

AI evaluationautonomous scientific researchpublished Aug 19, 2026 · arXiv:2608.17271

ASI-Bench: At the Dawn of Artificial Superintelligence

large language modelshypothesis rankingpublished Aug 19, 2026 · arXiv:2608.17270

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

document intelligencevisual document parsingpublished Aug 18, 2026 · arXiv:2608.15064

LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents

large language modelsparenting advicepublished Aug 18, 2026 · arXiv:2608.14622

A Human-Centred Approach to Benchmarking LLMs for Parenting Advice

AI moral reasoningmoral value problempublished Aug 18, 2026 · arXiv:2608.14566

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

agentic systemsmodel routingpublished Aug 18, 2026 · arXiv:2608.14641

Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks

large language modelsmechanical engineeringpublished Aug 18, 2026 · arXiv:2608.14615

Large Language Models and their Awareness of Mechanics and Spatial Geometry

anchoring effectcognitive biaspublished Aug 17, 2026 · arXiv:2608.14320

AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs

large language modelsmmWave radarpublished Aug 17, 2026 · arXiv:2608.14179

Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding

enterprise data interfacessemantic planningpublished Aug 17, 2026 · arXiv:2608.13612

SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data

AI evaluationhuman-AI collaborationpublished Aug 17, 2026 · arXiv:2608.13577

AI Evaluation Should Work With Humans

autonomous systemsfailure discoverypublished Aug 17, 2026 · arXiv:2608.13719

Coverage Aware Active Evaluation for Failure Discovery with Paired Systems

mechanistic interpretabilityrepresentation learningpublished Aug 17, 2026 · arXiv:2608.13626

A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure

machine learningconstitutive modelingpublished Aug 17, 2026 · arXiv:2608.14063

Benchmarking data-driven material models on the classic Treloar dataset

Auto-Researchresearch process evaluationpublished Aug 15, 2026 · arXiv:2608.12788

ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

LLM-as-a-judgereasoning trace evaluationpublished Aug 15, 2026 · arXiv:2608.12585

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

LinkedGPT-5.6 Luna
autonomous reasoningsymbolic AIpublished Aug 15, 2026 · arXiv:2608.12325

Position: Reasoning is a Learnable Rule-Based Process

research integrityLLM evaluationpublished Aug 15, 2026 · arXiv:2608.12345

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

medical visual question answeringvision-language modelspublished Aug 15, 2026 · arXiv:2608.12928

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

LinkedGPT-5.6 Luna
mathematical reasoningautomated theorem provingpublished Aug 13, 2026 · arXiv:2608.11941

OEIS Open: How many conjectures can language models turn into theorems?

LLM evaluationinference-time token budgetspublished Aug 13, 2026 · arXiv:2608.12150

Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

large language modelsrecommendation evaluationpublished Aug 13, 2026 · arXiv:2608.11493

From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation

AI evaluationreal-world benchmarkspublished Aug 13, 2026 · arXiv:2608.11341

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

LinkedGPT-5.6 Luna
cybersecuritysecurity detectionpublished Aug 13, 2026 · arXiv:2509.16749

Evaluating LLM Generated Detection Rules in Cybersecurity

rubric-based evaluationLLM judgespublished Aug 13, 2026 · arXiv:2608.12097

Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges

LinkedGPT-5.6 Luna
financial reasoningstructured datapublished Aug 12, 2026 · arXiv:2608.11047

V-FiLLM: Verified Financial LLM Reasoning Benchmark

human action recognitionhierarchical intentionspublished Aug 12, 2026 · arXiv:2608.10765

Compositional Benchmark Synthesis for Hierarchical Human Action Recognition

artificial lifeneuroevolutionpublished Aug 12, 2026 · arXiv:2608.10323

Neuroevolution Arena: Nested Ecological Evaluation of Update-and-Inheritance Regimes across Neural Architectures

defensive LLMssocial engineeringpublished Aug 12, 2026 · arXiv:2608.10239

Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction

interactive analytic dashboardsdashboard generationpublished Aug 12, 2026 · arXiv:2608.10567

DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation

LLM evaluationmoral reasoningpublished Aug 11, 2026 · arXiv:2608.08061

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

benchmark contaminationAI evaluationpublished Aug 11, 2026 · arXiv:2608.07914

When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

medical reasoningelectronic health recordspublished Aug 11, 2026 · arXiv:2608.07796

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

LLM benchmarksbenchmark verificationpublished Aug 11, 2026 · arXiv:2608.07762

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

LinkedDeepSeek-V4-Flash-0731GLM-5.3GPT-5.6 Luna
large language modelsretrieval-grounded systemspublished Aug 11, 2026 · arXiv:2608.07688

IntelliAudit: Using Large Language Models to Evaluate Audit Controls

large language modelsdocument repairpublished Aug 11, 2026 · arXiv:2608.07617

TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

large language model evaluationgeospatial reasoningpublished Aug 10, 2026 · arXiv:2608.07411

GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks

large language modelscomputational creativitypublished Aug 10, 2026 · arXiv:2608.07243

Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models

LLM evaluationjudge panelspublished Aug 10, 2026 · arXiv:2608.06940

Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps

automated item evaluationeducational assessmentpublished Aug 10, 2026 · arXiv:2608.06609

Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques

human-AI coordinationcyber-physical systemspublished Aug 10, 2026 · arXiv:2608.06657

TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure

financial LLMsinvestment decision-makingpublished Aug 7, 2026 · arXiv:2608.06108

Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents

personalized large language modelsbehavioral personalizationpublished Aug 7, 2026 · arXiv:2608.05246

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

AI tutoringlarge language modelspublished Aug 7, 2026 · arXiv:2608.05411

Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

LinkedGemini 3.7 FlashGPT-5.6
factuality evaluationclaim decompositionpublished Aug 7, 2026 · arXiv:2608.05228

TriQua: Reconciling Granularity and Context in Factuality Evaluation

embodied intelligencephysics simulationpublished Aug 7, 2026 · arXiv:2608.05948

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

NAVSIMre-simulation benchmarkspublished Aug 6, 2026 · arXiv:2608.04896

When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit

large language modelselectroencephalographypublished Aug 6, 2026 · arXiv:2608.04156

BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

large language modelschess commentarypublished Aug 6, 2026 · arXiv:2608.04240

Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary

music quality assessmentreference-free evaluationpublished Aug 6, 2026 · arXiv:2608.04142

InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion

AI evaluationsimulated userspublished Aug 6, 2026 · arXiv:2608.04205

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

chain-of-thought promptingzero-shot promptingpublished Aug 5, 2026 · arXiv:2608.03550

Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve

AI evaluationbenchmark infrastructurepublished Aug 5, 2026 · arXiv:2608.02996

On the missing benchmarks layer and a potential solution

electrocardiographyatrial fibrillation detectionpublished Aug 5, 2026 · arXiv:2608.03597

FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection

AI scientistsscientific reasoningpublished Aug 5, 2026 · arXiv:2608.03569

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

long-document question answeringbenchmark correctionpublished Aug 5, 2026 · arXiv:2608.03397

MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc

large language modelsforecastingpublished Aug 5, 2026 · arXiv:2608.03416

AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction

AI for researchautonomous experimental designpublished Aug 5, 2026 · arXiv:2608.03501

Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

benchmark evaluationcapability measurementpublished Aug 5, 2026 · arXiv:2608.03219

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

scientific equation discoverysymbolic regressionpublished Aug 3, 2026 · arXiv:2607.28684

Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery

computational creativityinterpretive perspectivespublished Aug 3, 2026 · arXiv:2607.28644

Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework

clinical diagnosisemergency department encounterspublished Aug 3, 2026 · arXiv:2607.28788

EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses

federated learningfoundation modelspublished Aug 3, 2026 · arXiv:2607.28658

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

large language model evaluationoptimization modelspublished Aug 3, 2026 · arXiv:2607.29431

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

LLM evaluationbenchmark auditingpublished Aug 3, 2026 · arXiv:2607.28801

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

LLM-as-a-Judge evaluationhallucination detectionpublished Aug 3, 2026 · arXiv:2607.28641

The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?

AI Scientist systemsautomated peer reviewpublished Aug 3, 2026 · arXiv:2607.28631

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

language-model evaluationbenchmark validitypublished Jul 31, 2026 · arXiv:2607.26191

Position: Evaluation Scores Are Perishable Knowledge Claims

medical imagingradiologypublished Jul 31, 2026 · arXiv:2607.25589

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

synthetic usersLLM evaluationpublished Jul 31, 2026 · arXiv:2607.26348

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

AI evaluationbenchmark validitypublished Jul 31, 2026 · arXiv:2607.26159

When benchmark inferences do not compose: Projectibility in AI evaluation

large language modelschain-of-thought evaluationpublished Jul 31, 2026 · arXiv:2607.26102

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models

large language modelseducational assessmentpublished Jul 31, 2026 · arXiv:2607.26067

The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty

masked diffusion language modelsevaluation protocolspublished Jul 30, 2026 · arXiv:2607.24763

CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models

large language modelsparaphrase robustnesspublished Jul 28, 2026 · arXiv:2607.22554

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

parallel programminghigh-performance computingpublished Jul 28, 2026 · arXiv:2607.22588

ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation

phone call assistantstarget-oriented dialogue systemspublished Jul 28, 2026 · arXiv:2607.22635

CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants

phonon stabilitycomputational materials sciencepublished Jul 28, 2026 · arXiv:2607.22573

PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability

Show more