Radar
Which research topics are accelerating, and where earlier papers have already landed in shipped models.
More papers match this topic — see “Show more” below.
Papers this week
50
Trending topic
Agents
▲ 28%papers this window vs prior window
Papers linked to models
152
linked by the desk, past 90 days
Median days paper → model
—
lower is faster
papers this window by topic · Δ vs prior window · a paper can carry more than one topic
Links are AI-inferred by the desk's LLM from tech reports, system cards and citations — each carries a confidence level and is not a claim by the authors.
EDATracer: An Agentic Framework for Large-Scale EDA Artifact Analysis
arXiv:2608.04032AI-inferred · highMatrAIx: Simulating the World with 8.3 Billion Persona Agents
arXiv:2608.04205AI-inferred · highHallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
arXiv:2608.04240AI-inferred · highOctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
More papers match this topic — see “Show more” below.
| Topic | Papers this window | Prior window | Δ |
|---|---|---|---|
| Agents | 139 | 109 | +28% |
| Evaluation & Benchmarks | 37 | 27 | +37% |
| Multimodal | 25 | 33 | -24% |
| Safety & Alignment | 24 | 15 | +60% |
| Efficiency & Inference | 21 | 20 | +5% |
| Training & Optimization | 21 | 14 | +50% |
| Reasoning | 13 | 11 | +18% |
| Interpretability | 11 | 15 | -27% |
| Reinforcement Learning | 10 | 14 | -29% |
| Privacy & Security | 8 | 2 | +300% |
| Retrieval & RAG | 8 | 11 | -27% |
| Code Generation | 8 | 11 | -27% |
| Generative Models | 7 | 15 | -53% |
| Robotics & Embodied AI | 7 | 7 | 0% |
| Language Understanding | 5 | 4 | +25% |
| Speech & Audio | 5 | 6 | -17% |
| Mixture of Experts | 3 | 2 | +50% |
| Long Context | 3 | 5 | -40% |
| Computer Vision | 2 | 8 | -75% |
| Data & Synthetic Data | 1 | 4 | -75% |
| Other | 162 | 147 | +10% |
Safety & Alignment
Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools
arXiv:2608.04719AI-inferred · highArgus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning
arXiv:2608.05144AI-inferred · highNuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts
arXiv:2608.04030AI-inferred · highRecursive Synthesis for Long-Horizon Terminal Tasks
arXiv:2608.05466AI-inferred · highOtter: A Time-Aware, History-Conditioned Human Chess AI
arXiv:2608.05206AI-inferred · highECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation
arXiv:2608.05893AI-inferred · highECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment
arXiv:2608.06110AI-inferred · highSearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
arXiv:2608.05212AI-inferred · high| Paper | Linked model | Relation | Confidence |
|---|---|---|---|
| EDATracer: An Agentic Framework for Large-Scale EDA Artifact Analysis (2608.04032) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 85% |
| MatrAIx: Simulating the World with 8.3 Billion Persona Agents (2608.04205) | GPT-5.6 | technique used | 95% |
| MatrAIx: Simulating the World with 8.3 Billion Persona Agents (2608.04205) | claude-fable-5, claude-sonnet-5, claude-opus-5 | technique used | 95% |
| Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary (2608.04240) | GPT-5.6 | evaluates | 100% |
| OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling (2608.05141) | OctoLong-Instruct | evaluates | 98% |
| Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools (2608.04719) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 99% |
| Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning (2608.05144) | GPT-5.6 | evaluates | 98% |
| NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts (2608.04030) | Gemini 3.7 Flash | evaluates | 99% |
| NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts (2608.04030) | GPT-5.6 | evaluates | 99% |
| Recursive Synthesis for Long-Horizon Terminal Tasks (2608.05466) | Qwen 3.8 Max | technique used | 98% |
| Otter: A Time-Aware, History-Conditioned Human Chess AI (2608.05206) | Otter | related | 100% |
| ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation (2608.05893) | GPT-5.6 | technique used | 98% |
| ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment (2608.06110) | GPT-5.6 | evaluates | 85% |
| SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents (2608.05212) | GPT-5.6 | evaluates | 98% |
| StepReflect: Structured UI Transition Reflection for Mobile GUI Agents (2608.05587) | GPT-5.6 | evaluates | 95% |
| Runtime Observability for Heterogeneous Attention Memory (2608.05863) | DeepSeek V4-Flash | related | 90% |
| Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index (2608.05411) | Gemini 3.7 Flash | evaluates | 99% |
| Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index (2608.05411) | GPT-5.6 | evaluates | 99% |
| C^3PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models (2608.05381) | Gemini 3.7 Flash | evaluates | 98% |
| SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation (2608.05628) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation (2608.05628) | GPT-5.6 | evaluates | 98% |
| Counterfactual Analysis via Large Language Models (2608.05367) | GPT-5.6 | evaluates | 98% |
| EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents (2608.05519) | GPT-5.6 | evaluates | 95% |
| Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints (2608.06949) | GPT-5.6 | evaluates | 98% |
| Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068) | Grok | evaluates | 96% |
| Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 96% |
| Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068) | Gemini 3.7 Flash | evaluates | 96% |
| Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068) | GPT-5.6 | evaluates | 96% |
| MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning (2608.22167) | GPT-5.6 | technique used | 95% |
| Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems (2608.22143) | Qwen 3.8 Max | evaluates | 98% |
| Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems (2608.22143) | Gemma 4 | evaluates | 98% |
| GenCoord: Skill-Path Commitments under Private Information (2608.22055) | Qwen 3.8 Max | technique used | 98% |
| More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning (2608.22048) | Qwen 3.8 Max | evaluates | 100% |
| Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents (2608.22191) | Qwen 3.8 Max | evaluates | 98% |
| Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents (2608.22191) | GPT-5.6 | evaluates | 98% |
| Redteaming Leading Arabic LLMs with ASAS (2608.21985) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| Redteaming Leading Arabic LLMs with ASAS (2608.21985) | GPT-5.6 | evaluates | 98% |
| Task-Driven 3D Printability Assistance via Geometry- and Knowledge-Grounded LLM Reasoning (2608.22128) | Gemini 3.7 Flash | evaluates | 95% |
| MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds (2608.22061) | GPT-5.6 | evaluates | 95% |
| HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews (2608.21868) | Qwen 3.8 Max | technique used | 99% |
| Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning (2608.21811) | Qwen 3.8 Max | evaluates | 98% |
| HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries (2608.21792) | Qwen 3.8 Max | technique used | 99% |
| ATHENA: Knowledge-guided agentic neural architecture search for AutoFormer-based electronic health record modeling (2608.21712) | Athena-Brain-8B | technique used | 99% |
| Context as an Environment: Programmatic Context Management for Long-Horizon Agents (2608.21690) | Qwen 3.8 Max | evaluates | 95% |
| From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668) | GPT-5.6 | evaluates | 98% |
| From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668) | Gemini 3.7 Flash | evaluates | 98% |
| Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning (2608.21501) | Qwen 3.8 Max | evaluates | 99% |
| KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference (2608.21362) | Qwen 3.8 Max | evaluates | 98% |
| K-Bench: measuring model performance on real scientific agent requests (2608.21601) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 100% |
| K-Bench: measuring model performance on real scientific agent requests (2608.21601) | GPT-5.6 | evaluates | 100% |
| IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents (2608.06735) | DeepSeek-V4-Flash-0731 | evaluates | 95% |
| NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs (2608.07167) | GPT-5.6 | evaluates | 98% |
| WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader (2608.06474) | GPT-5.6 | evaluates | 99% |
| Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 99% |
| Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578) | GPT-5.6 | evaluates | 99% |
| Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory (2608.07169) | GPT-5.6 Luna | technique used | 98% |
| Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills (2608.07885) | GPT-5.6 Luna | evaluates | 98% |
| SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents (2608.08055) | DeepSeek-V4-Flash-0731 | technique used | 98% |
| Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762) | DeepSeek-V4-Flash-0731 | evaluates | 98% |
| Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762) | GLM-5.3 | evaluates | 98% |
| Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762) | GPT-5.6 Luna | evaluates | 98% |
| An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651) | GPT-5.6 Luna | evaluates | 99% |
| An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651) | Gemini 3.7 Flash | evaluates | 99% |
| When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines (2608.07813) | DeepSeek-V4-Flash-0731 | evaluates | 95% |
| The Authority Expectancy Effect in Multi-User Conflict (2608.08026) | Grok 4.6 | evaluates | 95% |
| The Authority Expectancy Effect in Multi-User Conflict (2608.08026) | GPT-5.6 Luna | evaluates | 95% |
| The Authority Expectancy Effect in Multi-User Conflict (2608.08026) | Gemini 3.7 Flash | evaluates | 95% |
| The Authority Expectancy Effect in Multi-User Conflict (2608.08026) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 95% |
| Can Legal AI Know When It Is Wrong? And Do Students Know When It Is? (2608.21089) | GPT-5.6 Luna | evaluates | 99% |
| Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol (2608.20729) | Qwen 3.8 Max | evaluates | 98% |
| Why2Speak: Faithful Reasoning for Abstaining Action Policies (2608.20670) | Qwen 3.8 Max | evaluates | 98% |
| No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators (2608.20938) | Qwen 3.8 Max | evaluates | 99% |
| Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation (2608.20797) | Qwen 3.8 Max | technique used | 98% |
| Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768) | GPT-5.6 Luna | evaluates | 99% |
| Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768) | Qwen 3.8 Max | evaluates | 99% |
| Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768) | Gemma 4 | evaluates | 99% |
| Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design (2608.20755) | GPT-5.6 Luna | technique used | 95% |
| DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717) | Gemma 4 | evaluates | 95% |
| DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717) | Mistral OCR 4 | evaluates | 95% |
| DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717) | Qwen 3.8 Max | evaluates | 95% |
| Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory (2608.20397) | Qwen 3.8 Max | evaluates | 99% |
| Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378) | Gemma 4 | evaluates | 98% |
| Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378) | Qwen 3.8 Max | evaluates | 98% |
| PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure (2608.20342) | claude-fable-5, claude-sonnet-5, claude-opus-5 | technique used | 98% |
| Personalized Privacy Control in LLMs via Attention Head Intervention (2608.21209) | Qwen 3.8 Max | evaluates | 99% |
| TreeWY: Speculative Verification for Gated DeltaNet Hybrids (2608.20961) | Qwen 3.8 Max | evaluates | 98% |
| Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631) | Gemma 4 | evaluates | 99% |
| Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631) | Qwen 3.8 Max | evaluates | 99% |
| FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth (2608.20574) | Grok 4.6 | evaluates | 99% |
| StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 99% |
| StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414) | GPT-5.6 Luna | evaluates | 99% |
| SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation (2608.10775) | GPT-5.6 Luna | evaluates | 98% |
| CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation (2608.10090) | DeepSeek-V4-Flash-0731 | evaluates | 98% |
| DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? (2608.10366) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence (2608.11341) | GPT-5.6 Luna | evaluates | 98% |
| AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research (2608.11216) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| MBA: Multimodal Benchmark and Agents for Real-World Business Ideation (2608.11616) | GPT-5.6 Luna | technique used | 98% |
| When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs (2608.11403) | Qwen 3.8 27B | evaluates | 99% |
| XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication (2608.11676) | Qwen 3.8 27B | evaluates | 95% |
| Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges (2608.12097) | GPT-5.6 Luna | evaluates | 95% |
| HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting (2608.11692) | Qwen 3.8 27B | evaluates | 99% |
| Claim-Level Reliability Assessment for Efficient Test-Time Reasoning (2608.11994) | GPT-5.6 Luna | evaluates | 95% |
| From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate (2608.11381) | Qwen 3.8 27B | technique used | 98% |
| Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets (2608.11233) | Qwen 3.8 27B | technique used | 99% |
| Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 95% |
| Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232) | GPT-5.6 Luna | evaluates | 95% |
| FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents (2608.11683) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (2608.12036) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 86% |
| Harnessing agent memory to build lifelong AI partners for materials scientists (2608.11224) | GPT-5.6 Luna | evaluates | 95% |
| Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 99% |
| Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343) | GPT-5.6 Luna | evaluates | 99% |
| An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS (2608.12249) | claude-fable-5, claude-sonnet-5, claude-opus-5 | technique used | 95% |
| Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration (2608.11210) | Qwen 3.8 27B | evaluates | 95% |
| From Monolithic to Modular: Segment-level Automatic Prompt Optimization (2608.11219) | GPT-5.6 Luna | evaluates | 98% |
| RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle (2608.11241) | claude-fable-5, claude-sonnet-5, claude-opus-5 | technique used | 90% |
| Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585) | GPT-5.6 Luna | evaluates | 95% |
| Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes (2608.13420) | Qwen 3.8 27B | evaluates | 98% |
| Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing (2608.13156) | Qwen 3.8 27B | evaluates | 95% |
| Jointly Predicting Courses and Grades Using a Transformer-Based Model (2608.13409) | trace-supervised symbolic neural CPU | technique used | 99% |
| Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) (2608.13063) | Qwen 3.8 27B | evaluates | 99% |
| ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522) | AlphaEvolve | related | 94% |
| ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522) | GPT-5.6 Luna | evaluates | 98% |
| Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents (2608.12476) | Qwen 3.8 27B | evaluates | 99% |
| Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs (2608.12675) | Qwen 3.8 27B | evaluates | 95% |
| Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence (2608.12928) | GPT-5.6 Luna | evaluates | 99% |
| From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL (2608.13787) | GPT-5.6 Luna | evaluates | 95% |
| Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation (2608.13754) | GPT-5.6 Luna | evaluates | 99% |
| Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375) | GPT-5.6 Luna | evaluates | 95% |
| MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends (2608.13883) | GPT-5.6 Luna | evaluates | 95% |
| Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence (2608.13958) | GPT-5.6 Luna | evaluates | 100% |
| Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (2608.14290) | Qwen 3.8 Max | related | 95% |
| StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089) | DeepSeek-V4-Flash-0731 | evaluates | 99% |
| StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089) | GPT-5.6 Luna | evaluates | 95% |
| Large Language Models Show Metacognitive Sensitivity in Medical Reasoning (2608.14552) | GPT-5.6 Luna | evaluates | 99% |
| When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry (2608.14680) | DeepSeek-V4-Flash-0731 | evaluates | 98% |
| SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579) | GPT-5.6 Luna | technique used | 98% |
| ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145) | GPT-5.6 Luna | evaluates | 95% |
| The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558) | GPT-5.6 Luna | evaluates | 98% |
| The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588) | GPT-5.6 Luna | evaluates | 98% |
| LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks (2608.14927) | GPT-5.6 Luna | evaluates | 95% |
| Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338) | Qwen 3.8 Max | evaluates | 98% |
| Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338) | GPT-5.6 Luna | evaluates | 98% |
| When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models (2608.19230) | GPT-5.6 Luna | evaluates | 98% |
| MidTool: Mid-training Data Synthesis for Agentic Tool Use (2608.20314) | Qwen 3.8 Max | technique used | 98% |
| Automatic bioinformatic software named entity recognition from literature (2608.19201) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 90% |
| Automatic bioinformatic software named entity recognition from literature (2608.19201) | Grok | evaluates | 90% |
| Automatic bioinformatic software named entity recognition from literature (2608.19201) | Gemini 3.7 Flash | evaluates | 90% |
| Automatic bioinformatic software named entity recognition from literature (2608.19201) | GPT-5.6 Luna | evaluates | 90% |
| Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping (2608.19220) | GPT-5.6 Luna | technique used | 99% |
| A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to Deployment (2608.19199) | Athena-Brain-8B | evaluates | 98% |
| Phantom Gains: Auditing Self-Improvement Against a Measured Null (2608.20290) | Qwen 3.8 Max | evaluates | 98% |
| ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974) | Gemini 3.7 Flash | evaluates | 95% |
| ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974) | DeepSeek-V4-Flash-0731 | evaluates | 95% |
| PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861) | Gemini 3.7 Flash | evaluates | 99% |
| PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 99% |
| PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861) | GPT-5.6 Luna | evaluates | 99% |
| SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning (2608.19842) | Qwen 3.8 Max | evaluates | 98% |
| Can Agent Memory Systems Track Evolving State? (2608.19652) | Qwen 3.8 Max | evaluates | 98% |
| Can Agent Memory Systems Track Evolving State? (2608.19652) | DeepSeek-V4-Flash-0731 | evaluates | 98% |
| From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG (2608.19535) | Qwen 3.8 Max | evaluates | 99% |
| Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluation (2608.19299) | GPT-5.6 Luna | evaluates | 80% |
| When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909) | Grok | evaluates | 98% |
| Different Facets of Verbalised Overconfidence: an Interpretability Study (2608.18106) | Qwen 3.8 Max | evaluates | 99% |
| DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models (2608.18103) | DeepSeek V4-Flash | technique used | 99% |
| Abliteration Mitigation via Refusal Aliases (2608.18093) | Gemma 4 | evaluates | 99% |
| Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090) | Gemma 4 | evaluates | 85% |
| Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090) | Qwen 3.8 Max | evaluates | 85% |
| Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090) | Mistral OCR 4 | evaluates | 85% |
| Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089) | Qwen 3.8 Max | evaluates | 99% |
| Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089) | Mistral OCR 4 | evaluates | 99% |
| Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication (2608.19161) | Qwen 3.8 Max | evaluates | 98% |
| Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference (2608.18591) | Qwen 3.8 Max | technique used | 100% |
| A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389) | Qwen 3.8 Max | evaluates | 98% |
| A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 98% |
| Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair (2608.18324) | Qwen 3.8 Max | evaluates | 98% |
| Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu (2608.18142) | Mistral OCR 4 | evaluates | 99% |
| Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions (2608.18078) | DeepSeek V4-Flash | evaluates | 95% |
| Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS) (2608.18100) | GPT-5.6 Luna | evaluates | 99% |
| Breaking the weakest link to evade vision language models (2608.18938) | Qwen 3.8 Max | evaluates | 99% |
| Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models (2608.18884) | Qwen 3.8 Max | evaluates | 98% |
| CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence (2608.18613) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 80% |
| FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents (2608.18423) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 100% |
| Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 90% |
| Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336) | Qwen 3.8 Max | evaluates | 98% |
| Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 80% |
| Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111) | GPT-5.6 Luna | evaluates | 80% |
| ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307) | Qwen 3.8 Max | evaluates | 99% |
| ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307) | Gemini 3.7 Flash | evaluates | 99% |
| ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307) | GPT-5.6 Luna | evaluates | 99% |
| Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis (2311.06273) | GPT-5.6 Luna | evaluates | 95% |
| StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050) | GPT-5.6 Luna | evaluates | 98% |
| StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050) | Gemini 3.7 Flash | evaluates | 98% |
| Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch (2608.17684) | Qwen 3.8 Max | evaluates | 98% |
| The Price of Thinking: Reasoning Effort as a Model-Specific API Contract (2608.16956) | claude-fable-5, claude-sonnet-5, claude-opus-5 | evaluates | 95% |
| Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing (2608.17638) | GPT-5.6 Luna | evaluates | 95% |
| TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation (2608.17588) | GPT-5.6 Luna | evaluates | 95% |
| Agent Lightning v1.0: Towards Harnessed Agentic RL (2608.17528) | Qwen 3.8 Max | evaluates | 99% |
| LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (2608.17393) | Qwen 3.8 Max | evaluates | 99% |
| TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration (2608.17336) | Qwen 3.8 Max | evaluates | 98% |