Sens.aiAI signal desk
TodayRadarBriefing
Admin sign in
Sens.aiSensif.aiAI signal desk
ModelsResearchCountriesOpen vs ClosedComputeBetaCommentary
Loading…
TodayRadarBriefingAdmin

Radar

Research Radar

Which research topics are accelerating, and where earlier papers have already landed in shipped models.

Filtered to Efficiency & Inference — 75 papersClear filter

More papers match this topic — see “Show more” below.

Papers this week

50

Trending topic

Agents

▲ 19%

papers this window vs prior window

Papers linked to models

145

linked by the desk, past 90 days

Median days paper → model

—

lower is faster

Topic velocity

papers this window by topic · Δ vs prior window · a paper can carry more than one topic

AgentsAgents
132▲19%
Evaluation & BenchmarksEvaluation & Benchmarks
39▲44%
Safety & AlignmentSafety & Alignment
24▲60%
Multimodal

Topic momentum

prior window → this window

Agents

132 this window ▲19%

Evaluation & Benchmarks

39 this window ▲44%

Safety & Alignment

24 this window ▲60%

Where research lands: paper → model linkage

Links are AI-inferred by the desk's LLM from tech reports, system cards and citations — each carries a confidence level and is not a claim by the authors.

Runtime Observability for Heterogeneous Attention Memory

arXiv:2608.05863AI-inferred · high

Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

arXiv:2608.05411AI-inferred · high

C^3PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

arXiv:2608.05381AI-inferred · high

SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation

Filtered to Efficiency & Inference — 75 papersClear filter

More papers match this topic — see “Show more” below.

Papers on Efficiency & Inference

Clear filter
MambaMamba-2published Aug 25, 2026 · arXiv:2608.21952

SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality

large language modelskernel optimizationpublished Aug 25, 2026 · arXiv:2608.21836

LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization

Multimodal
23▼30%
Training & OptimizationTraining & Optimization
21▲50%
Efficiency & InferenceEfficiency & Inference
20▼5%
ReasoningReasoning
15▲36%
InterpretabilityInterpretability
12▼20%
Reinforcement LearningReinforcement Learning
11▼21%
Retrieval & RAGRetrieval & RAG
9▼18%
Code GenerationCode Generation
8▼27%
Privacy & SecurityPrivacy & Security
7▲250%
Generative ModelsGenerative Models
7▼53%
Robotics & Embodied AIRobotics & Embodied AI
6▼14%
Language UnderstandingLanguage Understanding
5▲25%
Speech & AudioSpeech & Audio
5▼17%
Mixture of ExpertsMixture of Experts
3▲50%
Long ContextLong Context
3▼40%
Computer VisionComputer Vision
3▼63%
Data & Synthetic DataData & Synthetic Data
1▼75%
OtherOther
163▲11%
View as table
TopicPapers this windowPrior windowΔ
Agents132111+19%
Evaluation & Benchmarks3927+44%
Safety & Alignment2415+60%
Multimodal2333-30%
Training & Optimization2114+50%
Efficiency & Inference2021-5%
Reasoning1511+36%
Interpretability1215-20%
Reinforcement Learning1114-21%
Retrieval & RAG911-18%
Code Generation811-27%
Privacy & Security72+250%
Generative Models715-53%
Robotics & Embodied AI67-14%
Language Understanding54+25%
Speech & Audio56-17%
Mixture of Experts32+50%
Long Context35-40%
Computer Vision38-62%
Data & Synthetic Data14-75%
Other163147+11%

Multimodal

23 this window ▼30%

Training & Optimization

21 this window ▲50%

Efficiency & Inference

20 this window ▼5%
arXiv:2608.05628AI-inferred · high

Counterfactual Analysis via Large Language Models

arXiv:2608.05367AI-inferred · high

EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents

arXiv:2608.05519AI-inferred · high

Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints

arXiv:2608.06949AI-inferred · high

Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays

arXiv:2608.22068AI-inferred · high

MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning

arXiv:2608.22167AI-inferred · high
DeepSeek V4-Flash1 link
Gemini 3.7 Flash3 links
GPT-5.67 links
claude-fable-5, claude-sonnet-5, claude-opus-52 links
Grok1 link
View as table
PaperLinked modelRelationConfidence
Runtime Observability for Heterogeneous Attention Memory (2608.05863)DeepSeek V4-Flashrelated90%
Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index (2608.05411)Gemini 3.7 Flashevaluates99%
Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index (2608.05411)GPT-5.6evaluates99%
C^3PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models (2608.05381)Gemini 3.7 Flashevaluates98%
SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation (2608.05628)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation (2608.05628)GPT-5.6evaluates98%
Counterfactual Analysis via Large Language Models (2608.05367)GPT-5.6evaluates98%
EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents (2608.05519)GPT-5.6evaluates95%
Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints (2608.06949)GPT-5.6evaluates98%
Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068)Grokevaluates96%
Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates96%
Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068)Gemini 3.7 Flashevaluates96%
Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays (2608.22068)GPT-5.6evaluates96%
MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning (2608.22167)GPT-5.6technique used95%
Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems (2608.22143)Qwen 3.8 Maxevaluates98%
Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems (2608.22143)Gemma 4evaluates98%
GenCoord: Skill-Path Commitments under Private Information (2608.22055)Qwen 3.8 Maxtechnique used98%
More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning (2608.22048)Qwen 3.8 Maxevaluates100%
Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents (2608.22191)Qwen 3.8 Maxevaluates98%
Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents (2608.22191)GPT-5.6evaluates98%
Redteaming Leading Arabic LLMs with ASAS (2608.21985)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Redteaming Leading Arabic LLMs with ASAS (2608.21985)GPT-5.6evaluates98%
Task-Driven 3D Printability Assistance via Geometry- and Knowledge-Grounded LLM Reasoning (2608.22128)Gemini 3.7 Flashevaluates95%
MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds (2608.22061)GPT-5.6evaluates95%
HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews (2608.21868)Qwen 3.8 Maxtechnique used99%
Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning (2608.21811)Qwen 3.8 Maxevaluates98%
HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries (2608.21792)Qwen 3.8 Maxtechnique used99%
ATHENA: Knowledge-guided agentic neural architecture search for AutoFormer-based electronic health record modeling (2608.21712)Athena-Brain-8Btechnique used99%
Context as an Environment: Programmatic Context Management for Long-Horizon Agents (2608.21690)Qwen 3.8 Maxevaluates95%
From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668)GPT-5.6evaluates98%
From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation (2608.21668)Gemini 3.7 Flashevaluates98%
Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning (2608.21501)Qwen 3.8 Maxevaluates99%
KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference (2608.21362)Qwen 3.8 Maxevaluates98%
K-Bench: measuring model performance on real scientific agent requests (2608.21601)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates100%
K-Bench: measuring model performance on real scientific agent requests (2608.21601)GPT-5.6evaluates100%
IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents (2608.06735)DeepSeek-V4-Flash-0731evaluates95%
NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs (2608.07167)GPT-5.6evaluates98%
WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader (2608.06474)GPT-5.6evaluates99%
Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
Divergent Response Modes in Frontier Language Models Under Steering Pressure (2608.06578)GPT-5.6evaluates99%
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory (2608.07169)GPT-5.6 Lunatechnique used98%
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills (2608.07885)GPT-5.6 Lunaevaluates98%
SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents (2608.08055)DeepSeek-V4-Flash-0731technique used98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)DeepSeek-V4-Flash-0731evaluates98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)GLM-5.3evaluates98%
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation (2608.07762)GPT-5.6 Lunaevaluates98%
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651)GPT-5.6 Lunaevaluates99%
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography (2608.07651)Gemini 3.7 Flashevaluates99%
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines (2608.07813)DeepSeek-V4-Flash-0731evaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)Grok 4.6evaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)GPT-5.6 Lunaevaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)Gemini 3.7 Flashevaluates95%
The Authority Expectancy Effect in Multi-User Conflict (2608.08026)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Can Legal AI Know When It Is Wrong? And Do Students Know When It Is? (2608.21089)GPT-5.6 Lunaevaluates99%
Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol (2608.20729)Qwen 3.8 Maxevaluates98%
Why2Speak: Faithful Reasoning for Abstaining Action Policies (2608.20670)Qwen 3.8 Maxevaluates98%
No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators (2608.20938)Qwen 3.8 Maxevaluates99%
Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation (2608.20797)Qwen 3.8 Maxtechnique used98%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)GPT-5.6 Lunaevaluates99%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)Qwen 3.8 Maxevaluates99%
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization (2608.20768)Gemma 4evaluates99%
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design (2608.20755)GPT-5.6 Lunatechnique used95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Gemma 4evaluates95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Mistral OCR 4evaluates95%
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning (2608.20717)Qwen 3.8 Maxevaluates95%
Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory (2608.20397)Qwen 3.8 Maxevaluates99%
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378)Gemma 4evaluates98%
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification (2608.20378)Qwen 3.8 Maxevaluates98%
PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure (2608.20342)claude-fable-5, claude-sonnet-5, claude-opus-5technique used98%
Personalized Privacy Control in LLMs via Attention Head Intervention (2608.21209)Qwen 3.8 Maxevaluates99%
TreeWY: Speculative Verification for Gated DeltaNet Hybrids (2608.20961)Qwen 3.8 Maxevaluates98%
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631)Gemma 4evaluates99%
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents (2608.20631)Qwen 3.8 Maxevaluates99%
FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth (2608.20574)Grok 4.6evaluates99%
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models (2608.20414)GPT-5.6 Lunaevaluates99%
SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation (2608.10775)GPT-5.6 Lunaevaluates98%
CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation (2608.10090)DeepSeek-V4-Flash-0731evaluates98%
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? (2608.10366)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence (2608.11341)GPT-5.6 Lunaevaluates98%
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research (2608.11216)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation (2608.11616)GPT-5.6 Lunatechnique used98%
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs (2608.11403)Qwen 3.8 27Bevaluates99%
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication (2608.11676)Qwen 3.8 27Bevaluates95%
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges (2608.12097)GPT-5.6 Lunaevaluates95%
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting (2608.11692)Qwen 3.8 27Bevaluates99%
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning (2608.11994)GPT-5.6 Lunaevaluates95%
From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate (2608.11381)Qwen 3.8 27Btechnique used98%
Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets (2608.11233)Qwen 3.8 27Btechnique used99%
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs (2608.11232)GPT-5.6 Lunaevaluates95%
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents (2608.11683)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence (2608.12036)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates86%
Harnessing agent memory to build lifelong AI partners for materials scientists (2608.11224)GPT-5.6 Lunaevaluates95%
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval (2608.11343)GPT-5.6 Lunaevaluates99%
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS (2608.12249)claude-fable-5, claude-sonnet-5, claude-opus-5technique used95%
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration (2608.11210)Qwen 3.8 27Bevaluates95%
From Monolithic to Modular: Segment-level Automatic Prompt Optimization (2608.11219)GPT-5.6 Lunaevaluates98%
RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle (2608.11241)claude-fable-5, claude-sonnet-5, claude-opus-5technique used90%
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces (2608.12585)GPT-5.6 Lunaevaluates95%
Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes (2608.13420)Qwen 3.8 27Bevaluates98%
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing (2608.13156)Qwen 3.8 27Bevaluates95%
Jointly Predicting Courses and Grades Using a Transformer-Based Model (2608.13409)trace-supervised symbolic neural CPUtechnique used99%
Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) (2608.13063)Qwen 3.8 27Bevaluates99%
ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522)AlphaEvolverelated94%
ε-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution (2608.12522)GPT-5.6 Lunaevaluates98%
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents (2608.12476)Qwen 3.8 27Bevaluates99%
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs (2608.12675)Qwen 3.8 27Bevaluates95%
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence (2608.12928)GPT-5.6 Lunaevaluates99%
From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL (2608.13787)GPT-5.6 Lunaevaluates95%
Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation (2608.13754)GPT-5.6 Lunaevaluates99%
Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (2608.14375)GPT-5.6 Lunaevaluates95%
MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends (2608.13883)GPT-5.6 Lunaevaluates95%
Implementing Computational Law in Wolfram Language for the Governance of Artificial Intelligence (2608.13958)GPT-5.6 Lunaevaluates100%
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (2608.14290)Qwen 3.8 Maxrelated95%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)DeepSeek-V4-Flash-0731evaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)GPT-5.6 Lunaevaluates95%
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning (2608.14552)GPT-5.6 Lunaevaluates99%
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry (2608.14680)DeepSeek-V4-Flash-0731evaluates98%
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization (2608.14579)GPT-5.6 Lunatechnique used98%
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models (2608.15145)GPT-5.6 Lunaevaluates95%
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning (2608.14558)GPT-5.6 Lunaevaluates98%
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588)GPT-5.6 Lunaevaluates98%
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks (2608.14927)GPT-5.6 Lunaevaluates95%
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338)Qwen 3.8 Maxevaluates98%
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability (2608.19338)GPT-5.6 Lunaevaluates98%
When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models (2608.19230)GPT-5.6 Lunaevaluates98%
MidTool: Mid-training Data Synthesis for Agentic Tool Use (2608.20314)Qwen 3.8 Maxtechnique used98%
Automatic bioinformatic software named entity recognition from literature (2608.19201)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)Grokevaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)Gemini 3.7 Flashevaluates90%
Automatic bioinformatic software named entity recognition from literature (2608.19201)GPT-5.6 Lunaevaluates90%
Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping (2608.19220)GPT-5.6 Lunatechnique used99%
A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to Deployment (2608.19199)Athena-Brain-8Bevaluates98%
Phantom Gains: Auditing Self-Improvement Against a Measured Null (2608.20290)Qwen 3.8 Maxevaluates98%
ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974)Gemini 3.7 Flashevaluates95%
ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2608.19974)DeepSeek-V4-Flash-0731evaluates95%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)Gemini 3.7 Flashevaluates99%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents (2608.19861)GPT-5.6 Lunaevaluates99%
SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning (2608.19842)Qwen 3.8 Maxevaluates98%
Can Agent Memory Systems Track Evolving State? (2608.19652)Qwen 3.8 Maxevaluates98%
Can Agent Memory Systems Track Evolving State? (2608.19652)DeepSeek-V4-Flash-0731evaluates98%
From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG (2608.19535)Qwen 3.8 Maxevaluates99%
Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluation (2608.19299)GPT-5.6 Lunaevaluates80%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Grokevaluates98%
Different Facets of Verbalised Overconfidence: an Interpretability Study (2608.18106)Qwen 3.8 Maxevaluates99%
DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models (2608.18103)DeepSeek V4-Flashtechnique used99%
Abliteration Mitigation via Refusal Aliases (2608.18093)Gemma 4evaluates99%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Gemma 4evaluates85%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Qwen 3.8 Maxevaluates85%
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities (2608.18090)Mistral OCR 4evaluates85%
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089)Qwen 3.8 Maxevaluates99%
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining (2608.18089)Mistral OCR 4evaluates99%
Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication (2608.19161)Qwen 3.8 Maxevaluates98%
Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference (2608.18591)Qwen 3.8 Maxtechnique used100%
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389)Qwen 3.8 Maxevaluates98%
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations (2608.18389)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates98%
Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair (2608.18324)Qwen 3.8 Maxevaluates98%
Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu (2608.18142)Mistral OCR 4evaluates99%
Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions (2608.18078)DeepSeek V4-Flashevaluates95%
Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS) (2608.18100)GPT-5.6 Lunaevaluates99%
Breaking the weakest link to evade vision language models (2608.18938)Qwen 3.8 Maxevaluates99%
Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models (2608.18884)Qwen 3.8 Maxevaluates98%
CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence (2608.18613)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates80%
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents (2608.18423)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates100%
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates90%
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme (2608.18336)Qwen 3.8 Maxevaluates98%
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates80%
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry (2608.18111)GPT-5.6 Lunaevaluates80%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)Qwen 3.8 Maxevaluates99%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)Gemini 3.7 Flashevaluates99%
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents (2608.18307)GPT-5.6 Lunaevaluates99%
Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis (2311.06273)GPT-5.6 Lunaevaluates95%
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050)GPT-5.6 Lunaevaluates98%
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents (2608.18050)Gemini 3.7 Flashevaluates98%
Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch (2608.17684)Qwen 3.8 Maxevaluates98%
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract (2608.16956)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates95%
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing (2608.17638)GPT-5.6 Lunaevaluates95%
TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation (2608.17588)GPT-5.6 Lunaevaluates95%
Agent Lightning v1.0: Towards Harnessed Agentic RL (2608.17528)Qwen 3.8 Maxevaluates99%
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (2608.17393)Qwen 3.8 Maxevaluates99%
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration (2608.17336)Qwen 3.8 Maxevaluates98%
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning (2608.17301)Qwen 3.8 Maxevaluates98%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Grok 4.6evaluates99%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)Gemini 3.7 Flashevaluates98%
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice (2608.16909)GPT-5.6 Lunaevaluates98%
FedPref: Federated Preference Learning for Structured Radiology Report Extraction (2608.16971)Qwen 3.8 Maxtechnique used99%
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification (2608.17247)GPT-5.6 Lunaevaluates98%
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents (2608.16890)GPT-5.6 Lunaevaluates99%
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents (2608.16890)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)DeepSeek V4-Flashevaluates99%
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling (2608.15089)GPT-5.6 Solevaluates99%
Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5 (2608.14992)claude-fable-5, claude-sonnet-5, claude-opus-5evaluates100%
Small Models Scout Bottleneck Order for Large-Model Data Control (2608.14936)Qwen 3.8 Maxevaluates95%
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks (2608.14927)GPT-5.6 Soltechnique used98%
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588)Qwen 3.8 Maxevaluates98%
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines (2608.14588)GPT-5.6 Solevaluates99%
KV cache reuseprefill latency reductionpublished Aug 25, 2026 · arXiv:2608.21362

KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference

LinkedQwen 3.8 Max
speculative decodingGated DeltaNetpublished Aug 24, 2026 · arXiv:2608.20961

TreeWY: Speculative Verification for Gated DeltaNet Hybrids

LinkedQwen 3.8 Max
causal inferencespillover effectspublished Aug 21, 2026 · arXiv:2608.19224

Causal Inference under Interference with Learned Exposure Mappings

multi-head attentionhead-wise context allocationpublished Aug 21, 2026 · arXiv:2608.19203

Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention

AI model routingvalue-of-information estimationpublished Aug 21, 2026 · arXiv:2608.20316

Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation

on-policy distillationlanguage-model post-trainingpublished Aug 21, 2026 · arXiv:2608.19408

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

KV cachingtransformer inferencepublished Aug 20, 2026 · arXiv:2608.18098

Fractional Decay KV-Cache: Ownership-Aware Memory Management for Improved Inference Relevancy in Dialog Systems

large language modelsreasoningpublished Aug 20, 2026 · arXiv:2608.18884

Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models

LinkedQwen 3.8 Max
long-context inferencemixed-precision attentionpublished Aug 19, 2026 · arXiv:2608.17336

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

LinkedQwen 3.8 Max
test-time inferencerollout pruningpublished Aug 18, 2026 · arXiv:2608.15065

Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning

FLOPs measurementAI efficiencypublished Aug 18, 2026 · arXiv:2608.14550

FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment

video world modelsinference-time adaptationpublished Aug 18, 2026 · arXiv:2608.15043

SCOPE: Score-Isolated Agentic Optimization for Video World Models

training-free post-training quantizationlow-bit W4A4 quantizationpublished Aug 17, 2026 · arXiv:2608.14149

QuaSAR: Quantization Compensation via Stable Activation-Aware Rank Truncation

deep neural network inferencemobile computingpublished Aug 17, 2026 · arXiv:2608.13863

Joint Optimization of Memory and Computing Frequency for Energy-Efficient DNN Inference

neural architecture searchhardware-aware NASpublished Aug 15, 2026 · arXiv:2608.13293

NAS-Driven Hardware Accelerator Exploration for Edge AI and Quantization Effects on the Pareto Space

speculative decodingedge computingpublished Aug 15, 2026 · arXiv:2608.13076

SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference

model mergingconflict-aware sparsificationpublished Aug 15, 2026 · arXiv:2608.12842

CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation

large language model inferenceprefill-decode separationpublished Aug 15, 2026 · arXiv:2608.12385

Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation

probabilities of causationpartial identificationpublished Aug 15, 2026 · arXiv:2608.12657

General Probabilities of Causation with Causal Knowledge

LLM reasoningprocess-level evaluationpublished Aug 15, 2026 · arXiv:2608.13221

TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems

large language modelsTransformerspublished Aug 15, 2026 · arXiv:2608.13156

Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing

LinkedQwen 3.8 27B
constrained decodingtrie automatapublished Aug 15, 2026 · arXiv:2608.12574

Trie Automata for Constrained Decoding over Large Finite Sets

self-consistencymajority votingpublished Aug 13, 2026 · arXiv:2608.11403

When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs

LinkedQwen 3.8 27B
vector quantizationAI infrastructurepublished Aug 13, 2026 · arXiv:2608.11240

VQ-bench: A Composable Vector Quantization Framework

position-independent cachinghybrid LLMspublished Aug 13, 2026 · arXiv:2608.11231

LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs

reinforcement learningLLM trainingpublished Aug 13, 2026 · arXiv:2608.11226

Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet

adaptive neuro-fuzzy inference systemshyperbolic geometrypublished Aug 13, 2026 · arXiv:2608.11768

HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry

temporal knowledge graphslink predictionpublished Aug 12, 2026 · arXiv:2608.10668

FITTER: Vocabulary-Agnostic Cross-Domain Inference on Temporal Knowledge Graphs

weight-only quantizationsignal-to-noise ratiopublished Aug 11, 2026 · arXiv:2608.08188

Quantization Degradation in Large Language Models: A Signal-Noise Perspective

on-policy self-distillationlarge language modelspublished Aug 11, 2026 · arXiv:2608.08176

Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation

active inferencehuman drivingpublished Aug 11, 2026 · arXiv:2608.07480

Emotion in an active inference model of human driving

probabilistic logic programmingstatistical relational artificial intelligencepublished Aug 10, 2026 · arXiv:2608.07230

From probability to causality in probabilistic logic programming

spiking neural networkspost-training quantizationpublished Aug 10, 2026 · arXiv:2608.07066

PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks

EEG foundation modelstoken poolingpublished Aug 10, 2026 · arXiv:2608.07033

ZIPBrain: Can EEG Foundation Models Be Faster, Locally Deployable, but Accurate?

post-training quantizationmodel compressionpublished Aug 10, 2026 · arXiv:2608.07019

ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization

knowledge graphslarge language modelspublished Aug 10, 2026 · arXiv:2608.06752

Mind the Gap: A Dual Knowledge Graph Framework for Unified Multi-task User Intent Inference

Mixture-of-Recursionsgenomicspublished Aug 10, 2026 · arXiv:2608.06727

bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning

survey-country metadatarandomized auditspublished Aug 7, 2026 · arXiv:2608.06085

Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference

large reasoning modelschain-of-thought reasoningpublished Aug 6, 2026 · arXiv:2608.04771

Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning

top-k selectionsparse modelspublished Aug 6, 2026 · arXiv:2608.04057

LaPrune: Controllable Differentiable Sparsity at Million Scale

post-training quantizationmulti-precision representationpublished Aug 6, 2026 · arXiv:2608.04048

Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs

interoceptionprecision allocationpublished Aug 6, 2026 · arXiv:2608.04232

Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent

federated learninglow-rank adaptationpublished Aug 5, 2026 · arXiv:2608.03605

FraQ: Efficient Coordinate-Space Recompression for Federated Low-Rank Adaptation

parameter-efficient fine-tuninglocal credit assignmentpublished Aug 5, 2026 · arXiv:2608.03020

LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment

LLM pruningcompression statisticspublished Aug 5, 2026 · arXiv:2608.02940

When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning

computational cognitive modelingBayesian inverse planningpublished Aug 3, 2026 · arXiv:2607.28894

Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design

KV-cache quantizationruntime observabilitypublished Aug 3, 2026 · arXiv:2607.28699

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization

knowledge distillationbiaspublished Aug 3, 2026 · arXiv:2607.28639

The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models

disaggregated LLM inferenceKV-cache transferpublished Aug 3, 2026 · arXiv:2607.28633

Topology-Aware Data Movement for Disaggregated GPU Inference

large language modelsagentic evolutionpublished Jul 31, 2026 · arXiv:2607.26661

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

KV-cache optimizationlong-context generationpublished Jul 30, 2026 · arXiv:2607.24788

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference

generative recommendationsemantic IDspublished Jul 30, 2026 · arXiv:2607.24995

Understanding Semantic IDs: From Item Representation to Item Selection in Generative Recommendation

large reasoning modelschain-of-thought internalizationpublished Jul 28, 2026 · arXiv:2607.22629

Masked Distillation: Internalizing the Chain-of-Thought in Language Models

mixture-of-expertsmixture-of-LLMspublished Jul 28, 2026 · arXiv:2607.22577

cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs

inference-time scalinglarge language model reasoningpublished Jul 28, 2026 · arXiv:2607.22602

DeepLook: Deeper Thinking with Lookahead

prefix cachingKV cache replicationpublished Jul 28, 2026 · arXiv:2607.22648

PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving

large language modelsstructured pruningpublished Jul 28, 2026 · arXiv:2607.22583

Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization

convolutional neural networksfeature-map pruningpublished Jul 28, 2026 · arXiv:2607.22564

Loss-Aware Feature-Map Pruning in Convolutional Neural Networks Using Multi-Armed Bandits

on-device LLMsprompt engineeringpublished Jul 28, 2026 · arXiv:2607.22568

Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting

LLM-as-a-judgeprogram distillationpublished Jul 28, 2026 · arXiv:2607.22561

Codifying the Judge: Scalable Evaluation via Program Distillation

text-attributed graphsgraph neural networkspublished Jul 24, 2026 · arXiv:2607.20477

Semi-Supervised Text-Attributed Graph Distillation

lightweight large language modelsedge and mobile deploymentpublished Jul 24, 2026 · arXiv:2607.20806

Profiling Lightweight Large Language Models

large language model inferenceprogram-of-thought reasoningpublished Jul 24, 2026 · arXiv:2607.20507

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference

electric power distribution systemsgrid topology identificationpublished Jul 24, 2026 · arXiv:2607.20480

Enabling Scalable Topology Inference in Distribution Systems via Constrained Multi-Source Inference

LLM inferencesamplingpublished Jul 24, 2026 · arXiv:2607.20475

SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification

single-cell datadataset distillationpublished Jul 23, 2026 · arXiv:2607.19426

Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min-Max Selection

bit-serial computingsparse inferencepublished Jul 23, 2026 · arXiv:2607.19431

BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator

Mathematics of Arraystransformer attentionpublished Jul 23, 2026 · arXiv:2607.19456

MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel

large reasoning modelsoverthinking mitigationpublished Jul 23, 2026 · arXiv:2607.19962

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization

small language modelsquantizationpublished Jul 23, 2026 · arXiv:2607.20129

CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning

distributed LLM inferencemalicious nodespublished Jul 23, 2026 · arXiv:2607.19490

Integrity of peer-to-peer distributed LLM inference under malicious nodes

LLM inferenceApple M5 Neural Acceleratorspublished Jul 23, 2026 · arXiv:2607.19438

BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators

Low-Rank Adaptationadaptive rank allocationpublished Jul 23, 2026 · arXiv:2607.19391

LAARA: Layer-Aware Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning

Show more