Today
A ranked briefing from the ai signal desk.
Today's summary
Custom inference silicon is becoming a direct lever on agent economics and latency
The AI scaling bottleneck is broadening from accelerators to the system around them
↳ Linked from your briefing
What changed: The proposed framework reports finer-grained ranking and diagnosis than aggregate scores, measures joint success across requirements, and supports cost-aware, training-free routing with 21.3% lower GPU-seconds per metric point than the cited ERNIE estimate.
What changed: According to the paper's reported results, a three-model medoid crowd outperformed voting across 25 models on two prediction benchmarks while reducing model calls by 88% and inference cost by approximately 80%.
What changed: The work provides a publicly released benchmark, data, and code for evaluating raw-signal reasoning rather than analysis of preprocessed features, and reports score improvements of 3.8 to 17.6 points from ReconPilot across tested backbone combinations.
What changed: The reported results indicate that reader-facing artifact design can materially alter measured performance, with matched-budget resolved packets outperforming recency-truncated raw dialogue by 42.4-72.6 points across 500 LongMemEval questions and nine models.
What changed: On the reported evaluations, the method recovered 93% of failed scenarios, increased Qwen3-4B-Instruct-2507 Pass^1 from 0.132 to 0.529, improved Gemma 4 E4B-it by 7.2 percentage points on BFCL v4 multi-turn, and improved deployed and on-device model performance.
What changed: The authors report that the method avoids task-specific teacher fine-tuning and improves performance over prior 3B-parameter reinforcement-learning baselines across seven question-answering benchmarks, including gains of 13.1% on HotpotQA and 8.5% on 2WikiMultihopQA.
What changed: It extends NL2SQL evaluation beyond simplified academic schemas by testing enterprise-scale complexity, cross-dialect generalization, and cases where executed queries return semantically incorrect results.
What changed: The reported system reduces active KV-cache memory by up to 3.50x relative to BF16 and 1.75x relative to FP8 while preserving live-request page addressability and maintaining near-control throughput in a limited canary test.
What changed: It introduces input- and task-dependent memory allocation across layer-stage nodes instead of static or random memory routing, reporting improved preservation of beneficial memory effects and suppression of regressions on biological and chemical reasoning benchmarks.
What changed: Under a shared evaluation harness, BRANCH improved accuracy in all 14 model-benchmark settings by an average of 5.98 percentage points and was the best-performing operator in 12 settings; the paper also identifies truncation recovery and paired scoring as important factors in evaluation outcomes.
What changed: The reported method achieves at least 0.997 held-out replay fitness across twelve public datasets, outperforms Agent Workflow Memory on matched next-step prediction datasets, and reaches held-out failure-prediction AUROC of up to 0.94.
What changed: It identifies detection gating and catalogue-derived metadata as dominant causal influences, quantifies the resulting tomographic redshift bias, and reports that the effect increases with model scale while withholding the detection channel removes it without measurable performance cost.
What changed: The proposed method aims to improve the accuracy-efficiency trade-off of agentic RAG without adding inference-time overhead from belief probing.
What changed: It proposes a shareable policy format and accompanying open-source tools intended to connect model alignment, testing, and runtime monitoring.
What changed: It provides a structured evaluation framework for detecting semantic drift in LLM-generated reproduction code, reporting a mean SAU score of 0.221 across 360 evaluations and 0.301 for the strongest reported configuration, Claude with PaperCoder.
What changed: EFR makes visual-change evidence explicit using action-location and candidate-region annotations, then filters that evidence before verifying task outcomes.
What changed: STRIVE adds explicit intermediate reasoning, a consistency gate, a validation agent and a progression-aware GRPO reward that distinguishes direction-preserving errors from direction reversals.
What changed: According to the paper, the agents generated results novel relative to prior literature on five construction problems, including new finite-field Kakeya sets, kissing configurations, records on Kakeya needle and sign uncertainty problems, and an improved Erdős minimum-overlap lower bound.
What changed: It reports that request difficulty falls into discrete denoising-step levels, short benchmarks understate serving variance, CPU dispatch accounts for most single-request wall-clock time, and synchronized batching can improve throughput by 16.0x at batch size 16 over a per-request-dispatch baseline.
What changed: The paper reports evidence that instrument choice substantially affects measured model preferences: rankings generalised across instruments with a coefficient of 0.348, and one-instrument preferences carried little information about what a second instrument would report.
What changed: The authors report that reinforcement-learning fine-tuning of Qwen3-32B with RobustTests improved absolute performance by 3% on LiveCodeBench compared with baseline methods.
What changed: It treats context construction as a distinct optimization layer and reports improvements over weaker and stronger hand-crafted baselines, with transfer to additional long-video benchmarks without further search.
What changed: TRACE uses parent-edit-child transitions and observed property deltas as feedback, and the paper reports a macro-average hit-rate increase from 18.13% to 25.96% against the LLEMA baseline in a controlled same-backbone comparison.
What changed: The authors report improvements including up to 1.85x faster curvature updates, 1.25x faster training, 1.38x speedups, and accuracy gains of up to 5.66% over state-of-the-art compression methods.
What changed: On the reported controlled benchmark, the method reduced the candidate space by approximately 480x, increased F1 from 0.08 to 0.66 in ablation, achieved macro F1 of 0.70 across seven enterprise data makes, and reduced estimated cost and expert effort.
What changed: It adds a contract-guided workflow with separate implementation-requirement and reference-evidence channels, and reports the highest mean score among same-backbone scaffolds on PaperBench Code-Dev under Claude-Sonnet-4.5 and Gemini-3-Flash.
What changed: It shifts evaluation from single outputs or aggregated success rates toward direct assessment of finite sets of candidate outputs using automatically verifiable task coverage.
What changed: It broadens visual-grounding evaluation beyond one-shot reference resolution and reports that current LVLMs perform below human task-level baselines, especially when they must proactively ask questions to resolve underspecified targets.
What changed: It introduces a cross-modal forecasting approach that uses visual in-context learning rather than specialised temporal architectures and reports competitive performance across epidemiology, meteorology, and power-system datasets.
What changed: The work treats explicit revisions to target-defining rules as a structured maintenance signal rather than as ordinary statistical concept drift, reporting 92.3% accuracy and 90.2% Macro-F1 while reprocessing 14.7% of historical data and reducing average update latency relative to complete retraining.
What changed: The study links measurable model behaviour to simulated systemic outcomes, reporting that simple adversarial inputs can alter AI recommendations and, under the paper's financial contagion model, increase bank failures and lower the threshold for cascading disruption.
What changed: It reports that attacks can enter through source data and prompts, propagate through agent communication, and affect final trading decisions across five assets, two backbones, and two target directions.
What changed: It extends agent evaluation beyond scripted interactions by measuring persona-dependent failure modes, trajectory-level brittleness, tool and infrastructure attack effects, and consistency across repeated runs.
What changed: It presents a structured method for evaluating explanation reliability and reports that high-performing classifiers can still produce degenerate or uninformative explanations, while overfitting can reduce the discriminative value of fidelity scores.
What changed: It reports that the proposed experimental designs can recover relevant mixture rankings after removing about 25% of the original proxy-training runs, while modelling relational and pairwise domain effects.
What changed: It identifies a handoff tax: full-trajectory escalation from a lower-capability model to a higher-capability model recovers less than half of the quality gap while adding substantial cost, whereas downshifting can offer a more favorable cost-quality trade-off.
What changed: The proposed method replaces log-probability-based confidence signals with an externally computed retrieval-grounding signal and reportedly improves accuracy by up to 5.4% and minority-correct question performance by up to 35% across four benchmarks and five LLMs.
What changed: It proposes a framework for translating between alignment protocols and desired constraints on individual or group welfare, and applies it to voting-by-issues, random dictatorship, and welfare-maximizing protocols.
What changed: It introduces tests of an LLM's implicit world model for stochastic limit order book dynamics and identifies biased estimates and spurious predictability in forecasting future events.
What changed: Compared with outcome-only RL and prior self-distillation baselines, the proposed method reportedly improves task success and reduces the training steps and interaction budgets needed to reach target performance.
What changed: The authors report 72.0% macro accuracy in sonar-only tasks and 68.7% under fusion, exceeding the strongest baselines by 34.4 and 25.1 percentage points respectively.
What changed: The proposed method jointly addresses unsupported claims and user-pressure-induced answer changes while keeping model weights frozen and avoiding intervention on every turn.
What changed: The authors report compression of coding-agent context to 25.7% of its original size while retaining 86.5% of uncompressed single-shot solve quality on 300 SWE-bench Lite instances; the model is described as a 264 MB adapter that can self-host on one 24 GB GPU.
What changed: Compared with the Qwen3.5-4B base model, the authors report a 9.95% improvement in SciFact label prediction accuracy, a 3.7% increase in evidence quality, and a 19.63% reduction in evidence hallucination rate.
What changed: It frames alignment-tax mitigation as a data-selection problem and combines reference-model log-probability margins, chosen-versus-rejected response length differences, and TF-IDF similarity to general-capability corpora into a composite risk score.
What changed: It moves beyond prompting and supervised fine-tuning by combining supervised initialization, GRPO with verifiable rewards, and policy-context perturbation for adaptive policy invocation.
What changed: It applies a unified activity-based model to AI-assisted large vessel occlusion detection in the CT stroke pathway and estimates a break-even threshold of approximately 3,992 patients per year, with positive first-year ROI at around 5,000 annual stroke patients under the stated assumptions.
What changed: The authors report that self-correction raised instruction-following performance to 4.23 versus 3.81 for a same-backbone agentic HTML pipeline, while operating at 1.75 times the speed and approximately 44% lower cost; 66% of cases stopped after one pass and rollback removed observed regressions.
What changed: It provides a systematic taxonomy and evaluation pipeline for runtime anomalies, reports broad vulnerability among evaluated agents, and distinguishes anomalies addressable through adversarial reinforcement learning from deeper reasoning limitations.
What changed: The authors report that marking untrusted spans as non-executable substantially reduces prompt-injection success while preserving utility and readability, including SEP separation rising from 24.3% to 96.5%, TensorTrust attack success falling from 34.8% to 6.6%, and zero compliance across four PIArena attack families.
What changed: The proposed method aims to improve fine-grained temporal alignment and robustness to noisy timestamp annotations while retaining speech transcription performance.
What changed: The protocol adds a second evaluation signal and reports asymmetric disagreement, including an 8.0% Type II pattern overall and substantially higher disagreement for high-scoring answers under heavy occlusion.
What changed: It reports improvements over outcome-only KTO and DPO on HumanEval(+), MBPP(+), BigCodeBench, and LiveCodeBench, and finds that execution-based labels outperform LLM-as-a-judge labels for this optimization task.
What changed: It reports that LLM-based evaluation is useful for scalable assessment but that reliability varies by metric and evaluation configuration, supporting hybrid pipelines that retain human oversight.
What changed: It reframes human-in-the-loop oversight as a design and organisational challenge requiring explicit support for overseer cognition, rather than treating human presence alone as sufficient.
What changed: It proposes 53 orthogonal binary evaluation dimensions across 10,671 samples, using controlled perturbations to isolate individual capability failures and balance positive and negative examples.
What changed: It consolidates scattered examples into a resource arguing that unconventional and potentially unpredictable behaviour is commonplace in modern AI, including in reinforcement-learning systems and internet-scale foundation models.
What changed: Across six judges and two domains, the reported results show that longer review windows increase catch rates and false rejection together, while one- or two-action units produce the highest informedness.
What changed: It reports an evaluation on four backend coding tasks using five frontier coding-CLI models, finding that two-agent AgentRoom runs abandoned fewer tasks and showed less variation than Solo, with one matched-compute comparison favouring AgentRoom over parallel-merge.
What changed: TRACE adds multilingual coverage, nine risk categories, ten attack strategies, evidence annotations, and evaluation of 18 guardrail models, extending safety assessment beyond prompts and final outputs.
What changed: It provides a quantified case-study measurement of autobiographical confabulation, reports replication with current named models, and evaluates a corpus-grounding remedy.
What changed: It reports that KV compression is 1.20x to 2.00x cheaper across the tested memory-relief settings, while tensor parallelism is the only tested lever that improves latency and remains necessary when model weights exceed a single device's memory.
What changed: The work extends language-model reasoning from text and code generation toward intervention-based experimentation, reporting more specific and actionable outputs than language-only reasoning in the described application setting.
What changed: The work presents multimodal large language models as a potential general-purpose alternative to specialist molecular encoders, with embeddings conditioned on both molecular information and natural-language context.
What changed: The benchmark grounds evaluation tasks and scoring rubrics in real experimental protocol modifications rather than tasks elicited solely from experts; Claude Opus 5 achieved the highest reported normalized rubric score of 59.2%.
What changed: On the CodeContests test split with Gemma 4, the authors report a 0.624 pass rate, 14.4 percentage points above direct prompting, while using fewer pipeline stages and lower wall-clock cost than CodeSIM.
What changed: It reports improvements over prior reward agents, including a 3.6-point F1 gain over the baseline average and a 4.23-point task-success gain.
What changed: It reports provable 1.28-to-1.36-fold sample-efficiency gains over rejection sampling and empirical parity with Best-of-N accuracy using substantially fewer generated tokens across several benchmarks.
What changed: The authors report that ExTS is competitive with or improves on task-specific tree-search baselines across several tasks, with an average relative gain of 5.5% using one fixed configuration.
What changed: It replaces static prompts reused across inference instances with item-specific prompts that adapt to context, constraints, evidence, and unresolved fields.
What changed: SAEM uses reasoning-stage detection, stage-aware caching, expert-aligned token repacking, and in-situ CPU execution to reduce memory pressure and data movement during inference.
What changed: The reported evaluation indicates that an optimised Whisper large-v3 pipeline reduced word error rate by 20.7% and 24.4% on recorded discussions, while MedGemma-RAG identified 2.3 times more MDT-concordant interventions than a proprietary cloud comparator without a significant difference in overall accuracy.
What changed: The reported evaluation moves kernel optimization beyond isolated benchmarks toward deployment-aware end-to-end testing, with geometric-mean latency speedups of 3.91x on A100 and 6.98x on H100 across ten language-model inference workloads.
What changed: It consolidates previously fragmented studies on the risks of recursively using AI-generated data to train subsequent models and identifies open research challenges.
What changed: It proposes signed, offline-verifiable decision records containing hashed references to inputs, outputs, and evidence, linked through a SHA-256 chain and supported by a reference implementation and two-language conformance kit.
What changed: The authors report improved EnterpriseRAG-Bench performance over the strongest published baseline at scales from 10 million to 618 million tokens, with gains ranging from 12.23% to 4.66%.
What changed: The authors report accuracy comparable to FP16 and up to 1.4x speedup in W4A8 and W8A8 configurations, with deployment validation on an Orin NX device.
What changed: RO-PnR models user health literacy and belief commitment and reports the highest cost-adjusted utility across three datasets and three base models, with 30% fewer turns than always-probe baselines.
What changed: On PlanBench, the authors report that the method increased average success from 35.5% for LLM+P to 70.8%, achieved 66.3% faithful success, and reduced semantic drift to 6.4%.
What changed: It presents a NotebookLM-enabled approach for rapidly generating quantitative geothermal benchmarks and describes an LLM-assisted case study for auto-parallelizing geothermal numerical models.
What changed: It replaces query-time SQL synthesis with selection among domain-aligned tools containing parameterized queries and server-side business logic, reporting a pooled mean score of 0.939 for the verticalized pack versus 0.666 for raw SQL and 0.605 for the generic pack.
What changed: The authors report that capability feedback increased paired local-information success from 50% to 100%, while multi-step commitments improved held-out-template success by 6.9 points and reduced model decisions by 32%; a short DSL also reduced peer traffic and commitment time relative to free-form communication.
What changed: KVBoost combines dual-hash cache keying, boundary repair, deviation-guided recomputation, KV quantization, adaptive chunking, and importance-weighted eviction; on Qwen2.5-3B it reports a 4.49x reduction in time-to-first-token compared with recomputation.
What changed: It finds that schema validation is the dominant quality filter, ontology grounding provides an additional filter, staging-logic validation rejected no records in the tested fully gated corpus, and retrieval augmentation has strongly model-dependent effects.
What changed: It adapts hack-verifiable environments to Terminal Bench, evaluates reward-hacking rates across frontier models, and releases the environments and agent traces.
What changed: The reported materials-science benchmark results show competitive answer accuracy with substantially fewer retrieved-context tokens, 2.7x lower end-to-end latency than prompt-all, and stronger tool selection, parameter validity, provenance, and licence grounding.
What changed: On EnterpriseRAG-Bench, the system reportedly improves the Overall score from 68.22 to 82.26 and reaches 86.50 Correctness across more than 500,000 documents and approximately 650 million tokens.
What changed: It provides a comparative framework-level analysis and reports that only one reviewed method supports frequency-domain explanations, only two evaluation metrics are time-series-specific, and identical XAI methods can produce substantially different explanations across frameworks.
What changed: It presents an integrated on-system alternative to cloud-GPU and conventional on-premises inference, claiming that sensitive data can remain within the LinuxONE hardware perimeter while achieving sub-two-second end-to-end RAG latency.
What changed: The reported experiment found that just-in-time warnings without cognitive forcing reduced compliance with AI-steered recommendations containing dark patterns from 71.7% to 53.7%, while awareness and flagging generally showed limited improvement.
What changed: The authors report that role assignment, particularly the final integrator, drives much of the observed performance and can produce a high-recall operating point without threshold tuning.
What changed: It provides a methodological framework linking geographic differences between source and target regions to transfer performance, finding that information shift and spatial shift have statistically significant and complementary explanatory power.
What changed: The reported system improves balanced accuracy over matched baselines on temporal, medication-adverse-event, and recorded-order verification, while reducing false-support predictions under evidence masking and recovering source-traceable event chains when intermediate events are hidden.
What changed: It introduces generative, missing-signal, missing-type, and temporal prompts to model missingness-aware intra- and cross-modal relationships in EHR data.
What changed: The approach replaces fixed memory compression with programmatic retrieval and transformation of persistent session state, while retaining lossless historical records and reporting strong results on LongMemEval_S, BEAM_10M, and LOCA_256K.
What changed: Encoder adaptation increased regional accuracy from approximately 39.03% for zero-shot CLIP to 75.94–82.10%, while full fine-tuning reduced mean distance to the predicted region centre from 12.30 km to 3.86 km.
What changed: It reports that the attack achieved an average adversary-aligned response rate of 91.2% across four downstream tasks, including 86.6% on GPT-5.5, while a proposed memory-boundary defence reduced the rate to 80.6%.
What changed: It reports that state-of-the-art vision-language models achieve only 24.1% average pair accuracy and presents Verdict Log-Odds Supervision as a post-training method that substantially improves open-weight backbones.
What changed: The authors report that directly optimised hybrid architectures, particularly MobileViTv2 models, outperformed heavier models and knowledge-distillation approaches for edge deployment, achieving up to 87.4% sensitivity, 86.5% specificity, and 97.2% negative predictive value against specialist labels.
What changed: It reports that priors improve learning when they effectively reach the policy, that HL-Gauss critic training adds 14.4 in-domain points over MSE, and that one commonly used in-domain slice negatively correlates with cross-domain transfer while the hardest slice predicts transfer strongly.
What changed: On the reported trade-finance and Tobacco-3482 evaluations, HIRA substantially improved Macro-F1 while reducing human corrections and LLM calls without updating model weights.
What changed: Compared with outcome-level contrastive RLVR methods, LACL-GUI adds preferences within successful and failed trajectories to favour concise successful executions and distinguish failures by their divergence from successful trajectories.
What changed: The authors report a formal method intended to distinguish responsibility among agents and human stakeholders, detect blame-shifting attempts and remain invariant under forged evidence.
What changed: It provides an expert-calibrated evaluator, LitJudge, that reportedly raises agreement with human experts from Spearman's rho 0.467 for existing LLM-as-a-judge methods to 0.78, while finding that leading systems win only 23.0% of decisive matches against human drafts.
What changed: It provides a large-scale benchmark and supporting historical box-office data for evaluating multimodal models on film narrative, reporting near-chance performance for many vision-language models and a maximum accuracy of 61.1% for audio-visual models.
What changed: It reports validation across five benchmark cases spanning four PDE classes, claiming lower long-time extrapolation error than the numerical prior and better performance than the strongest competing baseline in each case.
What changed: The authors report consistent improvements across large language models on the STATE-ToxiCN and TBO benchmarks, including state-of-the-art results for bilingual multi-tuple extraction.
What changed: It reports that distilled students can match teacher performance on adversarial text, reduce false alarms on harmless prompts, and achieve approximately 24 ms CPU classification latency for the encoder class.
What changed: It combines documentation-grounded code generation with executable feedback from ABB RobotStudio, exposing failures that static and semantic checks may miss.
What changed: It compares accuracy alongside runtime, energy use, output length, extraction failures, and statistical significance, finding that no model dominates and that Gemma3:4b produces more correct answers per watt-hour despite Qwen3:4b leading accuracy on two datasets.
What changed: On SWE-bench Verified, the authors report that routing arbitration increased the macro-average resolved rate from 44.9% to 48.2% for the gpt-oss family and improved over uniform selection on Qwen3.6.
What changed: The reported system improves Gemini 2.5 Flash-Lite material-selection accuracy from 37.5% with pure LLM reasoning to 90.0%, and achieves 75.0% printability across 96 physical validation trials with 88.9% task suitability among successfully printed samples.
What changed: It proposes linking persistent agent identities, cryptographically verified capabilities, execution performance, creator identities, and reputation within a single infrastructure.
What changed: It extends authorship obfuscation from independently protecting individual documents to addressing correlations across multiple texts, including cross-genre attribution and verification attacks.
What changed: It reports evaluations across 27 model configurations, identifies six recurring failure modes, and presents prompting and supervised fine-tuning methods that improve performance by up to 20 percent.
What changed: It defines a benchmark requiring models to extract requested textual attributes together with corresponding image bounding boxes, and reports that Q2IT improves on standalone vision-language models while still leaving a substantial performance gap.
What changed: The benchmark shifts evaluation toward underspecified real scientific requests with attachments and no ground-truth solutions, finding that no model cleared the acceptance threshold under all three judges and that 47.6% of scored judgments fell below the threshold.
What changed: It adds reusable environment- and rollout-orchestration layers that provision isolated tool environments and overlap long tool-call trajectories to improve GPU utilization, with integrations for veRL and slime.
What changed: It reports that most evaluated models show some activation-control ability and that this control can imperfectly evade several activation-based monitoring techniques in simple tasks.
What changed: The proposed SSKG approach produced a wider accuracy range of 44.1% to 85.2% and a monotone mastery gradient, compared with 96.8% to 100% accuracy from the tested prompt-based simulations.
What changed: It adds explicit multi-level orchestration, bounded evidence routing, specialist item scoring, profile auditing, targeted revision, and a persistent hierarchical evidence trace.
What changed: ATHENA combines a weight-sharing supernet with cross-hospital architecture priors and multi-agent LLM search, and reportedly matches or outperforms four NAS baselines in 9 of 12 hospital-task evaluations at a search budget of 30.
What changed: The authors report that a design generated by the framework passed China Classification Society Approval in Principle and reduced steel mass and unit capital cost by 8.1% versus the human-optimized TuQiang baseline.
What changed: Across 500 environments in six domains, the reported feasible-task rate increased from 48.6% to 94.8%, while trained PPO policies showed stronger performance and improved transfer to WebArena, WebShop, and MiniWoB++ without LLM calls during evaluation.
What changed: Across 57 repositories and 105 release transitions, the study finds that all evaluated transitions invalidate part of the prior skill set, while six frontier agents achieve only 29.9%–69.7% average-at-3 macro F1 on stale-content removal.
What changed: On the reported experiments, CompPO improved held-out accuracy over tuned GRPO on Qwen3-4B and improved greedy pass@1 on Qwen3-4B and Llama-3.1-8B-Instruct; the authors also report greater stability across PPO-grid runs.
What changed: The reported results indicate that residual representation can match dense retrieval quality with 30% fewer context tokens and lower scaling of per-query scoring work, while top-down progressive descent performs poorly for document selection.
What changed: The proposed evaluation reports per-level accuracy, magnitude-weighted stability, and family-specific collapse points instead of relying on aggregate accuracy, revealing recurring weaknesses involving conflicting instructions and impossible premises.
What changed: Verifier output is reused during memory retrieval, conflict resolution, summarization, and archival rather than being applied only as a one-time admission filter.
What changed: It adds a culturally grounded Arabic safety evaluation protocol and reports that most evaluated models failed to defend against roughly half of the unsafe prompts, particularly in weapons and illicit-substance categories.
What changed: The authors report that the approach can synthesize selection heuristics for LU factorization that replicate expert-designed strategies.
What changed: It reports that harness choices can change individual model scores from bands rather than points, account for 95.7% of adjacent-model score gaps on average, and select different leaderboard winners.
What changed: Evaluation is expanded from final game artifacts or isolated tasks to three lifecycle stages assessed through interaction, behavioral tests, product criteria, and regression checks.
What changed: The paper organizes the field around the representational and workflow requirements that determine when native volumetric modeling is needed, and proposes a Claim-Design-Validation framework for assessing technical, workflow, and clinical claims.
What changed: The method eliminates per-draft-position recurrent-state snapshots, storing a smaller pseudo-value matrix and reconstructing only the accepted state during commit.
What changed: On a 10-target held-out split, five sampled GPT-4o policies achieved 0.589 Recall@10 versus 0.571 for the strongest single-feature Protenix binder ipTM baseline; target-conditioned GPT-5.4 policies reached 0.519 Recall@10 and 0.583 NDCG@10 on a three-target subset.
What changed: It reports that scheduled context variation matches or outperforms static-context baselines in out-of-distribution evaluations across CartPole, BipedalWalker and VehicleRacing, with additional in-distribution gains on the more complex environments.
What changed: The proposed ensembles reportedly improve modified F1 by 4.3% over a Multimodal Transformer and accuracy by 3.0% over the Interpretable Multimodal Routing baseline, while producing feature-importance measures with higher human-annotator agreement on IEMOCAP and CMU-MOSI.
What changed: It reports a capability-auditability trade-off and identifies limitations in reinforcement-learning signals, probability-based metrics, probes, and reasoning ablations used for oversight.
What changed: In an immersed-boundary CFD benchmark, the authors report mean paired reductions of 31.1% in relative-L2 field error and 24.7% in normalized pressure-derived load error across twelve backbone architectures, three regimes, and two lead ranges.
What changed: It replaces direct averaging or heuristic aggregation of confidence estimates with soft evidence over candidate answers and a null state, and reports improved calibration on GSM8K, SVAMP, and GSM-Hard using Qwen, Mistral, and Gemma models.
What changed: It applies difference-graph discovery methods to linear discrete-time dynamic structural causal models to identify variables whose causal coefficients change across regimes.
What changed: Across more than 78,000 measurements, the tested models reportedly performed no better than chance at distinguishing real interventions from sham conditions, although fine-tuning and probes could recover the intervention signal from internal activations.
What changed: The authors report 98.2% attribute-value accuracy at 74.7% coverage in an offline human evaluation, a 90.4% increase in impression-weighted enrichment coverage in production, and a 0.48% checkout-conversion increase in an online experiment.
What changed: In experiments using Qwen2.5-14B-Instruct Q4_K_M on Apple-silicon unified memory, the authors report approximately 80% main-context token savings, a 1.66x faster first-argument token at 250 tools, and 1.1-1.7x TTFT speedups at moderate context depth.
What changed: It extends certified robustness beyond single-turn inputs using compositional bounds, safety-persistence analysis, information-theoretic upper bounds, and a unified algorithm; experiments on six unspecified LLMs reportedly show empirical safety above the certified bounds.
What changed: It provides standardized metrics for within-section cellular representation advantage and cross-section, dataset, and organ transferability, and applies them across 30 pathology-specific and general-purpose foundation models in 304,920 evaluation runs.
What changed: It adds a reproducible, executable-ground-truth evaluation framework for open-ended model outputs and reports Grok 4.6 with the highest point estimate at 65.1, although only 101 of 351 pairwise model comparisons were statistically resolved.
What changed: It proposes shifting engineering discipline and governance upstream into precise specifications, explicit quality gates, quantitative measures, and auditable provenance as coding agents handle larger portions of implementation.
What changed: It expands long-context evaluation beyond document-level, high-resource-language testing by measuring word-, sentence-, paragraph-, and document-level comprehension across document positions.