scientific artificial intelligenceagent evaluationpublished Aug 25, 2026 · arXiv:2608.21601
K-Bench: measuring model performance on real scientific agent requests
tool useretrieval-augmented systemspublished Aug 18, 2026 · arXiv:2608.14992
Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5