arXiv cs.AI - 2026-08-05 ​
416 items collected.
1. ISEE: Interactive Semantic Enrichment for Database Fields ​
Author: Yuan Tian, Yiru Chen, Rakesh R. Menon, Zifan Liu, Ting Cai, Fei Wu, Anudeep Chimakurthi, Prashanthi Ramamurthy, Sridevi Aishwariya Ganesan, Kun Qian, Yunyao Li
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.IR
arXiv:2608.02604v1 Announce Type: new Abstract: LLM-based agents are increasingly being deployed for data-related tasks, including data sense-making, exploration, and retrieval. However, their performance heavily depends on the clarity and completeness of data semantics. In practice, many field desc...
2. Self-Organising Digital Circuits ​
Author: Marcello Barylli, Gabriel B'ena, Alexander Mordvintsev, Eleni Nisioti, Sebastian Risi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2608.02606v1 Announce Type: new Abstract: Fault tolerance in classical computing has traditionally relied on static strategies like hardware redundancy and error-correcting codes. Biological systems, in contrast, exhibit adaptive plasticity, maintaining function through dynamic re-organisation...
3. Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling ​
Author: Tairan Fu, Javier Conde, Carlos Arriaga, Gonzalo Mart'inez, Pedro Reviriego, Javier Coronado-Bl'azquez
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.02618v1 Announce Type: new Abstract: Recent studies have identified an ``Artificial Hivemind'' effect in Large Language Models (LLMs) causing models to converge on a narrow, homogenized consensus even for open questions. This semantic collapse limits the diversity of AI, resulting in high...
4. PULSE: An Executable Contract Language for Spatiotemporal Knowledge Graph Engineering ​
Author: Dongxu Yang, Ziyi Liang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.DB, cs.PL
arXiv:2608.02630v1 Announce Type: new Abstract: Knowledge graph engineering often distributes accepted state, observations, constraints, processes, and hypothetical scenarios across artifacts whose combined execution contract remains external. We present PULSE, an Object-Process-Methodology-inspired...
5. HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents ​
Author: Zian Zhai, Xingyu Tan, Gaowang Zou, Xiaoyang Wang, Wenjie Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.02650v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external tools to complete complex real-world tasks. However, reliable tool-use planning remains challenging due to the limitations of implicit reasoning and the evolving nature of real-world execu...
6. Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap ​
Author: Benjamin Fresz, Elena Dubovitskaya, Marco F. Huber
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG
arXiv:2608.02699v1 Announce Type: new Abstract: When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right to Explanation. Yet whether (and how) Explainable AI (XAI) can satisfy this right in practice remains...
7. Predictive Set Theory: A Generative Framework for Cognitive Architecture with Operationalized Core Mechanisms ​
Author: Yiyang Yu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LO, q-bio.NC
arXiv:2608.02704v1 Announce Type: new Abstract: Predictive processing theories portray the brain as a hierarchical prediction engine that minimizes prediction error, yet they lack operational definitions for the structure of a "prediction," the standardized response to a prediction error, and the me...
8. Towards a new paradigm of scientific discovery with socialized artificial intelligence ​
Author: Xinjie Yao, Xingxin Xu, Xiyuan Gao, Zhoupeng Guo, Kunlong Yang, Dengyu Zhao, Siqi Zhao, Zhihe Fan, Yichen Dong, Xin Li, Jiekang Feng, Jiahe Wu, Sen Wang, Beiming Yu, Kejia Zhao, Ruipu Zhao, Jiaqi Zhou, Heyang Li, Jianjun Chen, Anbo Dai, Xin Liu, Zhengtao Yu, Qinghua Hu, Pengfei Zhu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02775v1 Announce Type: new Abstract: Scientific discovery has advanced through successive transformations in the organization of knowledge. Observation and experimentation established the empirical foundations of science. Theory made it possible to derive general principles from particula...
9. BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL ​
Author: Chong Peng, Pin Qian, Su Wang, Yihang Chen, Varun Sah
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02876v1 Announce Type: new Abstract: Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context and database work before useful evidence appears, while post-hoc compression cannot recover omitted...
10. VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space ​
Author: Yu-Tung Liu, Cunxi Yu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02878v1 Announce Type: new Abstract: Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% accuracy on standard benchmarks. We trace this ceiling to an incomplete debugging action space: existing systems res...
11. Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes ​
Author: Maryam Rezaee, Pooriya Safaei, Maryam Asgarinezhad, Fatemeh Seyyedsalehi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02879v1 Announce Type: new Abstract: The widespread adoption of proprietary Large Language Models (LLMs) accessed strictly through closed APIs has created a critical challenge for responsible deployment: a fundamental lack of interpretability. To address this, we propose a model-agnostic,...
12. Hypercubes, Hyperplanes, and Constraint-Induced Complexity Collapse in Atomic Concept Learning ​
Author: Irene Tsapara
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LO
arXiv:2608.02930v1 Announce Type: new Abstract: We revisit higher-arity atomic concept learning through the geometry of hypercubes and hyperplanes of ground instances. Our starting point is the observation that the ambient r-dimensional hypercube of ground atoms is not structurally uniform. Its logi...
13. When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning ​
Author: Andrew Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02940v1 Announce Type: new Abstract: A reproducible compression statistic can still select the wrong candidate. A dense pruning score with 0.906 split-half reliability predicted a 16.1% gain. Its selected endpoint was 6.0% and 7.7% worse than two controls. We model the gap through informa...
14. On the missing data layer and a potential solution ​
Author: Francis F Daniel, Mauro Iba~nez, Francis Perelman, Marian Basti
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.DB
arXiv:2608.02949v1 Announce Type: new Abstract: Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the dataset layer. The dataset layer faces two compounding problems: discovery and supply. Latin American AI datasets e...
15. Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Reinforcement Learning ​
Author: Subrat Prasad Panda, Blaise Genest, Arvind Easwaran
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.02993v1 Announce Type: new Abstract: (Flat) Reinforcement Learning (RL) agents face significant challenges in environments with sparse rewards that require long-horizon reasoning. A compelling approach to improve sample efficiency is to incorporate knowledge into learning and decision-mak...
16. On the missing benchmarks layer and a potential solution ​
Author: Francis F Daniel, Mauro Iba~nez, Francis Perelman, Marian Basti
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02996v1 Announce Type: new Abstract: Latin America is missing a foundational layer for native AI development: the benchmark layer. The benchmark layer does two things no other layer can - it audits AI systems against regional social requirements and it directs AI optimization in economica...
17. ProPRL: Property-Aware Prerequisite Relation Learning in Educational Knowledge Graphs ​
Author: Xinghe Cheng, Jiapu Wang, Chaobo He, Ruihai Dong, Quanlong Guan
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03006v1 Announce Type: new Abstract: Prerequisite relation learning is central to adaptive instruction, yet existing methods often formulate it as conventional link prediction, limiting their ability to adaptively integrate complementary educational evidence for individual candidate pairs...
18. UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks ​
Author: Jiayu Cao, Xingyuan Zeng, feiyu Li, Zhijing Huang, Xujie Yuan, Rongxiang Chen, Shimin Di, Libin Zheng, Jian Yin
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragmented and have little interoperability, placing a heavy operational burden on users. Existing digital p...
19. LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment ​
Author: Linhan Xia, Rui Liu, Zhaofeng Zhang, Yihao Wang, Binrui Shen, Shengxin Zhu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03020v1 Announce Type: new Abstract: Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation through the frozen backbone. Every adaptation step therefore needs backward-capable hardware and must store or recomput...
20. DiffImaginE: Imagine to Verify Entity Types with Diffusio ​
Author: Feng Zhang, Feiyu Han, Rongxin Yang, Yang Liu, Yancheng Chen, Rui Wang, Yingguang Yang, Tian Xueyun, Chongyang Zhang, Hao Zheng, Xu Kefu, Congjing Ran, Fuhai Chen, Bin Chong
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03025v1 Announce Type: new Abstract: Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visual fe...
21. Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning ​
Author: Zhitian Hou, Yuhang Liu, Pengkai Wang, Zeyu Liu, Guanghao Zhu, Zheng Liu, Shuo Cai, Congkai Xie, Zhijie Sang, Kun Zeng, Hongxia Yang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03028v1 Announce Type: new Abstract: Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations largely use isolated and fixed scenarios. A model may therefore answer correctly by recalling a dru...
22. CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting ​
Author: Xiaoyu Tao, Mingyue Cheng, Bokai Pan, Chuang Jiang, Huanjian Zhang, Tian Gao, Yaguo Liu, Qi Liu, Enhong Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03031v1 Announce Type: new Abstract: Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features. Recent advances in large language models (LLMs) have extend...
23. TraceCAD: Trace-Guided Repair for Agentic CAD Generation ​
Author: Fengxiao Fan, Jingzhe Ni, Fan Sang, Xiaolong Yin, Yu Liu, Ruofeng Tong, Min Tang, Peng Du
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.GR, cs.SE
arXiv:2608.03062v1 Announce Type: new Abstract: LLM-based CAD agents produce executable parametric programs, but their correction loops may lose evidence about satisfied requirements, faulty operations, and prior repairs. We introduce TraceCAD, a recovery layer that links requested features, modelin...
24. Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls ​
Author: Guoyao Yu, Xiaoqing Sun, Ziqi Huang, Shaojing Fan, Zhongyi Zhang, Xiaomeng Hu, Xiaobo Xue, Yangyang Shi, Xiong Xiao, Yang Song, Biao Lyu, Rong Wen, Xing Li, Qinming He, Shunming Zhu, Zhenguang Liu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03071v1 Announce Type: new Abstract: Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool and orchestrating the order of calls. However, correctly filling the parameters of a tool call is e...
25. AI Agent Economics: Can Autonomous Economic Behavior Emerge among AI Agents under Minimal External Conditions? ​
Author: Lingyun Zhang, Shang Shang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03076v1 Announce Type: new Abstract: Multi-agent studies commonly place AI agents in predefined games, markets, or roles, making it difficult to distinguish endogenous economic organization from behavior inherited from the scenario. We ask whether economic relations emerge when agents rec...
26. Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR ​
Author: Yongshi Ye, Liang Zhang, Yidong Chen, Xiaodong Shi, Biao Fu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03119v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR replace gold supervision with answer-level consensus from model samples. ...
27. Beyond Average Performance: Dynamic Instance Clustering and Specialized Algorithm Design in LLM-Assisted Evolutionary Search ​
Author: Qinglong Hu, Qingfu Zhang, Fei Liu, Xialiang Tong, Kun Mao, Mingxuan Yuan
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03129v1 Announce Type: new Abstract: Large Language Model-assisted Evolutionary Search (LES) has emerged as a powerful paradigm for automated algorithm design. However, existing LES methods primarily optimize for average performance, inherently directing search effort toward instances tha...
28. Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents ​
Author: Xiaolong Sun, Qichao Wang, Hangyu Li, Liang Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03137v1 Announce Type: new Abstract: Large language model (LLM) agents must retain reusable information, control a bounded active context, and recover earlier evidence during long-horizon interaction. Existing methods commonly optimize long-term memory (LTM) and short-term memory (STM) se...
29. Spatial proteomics guided by H&E-based AI reveals recurrence-risk niches in triple-negative breast cancer ​
Author: Yesung Cho, Ji Hwan Park, Chanil Kim, Hyewon Kim, Honglan Li, Yumin Lee, Geongyu Lee, Sujeong Hong, Seong Min Park, Yoonyoung Lee, Hee Sool Rho, Sumin Lee, Amos Chungwon Lee, Changhwan Lee, Hwanyoung Shim, Hyunwook Kim, Hyeji Shin, Sanha Park, Jihoon Yu, Yoon Hee Shin, Sooheon Kim, Hyunjin Park, Seung Min Park, Sangwan Kim, Yujung Kim, Sung-Im Do, Eun-Young Kim, Dongmyung Shin, Jongbae Park, In-Gu Do
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, q-bio.QM
arXiv:2608.03145v1 Announce Type: new Abstract: Deep learning models can predict cancer recurrence from H&E stained slides, but the localized molecular states underlying these predictions remain largely obscured. Here, we developed an outcome informed spatial pathology framework in TNBC that integra...
30. UniGD: A Unified Generative-Discriminative Framework for Industrial Retrieval ​
Author: Shujie Ji, Yawei Kong, Yilin Zhao, Li Wang, Xialong Liu, Peng Jiang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03150v1 Announce Type: new Abstract: Generative retrieval (GR) is a promising paradigm for industrial search advertising, yet its deployment is constrained by strict relevance and latency requirements. Existing systems cascade GR with an independent relevance model, decoupling the generat...
31. Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning ​
Author: Sahil Al Farib, Momota Ahsana Meem, Sheikh Redwanul Islam, Md. Tanvir Raihan
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.03161v1 Announce Type: new Abstract: Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval does not fully preserve. This paper presents an evidence-grounded multimodal pipeline that transcribes lectures,...
32. Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation ​
Author: Saqib Shouqi, Abdullah Nazly, Januki Wanniarachchi, Ravisha De Alwis
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03166v1 Announce Type: new Abstract: Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and education, where maintaining consistent personas, ethical constraints, and behavioral coherence under advers...
33. Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study ​
Author: Qiming Bao, Sherry J. H. Feng, Kim Chester Eugenio, Meng Fon
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.03172v1 Announce Type: new Abstract: Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes "Maria S.", not [NAME] -- so that clinical text stays fluent and downstream tools keep working. But this only he...
34. Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Domain QA ​
Author: Jiwon Lee, Yong-chan Park, Jungin Hong, U Kang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03177v1 Announce Type: new Abstract: How can question answering (QA) systems determine whether a query is ambiguous? Ambiguity detection is essential in open-domain QA, as misclassification leads to answering the wrong interpretation or unnecessary clarification. However, existing methods...
35. TumorBoard: Evidence-Grounded Multi-Agent Decision Support for Longitudinal Neuro-Oncology ​
Author: Yantong Liu, Zheyu Zhang, Runpeng Liu, Mu Xitang, Seong-Yoon Shin, Hyun-Ae Lee
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03190v1 Announce Type: new Abstract: Neuro-oncology decisions require coordinated interpretation of serial MRI, pathology, molecular markers, treatment history, performance status, and evolving guidelines. We present TumorBoard, a multi-agent decision-support system built around a shared ...
36. When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models ​
Author: Yu Feng, Chunting Zang, Chen Shen, Rui Miao, Ge Teng, Weidong Cai, Jieping Ye
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03201v1 Announce Type: new Abstract: Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit two widely used safety-guard training datasets, WildGuardMix and GR-Train, and find that among respo...
37. The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agentic Systems ​
Author: Ankur Sharma, Deep Shah
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03214v1 Announce Type: new Abstract: Large language models have transformed artificial intelligence from isolated prediction services into components of long-running, distributed systems that reason, invoke tools, retrieve external state, delegate tasks, and act on behalf of users and org...
38. Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains ​
Author: Yanchao Li, Wanhao Liu, Jiaqing Xie, Ben Gao, Yanbo Wang, Tianfan Fu, Yuqiang Li
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.03219v1 Announce Type: new Abstract: Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach new answers, or produce answers that were already within reach. Aggregate scores do not distin...
39. UniNav: A Unified World-Action Diffusion Model for Visual Navigation ​
Author: Changqing Zhou, Yueru Luo, Zeyu Jiang, Changhao Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03244v1 Announce Type: new Abstract: Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, while navigation world models can anticipate future observations but often r...
40. One Knob to Rule Them All: A Unified Optimal Transport View of Cold-Start Active Learning ​
Author: Ning Zhu, Xiaochuan Ma, Juntao Xu, Jingze Liang, Mengfei Zhao, An Chen, Liang-Jian Deng
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03249v1 Announce Type: new Abstract: Cold-Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance. Existing methods take diverse routes based on typicality, coverage, or diversity. Each rests on its own inductive...
41. TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning ​
Author: Wonpyo Park, Seung-won Hwang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03276v1 Announce Type: new Abstract: Long-context inference with large language models is constrained by the linear growth of the key-value cache to sequence length. While pruning offers mitigation, prevailing methods determine query-specific token importance that cannot be reused across ...
42. AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions ​
Author: Zhiyao Cui, Qianyi Wang, Haoyang Yan, Yiqun Zhang, Siyue Ren, Hangfan Zhang, Zelin Tan, Hao Li, Chunjiang Mu, Dexian Cai, Shao Zhang, Chen Zhang, Meng Li, Jianan Chai, Yuting Fan, Zichao Ye, Xiaolei Yang, Xinyao Lu, Yuyang Yu, Wenjie Lou, Xiaosong Wang, Fenghua Ling, Shiyang Feng, Mao Su, Qiaosheng Zhang, Bo Zhang, Yang Chen, Lei Bai, Shuyue Hu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03283v1 Announce Type: new Abstract: Identifying promising scientific ideas remains an important challenge in research practice. Researchers commonly rely on small-group discussions or one-to-one interactions with a single large language model, yet these approaches often expose them to on...
43. DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning ​
Author: Le Xiang, Zhicheng Guan, Hong Chen, Xiaocong Lin, Zhenghua Lei, Teng Hu, Bolei He, Long Zeng
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03292v1 Announce Type: new Abstract: Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distributed across multiple pages. Existing approaches, including end-to-end ML...
44. Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks ​
Author: Mohsen Arjmandi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.03297v1 Announce Type: new Abstract: A standard claim in the literature on retrieval-augmented and memory-augmented language models is that shorter context is better when the relevant information is preserved. We test this claim by running every sample of two long-context benchmarks -- BA...
45. SeaSlides: Semantic Abstraction Layer for Agentic Slide Generation ​
Author: Shengjun Fang, Chenyang Wu, Zongzhang Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03298v1 Announce Type: new Abstract: Agentic presentation generation must preserve source content, maintain coherent visual design, render specialized objects, and produce usable artifacts. Existing systems meet only part of this requirement: templates preserve regularity but restrict ada...
46. Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents ​
Author: Siqi Fan, Minghao Li, Xiaoqian Ma, Wenhui Tan, Xiusheng Huang, Juntong Wu, Liujie Zhang, Shuo Shang, Weihang Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03327v1 Announce Type: new Abstract: Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way the effect goes. Under one identical GUI-MCP harness on the OSWorld-MCP benchmark (309 tasks), the same MCP tools ...
47. Long-term Traffic Scene Prediction via Polynomial Representations in Autonomous Driving ​
Author: Yue Yao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03330v1 Announce Type: new Abstract: This thesis addresses fundamental challenges in traffic scene prediction for autonomous driving by introducing robust and computationally efficient models based on polynomial representations. While conventional sequence-based representations often stru...
48. Traceable Multi-Agent System for Knowledge-Based Forecasting ​
Author: Junhyeok Kang, Sangjun Han, Hyeokjun Choe, Soonyoung Lee
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03339v1 Announce Type: new Abstract: Enterprise forecasting increasingly relies on autonomous agents that interpret documents, search for data, generate code, and revise models. While this autonomy helps build adaptive forecasting pipelines, it also makes it difficult for practitioners to...
49. MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc ​
Author: Mingtian Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03397v1 Announce Type: new Abstract: MMLongBench-Doc is a long-document QA benchmark of 1,082 questions over 135 PDFs. Two properties of it push measured scores away from the quantity they are meant to capture: the reference metric compares extracted answers, so 1,358,000 loses to 1358000...
50. Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance ​
Author: Can Wang, Haoran Chen, Li Yu, Ding Hao, Bohai Zhao, Zhaoyang Liu, Zhiying Tu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03403v1 Announce Type: new Abstract: The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central role as the primary interface through which agents interact with external environments, yet existin...
51. Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems ​
Author: Zuojun Max Shen, Yuan Qu, Pujun Zhang, Anbang Liu, Yunhao Liang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.ET
arXiv:2608.03413v1 Announce Type: new Abstract: As artificial intelligence (AI) continues to evolve and mature, recent AI practices have moved beyond large language models (LLMs) and text or image generation tasks, increasingly integrating tools, agents, and harnesses to solve real business and indu...
52. AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction ​
Author: Jonaid Shianifar, Iias Faiud
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules. This paper reports the c...
53. Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory ​
Author: Jakub Rada (AI Center, Department of Computer Science, Faculty of Electrical Engineering, Czech Technical University in Prague), Viliam Lis'y (AI Center, Department of Computer Science, Faculty of Electrical Engineering, Czech Technical University in Prague)
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03420v1 Announce Type: new Abstract: Large language models have improved substantially on single-shot reasoning tasks, but their performance in sequential decision-making is less well understood. We study this on fully-observable two-player zero-sum games, which provide ground-truth evalu...
54. State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking ​
Author: Xiaohe Li, Yang Lu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03425v1 Announce Type: new Abstract: Transformer-based architectures have dominated sequence modeling, largely due to the expressive power of attention mechanisms. However, for a class of deterministic state tracking tasks---such as parity checking, modular counting, and parenthesis match...
55. DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces ​
Author: Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03451v1 Announce Type: new Abstract: Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured querying, retrieva...
56. LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models ​
Author: Fengqi Zhu, Shaoxuan Xu, Jingyang Ou, Zebin You, Yipeng Xing, Huabin Liu, Xiaolu Zhang, Jun Zhou, Zhenzhong Lan, Yankai Lin, Wayne Xin Zhao, Jianguo Li, Chongxuan Li, Ji-Rong Wen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03457v1 Announce Type: new Abstract: Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, comp...
57. Solver-Aware Decompositions for Programming-by-Example: When Dividing Requires Knowing how to Conquer ​
Author: Janis Zenkner, Tobias Sesterhenn, Tim Grams, Christian Bartelt
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03461v1 Announce Type: new Abstract: Decomposition-based Programming-by-example (PBE) scales performance by splitting tasks into subtasks that a learned synthesizer solves: a decomposer predicts intermediate subgoals, and a synthesizer generates programs conditioned on them. Current appro...
58. LeanMem: Simple and Efficient Long-Term Memory for LLM Agents ​
Author: Yuxin Liao, Le Wu, Min Hou, Hao Liu, Han Wu, Zishu Wang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03463v1 Announce Type: new Abstract: Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history. However, existing memory systems typically process heterogeneous dialogue content through a uniform summarization and retrieval pipeline, ...
59. ChartAnno: Evaluating MLLMs for Chart Annotation Generation ​
Author: Zhenghan Chen, Zekai Shao, Lidan Tan, Xin Lin, Xingchen Zeng, Yi Shan, Ziyue Lin, Xiaoliang Fu, Xinyuan Liu, Yuetong Guo, Fen Wang, Bongshin Lee, Siming Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC
arXiv:2608.03464v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored. Annotating charts is a common yet challenging communicative ta...
60. When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO ​
Author: Zhe Cao, Miaowen Wen, Fangjiong Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.03467v2 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. In GRPO, this completion-level uniformity creates structure-level skew: recurring correct solution forms accumulate po...
61. ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning ​
Author: Xiuhui You, Jiayi Luo, Zichao Shen, Qingyun Sun, Ziwei Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03468v1 Announce Type: new Abstract: Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage. Existing approaches directly construct tool-level graphs from these trajectories, but the resulting graphs remain tied...
62. WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks ​
Author: Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang, Xiyang Hu, Yue Zhao, Shuli Jiang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03499v1 Announce Type: new Abstract: Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents th...
63. Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design ​
Author: Zejun Liu, Jian Wu, Ru Peng, Yuliang Ji, Dongyuan Li, Renhe Jiang, Yue Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03501v1 Announce Type: new Abstract: AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the impor...
64. Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks ​
Author: Christophe D. Hounwanou, John Emeka Eze, Ya'e Ulrich Gaba
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2608.03502v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents. However, LLM-based agents struggle with long-horizon sequential decision tasks that require precise acti...
65. When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs ​
Author: Omatharv Bharat Vaidya, Connor Thomas Jerzak, Zayne Rea Sprague, Fangcong Yin, Nhat Ho
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, stat.ML
arXiv:2608.03506v1 Announce Type: new Abstract: Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an ...
66. Reversing Arrows in Large Language Models ​
Author: Sefika Efeoglu, Adrian Paschke
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03512v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance on text-to-knowledge graph generation and related tasks. Nevertheless, it is still unclear whether they accurately model the direction-dependent semantics of inverse relations, in which reve...
67. Dr. AGENTONOMICS: A Didactic Experiment of AGENTONOMICS ​
Author: Fengjunjie Pan, Alois Knoll
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03524v1 Announce Type: new Abstract: AGENTONOMICS is a framework that treats AI agents as economic entities that can be designed, managed, and governed through an integrated management architecture. Dr. AGENTONOMICS is its first application: a lecture agent developed in the context of the...
68. Behaviorally Adaptive Visual Diversion for Inclusive and Resilient Digital Assessment Delivery ​
Author: Gupta Lovi Raj, kaur Kamalpreet, Dama Sriram, Parali Prajithaa
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.03531v1 Announce Type: new Abstract: Institutions increasingly rely on browser lockdown, webcam monitoring, and behavioral analytics to secure high-stakes digital assessments, yet these mechanisms are commonly designed and evaluated independently and often overlook learner accessibility. ...
69. Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve ​
Author: Denys Pushkin, Albert Q. Jiang, Aryo Lotfi, Colin Sandon, Emmanuel Abb'e
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03550v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities. Originally, this technique was introduced to elicit step-by-step reasoning from large language models (LLMs), which would otherwise tend to direc...
70. Enhancing Tabular Learners with Context-Aware Semantic Embeddings ​
Author: G"unther Schindler, Maximilian Schambach, Johannes H"ohne
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.03565v1 Announce Type: new Abstract: While modern tabular learners excel at capturing statistical patterns, they frequently operate in a semantic vacuum, treating textual features as discrete symbols, ignoring the rich semantics inherent in feature names or cell entries. We propose CASE (...
71. Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities ​
Author: William Bolton, Philip Torr
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG, cs.MA
arXiv:2608.03569v1 Announce Type: new Abstract: Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in evaluating scientific reasoning and research replication, but often rely on synthetic tasks or retrospec...
72. Policy Fragmentation or Institutional Alignment? Institutional Governance of AI in Universities and Business Schools ​
Author: Lydia Manikonda, Dominique Outlaw
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.ET
arXiv:2608.03584v1 Announce Type: new Abstract: Artificial intelligence (AI) is rapidly transforming high-skilled domains, requiring higher education institutions (HEI) to balance the teaching of foundational principles with the integration of emerging tools to ensure workforce readiness. While HEI ...
73. From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities ​
Author: Mengying Zhou, Yongjie Yin, Yang Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.03585v1 Announce Type: new Abstract: Open-source software communities are a form of digital public infrastructure that not only produces code, but also generates public knowledge and interpersonal relationships through visible collaboration. Generative coding agents (CAs) are an advanced ...
74. FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection ​
Author: Amirhossein Taleshinosrati, Yangyang Wang, Atitaya Phoemsuk, Vahid Abolghasemi, Naser Hossein Motlagh, Sadasivan Puthusserypady, Daniel Teichmann, Abdolrahman Peimankar
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.03597v1 Announce Type: new Abstract: Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and is associated with increased risks of stroke, heart failure, and mortality. Recent ECG foundation models offer transferable representations for automated AF detection. However...
75. Large language models for partial differential equation workflows ​
Author: Han Wan, Rui Zhang, Hao Sun
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03600v1 Announce Type: new Abstract: Partial differential equations (PDEs) become actionable in science and engineering not as isolated formulae, but as executable workflows that connect modelling assumptions, governing equations, numerical solvers, diagnostics, and decisions. Large langu...
76. FraQ: Efficient Coordinate-Space Recompression for Federated Low-Rank Adaptation ​
Author: Shenghui Li, Thiemo Voigt
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03605v1 Announce Type: new Abstract: Federated fine-tuning with Low-Rank Adaptation (LoRA) enables efficient collaborative adaptation of Large Language Models (LLMs) without centralizing private data. However, LoRA's two-factor parameterization creates an aggregation mismatch across clien...
77. Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents ​
Author: William Bolton, Philip Torr
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.03606v1 Announce Type: new Abstract: Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence. We study this setting by framing oncology clinical development as an offline decision-making problem...
78. Formal Verification of Agentic Systems over Operational Data ​
Author: Alejandro J. Mercado, Alessio Lomuscio
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03609v1 Announce Type: new Abstract: Agentic systems driven by large language models (LLMs) are increasingly deployed in real-world workflows where they act on persistent operational data. Before deployment, these systems need to be verified against business requirements that govern workf...
79. Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations ​
Author: Chunlei Meng, Jacqueline J. Pang, Pengbin Feng, Zhenyu Yu, Chun Ouyang, Zhongxue Gan
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.MM
arXiv:2608.03611v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are often incomplete. Existing methods for incomplete-observation MSA mainly follow two paradigms. Reconstruction-based...
80. Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection ​
Author: Razieh Chalehchaleh, Reza Farahbakhsh, Noel Crespi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03627v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for automated fact-checking, yet their susceptibility to gender bias in this context remains underexplored. This study presents the first systematic investigation of gender bias in LLM-based fake news ...
81. Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model ​
Author: Abdallah Khemais
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.03629v1 Announce Type: new Abstract: A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream. For the one composition in that model where two carriers are ...
82. When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation ​
Author: Yinuo Jiang, Yongjie Ye, Zhou Tao, Xiang Zhuang, Qiang Zhang, Huajun Chen, Tiankai Li
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03632v1 Announce Type: new Abstract: On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, o...
83. Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details ​
Author: Maksymilian Wolski, Nicholas Hoernle, Johannes Forkel, Jakob Foerster
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.03644v1 Announce Type: new Abstract: AI agents deployed in real-world settings must be capable of coordinating with humans and other AI agents they have not encountered before. Zero-shot coordination (ZSC) algorithms aim to achieve this by specifying high-level learning rules such that in...
84. AutoSND: From Execution Evidence to Structural Policies for Automated Network Dismantling Heuristic Discovery ​
Author: Zhijing Hu, Changjun Fan, Yufan Deng, Zhiguang Cao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03653v1 Announce Type: new Abstract: Network dismantling is fundamental to analyzing the robustness and vulnerability of complex systems, yet practical heuristics must balance effectiveness and computational efficiency, and are usually designed manually by researchers. Existing large lang...
85. Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training ​
Author: Yibei Liu, Jiajun Chen, Qianle Zhang, Tangyue Jin, Mengying Zhu, Meng Xi, Yangyang Wu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03660v1 Announce Type: new Abstract: Reinforcement fine-tuning (RFT) is widely believed to inherently resist catastrophic forgetting in continual post-training of multimodal large language models. Under pronounced task distributional shifts, however, forgetting across representative RFT a...
86. Shielding for Higher-Order Safety ​
Author: Filip Cano, Thomas A. Henzinger, Konstantin Kueffner
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03662v1 Announce Type: new Abstract: Safety shields are runtime enforcement mechanisms that restrict the actions of a controller to guarantee safety. Classical shields are usually synthesised for state predicates: the current physical state is either safe or unsafe, and the shield disable...
87. PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud ​
Author: Chenghua Wang, Daliang Xu, Dongqi Cai, Duojin Sun, Hao Zhang, Haoze Qian, Huaiyuan Zhang, Jinshuo Cui, Junbo Cui, Kezhao Zhao, Longxi Gao, Mengwei Xu, Rongjie Yi, Tam Sikyuen, Tianyue Zhang, Weikai Xie, Xuanzhe Liu, Yingying Qin, Yiwen Lu, Yuan Yao, Yuezhi Zu, Yunhan Guo, Yuxin Zheng, Ziqi Guo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.03682v2 Announce Type: new Abstract: Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they ...
88. LiveEvalBench: Toward Open-World Evaluation for Web Generation ​
Author: Yiyao Wang, Zhen Wen, Yinghao Tang, Yixiao Fu, Lin Yuan, Xiaolau Zhang, Jun Zhou, Wei Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.03689v1 Announce Type: new Abstract: Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem. We argue that frontend artifacts demand a different paradigm: they are inter...
89. TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management in Long-Term Agents ​
Author: Han Xiao, Hongjun Xu, Xin Zhang, Yidong Chen, Xiaodong Shi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03699v1 Announce Type: new Abstract: Persistent memory helps long-term agents retain knowledge, yet a single update error can repeatedly distort future retrieval and reasoning. Most existing systems reduce memory updating to a binary Write/Hold decision, which cannot distinguish whether n...
90. Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchanges ​
Author: Jonaid Shianifar, Blaz Mramor, Fangda Zou, Matthieu C. Martin, Xingsheng Guo, Zhihua Zhu, Rong Zhou, Bichen Shi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.03705v1 Announce Type: new Abstract: Real-time bidding (RTB) ad exchanges typically forward nearly all incoming requests to demand-side platforms (DSPs), even though only a small fraction receive bids. This over-distribution weakens auction outcomes: DSPs throttle participation under comp...
91. When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Diagnostic for Machine Collectives ​
Author: Molood Arman
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.03722v2 Announce Type: new Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking argument...
92. SAT-Edge-Agent: Hardware-in-the-Loop Edge-Agent Orchestration for Onboard Satellite Intelligence ​
Author: Longji He, Jeto Xu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03728v1 Announce Type: new Abstract: Onboard satellite intelligence requires a task layer that translates mission intent into local tool calls, exposes execution state, and returns machine-consumable artifacts under communication and power constraints. We present SAT-Edge-Agent, a hardwar...
93. CARE-Bench: Benchmarking Patient-Facing LLM Triage ​
Author: Yining Hua, Hongbin Na, Cyrus Ayubcha
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03731v1 Announce Type: new Abstract: Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the user should take next. We introduce CARE-Bench, a source-grounded benchmark that evaluates sequentia...
94. Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement ​
Author: Chunyang Jiang, Pingping Zhang, Yuzhi Zhao, Wenao Ma, Zhijian Hou, Mengyang Wu, Yiyang Cai, Senkang Hu, Sitong Cheng, Chi-Min Chan, Wei Xue, Yike Guo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03733v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate. Self-augmentation offers a promisin...
95. AgenticECO: An Agentic Framework for ECO on 3D Integrated Circuits ​
Author: Shuo Ren, Yaohui Han, Libo Shen, Zhiqiang Jia, Rongliang Fu, Bei Yu, Tsung-Yi Ho
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03738v1 Announce Type: new Abstract: As Moore's law slows, the industry is turning to three-dimensional integration; yet in merged 3D-IC flows, routed designs expose bond-level defects with no 2D analogue, and post-route engineering change orders (ECO) remain manual, expertise-bound work....
96. MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models ​
Author: Yu Ran, Wentao Zhao, Xin Zhang, Yi Pan
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03740v1 Announce Type: new Abstract: Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks. The security implications of this coordinate generation process have been largely overloo...
97. Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems ​
Author: Sebasti'an Andr'es Cajas Ord'o~nez, Agastya Munnangi, Aldo Marzullo, Felipe Ocampo Osorio, Quang Bui, Mohammad Shahin, Armaan Grewal, Emmanuel Paul Kwesiga, Anqi Peter Li, Josephine Nanyonjo, Aaditya Panchal, Arshnoor Bhutani, Nikhil Jaiswal, Milit S. Patel, Maximin Lange, Leo Anthony Celi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03744v1 Announce Type: new Abstract: Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on ...
98. Risky Business: Measuring The Faithfulness-Safety Tension ​
Author: Dominik Meier, Luca Joshua Francis, Marco Bernhard Kaiser, Terry Ruas, Jan Philip Wahle, Bela Gipp
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.03745v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output strictly derives from its reasoning trace. We identify an alignment tension where a model must be faith...
99. GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks ​
Author: Leijun Zhou, Zhihao Liu, Xiang Qu, Chenxu Liu, Yifei Liu, Yanke Yu, Jingzhe Xu, Xuejun Wu, Buyue Qian, Xi Chen, Yaowei Zheng, Junhao Hu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03764v1 Announce Type: new Abstract: Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valuable task dom...
100. Computing Actual Causes for Neural Network Predictions under Structured Causal Inputs ​
Author: Jannick Strobel, Muqsit Azeem, Stefan Leue
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.LO
arXiv:2608.03772v1 Announce Type: new Abstract: Explaining the predictions of neural networks is a central challenge in trustworthy AI. Existing explanation methods, such as those based on feature attribution or minimal sufficient sets, typically treat input features as independent, which can yield ...
101. KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation ​
Author: Ruihan Li, Jiyang Tan, Kailin Jiang, Huining Li, Hengyang Lu, Yu Huang, Qian Li, Yuntao Du
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03782v1 Announce Type: new Abstract: Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks mainly focus on entity, attribute, and relation hallucinations, knowledge-related failures are often investigated ...
102. Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation ​
Author: Chunlin Liu, Junnian Chen, Haitong Jiang, Jianyu Zhao, Yingsen Pang, Jingchen Li, Jiabiao He, Youming Lu, Jinhe Bi, Yuntao Du
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03791v1 Announce Type: new Abstract: Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pretraining corpora. Removing such knowledge is essential for building trustworthy AI systems. However, existing stud...
103. LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards ​
Author: Zhinan Liu, Jie Li, Mingyu Kang, Jiayi Ji
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03838v1 Announce Type: new Abstract: Reasoning-based guard models improve LLM safeguards, but decoding explicit rationales for every interaction makes them costly to deploy. Although latent-reasoning methods reduce token generation by moving reasoning into continuous states, they remain u...
104. Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes ​
Author: Tao Jin, Phuong Minh Nguyen, Zhenzhu Yan, Teeradaj Racharak, Naoya Inoue
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03839v1 Announce Type: new Abstract: Training-free speculative decoding drafts by matching an exact suffix of the context against a pool of earlier context. That lookup misses correct drafts already in the pool, most visibly on tool-calling traffic, where a request repeats almost everythi...
105. MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents ​
Author: Jiaming Chen, Yisen Gao, Yanping Li, Zifan Liu, Yumeng Zhang, Jun Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03844v1 Announce Type: new Abstract: Memory-augmented LLM agents rely on rich context for long-horizon reasoning and acting, yet their memory modules expose a persistent attack surface for malicious records, making the study of memory poisoning threats imperative. However, existing query-...
106. ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories ​
Author: Yash Misra, Javal Vyas, Siddharth Gutta, Mehmet Mercang"oz
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY
arXiv:2608.03866v1 Announce Type: new Abstract: This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action. The framework implements a versioned, safety-governed evaluation contract that checks whether a recommendation is ...
107. ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? ​
Author: Tianyi Guan, Yiding Wang, Haotong Yang, Siyuan Cao, Shirui Liu, Yi Hu, Jiaqi Li, Muhan Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.03874v1 Announce Type: new Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabi...
108. Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition ​
Author: Michal Mr'az, Justin Shenk
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03892v1 Announce Type: new Abstract: We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-forced temporal-choi...
109. When Efficiency Becomes Fragility: Exploiting Dynamic Routing Vulnerabilities in Adaptive UAV Tracking ​
Author: Shaofeng Liang, Runwei Guan, Wenshuo Chen, Jiemin Wu, Bowen Tian, Haozhe Jia, Kaishen Yuan, Songning Lai, Daizong Liu, Yutao Yue
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03902v1 Announce Type: new Abstract: Resource constraints on UAV platforms have driven a paradigm shift in aerial tracking, from pursuing performance toward balancing accuracy with efficiency. Adaptive Transformer Trackers, which leverage an input-dependent dynamic routing architecture, h...
110. Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory ​
Author: Matt Ratto, Abhishek Moturu, Daniel Silver
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2608.03910v1 Announce Type: new Abstract: As AI systems are deployed across increasingly diverse social contexts, alignment can no longer be framed as the optimization of a single, unified set of values. Instead, systems must be able to recognize, represent, and respond to multiple legitimate ...
111. Implementing Causal Perception: Competing SCMs and Situated Fairness ​
Author: Jose M. 'Alvarez
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03917v1 Announce Type: new Abstract: Causal perception occurs when agents with competing Structural Causal Models (SCMs) of the same system infer different probability distributions, including the hypothetical distributions implied by each agent's SCM under the same set of interventions. ...
112. The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections ​
Author: Marco Giunti, Fabrizia Giulia Garavaglia
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2608.03921v1 Announce Type: new Abstract: This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models merely reproduce statistical regularities learned in training, we argue that Transformers construct and apply pr...
113. TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring ​
Author: Dongjie Yang, Siyan Lin, Leixian Shen, Rui Sheng, Huamin Qu, Zixin Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03952v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-language (ESL) learners. Effective ESL tutoring, however, requires more than fluent response generation: a tutor must select an appropriate pe...
114. A game theory for foundation models shows new paths to rational cooperation through similarity inference ​
Author: Alexander Meulemans, Maciej Wo{\l}czyk, Marissa A. Weis, Rajai Nasser, Roberta Rocca, Seijin Kobayashi, Guillaume Lajoie, Angelika Steger, Blake Richards, Marcus Hutter, James Manyika, Rif A. Saurous, Jo~ao Sacramento, Blaise Ag"uera y Arcas
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03958v1 Announce Type: new Abstract: As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, th...
115. Interpretable Adaptive Sampling for LLM Test-Time Scaling ​
Author: Mobina Kashaniyan, Ali Jannesari
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03961v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query budgets that spend the same compute on easy and difficult prompts. These fixed budgets are also difficult to inspe...
116. Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations ​
Author: Zizhao Hu, Nathan Elijah Segura, Mohammad Rostami, Jesse Thomason
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03970v1 Announce Type: new Abstract: Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional transcription and restructuring from AI-backed dictation tools. How do th...
117. ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning ​
Author: Jinhe Bi, Chennan Zhou, Zengjie Jin, Aniri, Shuo Lu, Wenke Huang, Hu Cao, Xun Xiao, Zhihong Zhu, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03972v1 Announce Type: new Abstract: On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder pr...
118. Multi-Camera Trajectory Forecasting with Trajectory Tensors ​
Author: Olly Styles, Tanaya Guha, Victor Sanchez
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2108.04694v2 Announce Type: cross Abstract: We introduce the problem of multi-camera trajectory forecasting (MCTF), which involves predicting the trajectory of a moving object across a network of cameras. While multi-camera setups are widespread for applications such as surveillance and traffi...
119. CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification ​
Author: Yiming Ma, Victor Sanchez, Tanaya Guha
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2403.09281v3 Announce Type: cross Abstract: We propose CLIP-EBC, the first fully CLIP-based model for accurate crowd density estimation. While the CLIP model has demonstrated remarkable success in addressing recognition tasks such as zero-shot image classification, its potential for counting h...
120. UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture Search ​
Author: Xiaobin Rong, Leyan Yang, Dahan Wang, Yuxiang Hu, Changbao Zhu, Kai Chen, Jing Lu
Published: 8/5/2026, 4:00:00 AM
Categories: eess.AS, cs.AI
arXiv:2503.00340v2 Announce Type: cross Abstract: Lightweight models are essential for real-time speech enhancement applications. In recent years, there has been a growing trend toward developing increasingly compact models for speech enhancement. In this paper, we propose an Ultra-Lightweight U-net...
121. Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions ​
Author: Wolfgang Mack, Nezih Topaloglu, Laura Lechler, Ivana Bali'c, Alexandra Craciun, Mansur Yesilbursa, Kamil Wojcicki
Published: 8/5/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD, eess.SP
arXiv:2509.24457v1 Announce Type: cross Abstract: Objective speech-quality metrics are widely used to assess codec performance. However, for neural codecs, it is often unclear which metrics provide reliable quality estimates. To address this, we evaluated 45 objective metrics by correlating their sc...
122. PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement ​
Author: Xiaobin Rong, Qinwen Hu, Mansur Yesilbursa, Kamil Wojcicki, Jing Lu
Published: 8/5/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD
arXiv:2511.13300v1 Announce Type: cross Abstract: Generative models have shown remarkable performance in speech enhancement (SE), achieving superior perceptual quality over traditional discriminative approaches. However, existing generative SE approaches often overlook the risk of hallucination unde...
123. StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement ​
Author: Xiaobin Rong, Jun Gao, Zheng Wang, Mansur Yesilbursa, Kamil Wojcicki, Jing Lu
Published: 8/5/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD
arXiv:2603.09234v2 Announce Type: cross Abstract: Achieving high perceptual quality without hallucination remains a challenge in generative speech enhancement (SE). A representative approach, PASE, is robust to hallucination but has limited perceptual quality under adverse conditions. We propose Stu...
124. GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement ​
Author: Xiaobin Rong, Yushi Wang, Zheng Wang, Jing Lu
Published: 8/5/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD
arXiv:2604.01832v1 Announce Type: cross Abstract: We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system integrates a generative branch, which performs full-stack speech restoration in a self-supervised representation ...
125. UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations ​
Author: Xiaobin Rong, Zheng Wang, Yushi Wang, Jun Gao, Jing Lu
Published: 8/5/2026, 4:00:00 AM
Categories: eess.AS, cs.AI
arXiv:2604.14606v2 Announce Type: cross Abstract: Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates. We propose UniPASE, an extension of the low-hallucination PASE framework tailored for USE. At its core is DeWavLM-Omni, a unifi...
126. KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization ​
Author: Shuai Che, Gang Peng
Published: 8/5/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG, cs.SE
arXiv:2608.02611v1 Announce Type: cross Abstract: Automating GPU kernel optimization remains difficult in practice: generated variants can violate correctness constraints, runtime measurements are noisy, and search often stalls early. We present a practical optimization agent that combines LLM-guide...
127. MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale ​
Author: Jiadong Zhang, Xiaosong Ma
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.MA
arXiv:2608.02613v1 Announce Type: cross Abstract: Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and c...
128. OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning ​
Author: Ahnaf Munir, Dannong Wang, Michael W. McDonald, Mubarak Shah, Pegah Khosravi, Yu Tian
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.02615v1 Announce Type: cross Abstract: Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata. However, most medical large language model (LLM) and vision-language model (VLM) benchmarks focus on isolated...
129. Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks ​
Author: Rohith Uppala
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.02616v1 Announce Type: cross Abstract: We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synthetic benchmarks spanning 22 languages and 5 domains. Zero-shot, OPF achieves F1=0.855 on AI4Privacy ...
130. Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety ​
Author: Fay Elhassan, David Sasu, Alexandra Kulinkina, Lars Henning Klein, Mary-Anne Hartley
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.02617v1 Announce Type: cross Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collec...
131. Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks ​
Author: Hsien-Jyh Liao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.02621v1 Announce Type: cross Abstract: Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy for authority grounding. Under ordinary reasoning prompts that did not request statutory citations, fo...
132. Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models ​
Author: Brian K Chen, Chong Wu, Kenji Kawaguchi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.02625v1 Announce Type: cross Abstract: Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by producing text block by block. We study a simple plug-and-play inference pattern: first generate a co...
133. Deep Divide-and-Reduce in Symbolic Regression ​
Author: Yusong Deng, Yanjie Li, Weijun Li
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02628v1 Announce Type: cross Abstract: Symbolic regression (SR) is the task of discovering underlying patterns from data and representing them using mathematical expressions. Current machine learning approaches to SR often lack a profound understanding of the intrinsic mathematical and ph...
134. Multimodal Auto-regressive Transformer Surrogate for Modeling Variable Operations and Quantifying Uncertainty in Geological Carbon Storage ​
Author: Yifu Han, Louis J. Durlofsky
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02629v1 Announce Type: cross Abstract: The use of variable well perforation and injection strategies can improve the efficiency of geological carbon storage operations. We develop a new multimodal auto-regressive transformer surrogate to model these operations under geological uncertainty...
135. Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds ​
Author: Yuxuan Liu, Zhaochen Su, Yuhao Zhang, Jiahe Guo, Zhongwei Xie, Huihao Jing, Lingyun Xie, Qing Zong, Yauwai Yim, Zhixiong Zhang, Haoran Li, Yangqiu Song
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.02636v1 Announce Type: cross Abstract: Self-evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the underlying model. Yet it remains unclear when further evolution helps, how successful and failed trajectories shape...
136. Studying, Identifying, and Fixing Hidden Technical Debt in AI-Intensive Cyber-Physical Systems ​
Author: Beena
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.02638v1 Announce Type: cross Abstract: Artificial Intelligence (AI) components are increasingly pervasive in several software systems, including Cyber-Physical Systems (CPSs). AI-CPS are used in several domains, including autonomous vehicles, industry, home automation, robotics, and healt...
137. Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation ​
Author: Atul Anand, Sourav Chattaraj
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.02639v1 Announce Type: cross Abstract: Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed tone at the same time. We study how instruction-following degrades as such constraints accumulate. We introduc...
138. IR2Solve: Structured Intermediate Representations for Cost-Efficient Optimization Autoformulation ​
Author: Penglin Zhu, Linhai Zhang, Jungang Xu, Xinchi Wei, Xiuqi Wu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.02641v1 Announce Type: cross Abstract: Large language models (LLMs) can translate natural-language optimization problems into solver-ready formulations, but direct code generation is brittle: schema, indexing, and semantic errors can cause compilation failures, infeasible models, or incor...
139. MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows ​
Author: Nithishwer Mouroug Anand, Wei-Tse Hsu, Kyle Vaccaro, Eden James Gage, Jonathan David Colburn, Linda Xi Phan, Minjoon Seo, Kevin Guan, Philip C. Biggin
Published: 8/5/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.AI
arXiv:2608.02642v1 Announce Type: cross Abstract: Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out as a particularly promising target within this broader effort. Coding agents promise to automate significant po...
140. CUADebug: Diagnosing and Repairing Computer-Use Agent Failures ​
Author: Weijia Zhang, Kunlun Zhu, Zeyi Liu, Yinting Chen, Tianyi Ma, Jiateng Liu, Jiaxun Zhang, Bingxuan Li, Xiangru Tang, Heng Ji, Jiaxuan You
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.02643v1 Announce Type: cross Abstract: Computer-use agents (CUAs) operate real desktop and web interfaces through screenshots, mouse and keyboard actions, and stateful UI feedback, yet their failures remain difficult to diagnose and repair. Unlike text-only agents, CUA failures arise from...
141. Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures ​
Author: Isham Kalappurackal Mansoor, Abhishek Phadke, Pratip Rana
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.02645v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on external tools to perform multistage tasks. Existing agent frameworks typically assume that tool calls are atomic and return binary success or failure signals. However, real-world systems exhibit non-atomic b...
142. Cross-Anesthetic ECoG State Decoding Fails at the Decision Threshold, Not the Representation ​
Author: Kunkun Zhang, Qianwei Zhou
Published: 8/5/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI
arXiv:2608.02646v1 Announce Type: cross Abstract: Decoders of anesthetic state from cortical activity fail across drug classes, most notoriously ketamine, but reported accuracy cannot say whether the neural representation or only the decision threshold has failed; we separate the two in a controlled...
143. Secure AI Watermarking Framework for IP Protection in Multi-Tenant Cloud Platforms ​
Author: M Anjan Kumar, Kishor Kumar Gajula, Ch Prathima
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02656v1 Announce Type: cross Abstract: The Secured data safe guard transaction with multi-tenant environments run on private-protected authenticate platforms runs by secured handed environments that emerges with the expansion of cloud-based AI services. To enhanced this secured leakage ad...
144. Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure ​
Author: Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Xu Peng, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02657v1 Announce Type: cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address the threats, little is known about the internals of agentic LLMs when they a...
145. AI Alignment and Fiduciary Obligation ​
Author: Benjamin Lange
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2608.02660v1 Announce Type: cross Abstract: Advanced AI assistants engage users in extended interactions across a widening range of roles, including advice, decision support, collaboration, learning, emotional support, and companionship among others. Current alignment efforts consider what ali...
146. Verifier-Guided Model Discovery for Physical Dynamical Systems with Pretrained Symbolic Transformers ​
Author: Farbod Faraji, Francesco Belardinelli
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02662v1 Announce Type: cross Abstract: Reliable forecasting of nonlinear physical systems underpins scientific discovery and engineering decision-making. Yet high-fidelity simulations are prohibitively costly, and machine-learning surrogates can be opaque and encode assumptions about syst...
147. CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study ​
Author: Mohammad Nasir Uddin, Rahnuma Tabassum Orpita, Asaduzzaman Anik, Eklachur Rahman Bhuiyan, Marjahan Risalat, SM Wali Ullah, Asif Ahamed
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02663v1 Announce Type: cross Abstract: Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types. Existing sequence models handle irregular sampling but ignore typed relational structure; existing graph models assume fixed-interv...
148. ZK-SR117: A Chunked Zero-Knowledge Attestation Design for Aggregated Fair-Lending Metrics, with a Control Mapping toward Full SR 11-7 Coverage ​
Author: Mohammad Nasir Uddin, Rahnuma Tabassum Orpita, Eklachur Rahman Bhuiyan, Asaduzzaman Anik
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02664v1 Announce Type: cross Abstract: Deploying ML models in regulated decision-making (credit underwriting, fraud detection, loan approval) requires demonstrating fairness and robustness to auditors without exposing model weights or customer data. We address this attestation problem for...
149. Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity ​
Author: Yongxi Zhou, Junwei Yao, Yuanzhe Liu, Zihan Dong, Wenbo Ye, Jiaxi Wen, Lai Yun Choi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving surface form varies, do...
150. Sphere Retraction Normalizations ​
Author: Jie Zhang, Cheng-Fang Su, Yi-Jui Huang, Min-Te Sun
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.02668v1 Announce Type: cross Abstract: Residual connections are the de facto mechanism for training deep neural networks stably. Geodesic Normalization (GeoNorm) recasts them on a Riemannian manifold, orthogonalizing each layer output against the current hidden state and applying the resu...
151. Vulnerabilities, Secrets and Misconfiguration in the Highest-Exposure Docker Hub Images ​
Author: Cristhian Kapelinski, Beatriz Machado, Diego Kreutz
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CY, cs.OS, cs.SE
arXiv:2608.02669v1 Announce Type: cross Abstract: Docker Hub is the registry underneath most container deployments, and a flaw in a widely reused base image is inherited by every image built on it. Prior ecosystem-scale measurements each rely on a single detector, leaving the tool-dependence of thei...
152. Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments ​
Author: Dotan Davidovich, Yair Amar, Hai Rozencwajg, Or Hiltch
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02670v1 Announce Type: cross Abstract: Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems, non-root execution) constrain them like any other software. Existing benchmarks, however, evaluate agents almos...
153. Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC Generation ​
Author: Francis Luis Santos Vargas, Rodrigo Brand~ao Mansilha, Diego Kreutz
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.ET, cs.SE
arXiv:2608.02672v1 Announce Type: cross Abstract: Cloud misconfiguration remains a leading cause of security incidents, yet whether LLMs and SLMs can generate security-compliant Infrastructure-as-Code is an open question. We benchmark seven models, three closed LLMs (Claude Opus 4, GPT-5.4, Gemini 2...
154. dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model ​
Author: Hankun Wang, Bohan Li, Shi Lian, Xiaoyu Gu, Jing Peng, Da Zheng, Colin Zhang, Kai Yu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL, eess.AS
arXiv:2608.02673v1 Announce Type: cross Abstract: Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface for expressing edit requests, but its ambiguity may leave the intended o...
155. Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks ​
Author: Shangze Li, Chuancheng Shi, Simiao Xie, Lingzhi He, Cheng Ji, Zifeng Cheng, Fei Shen, Chao Wu, Tat-Seng Chua
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02674v1 Announce Type: cross Abstract: With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward white-box attacks that directly identify and disrupt internal safety neurons or routes. However, exist...
156. When Policies Change Probabilities: Modular Decision-Making for LLM Code Review ​
Author: Rasvik Kudum, Max Corbett, Hitansh Paliwal, Romaisa Fatima, Thomas Jiralerspong, Sneheel Sarangi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.MA
arXiv:2608.02677v1 Announce Type: cross Abstract: LLM code reviewers often estimate patch risk and make approval decisions in one prompt. A probability should depend on evidence; costs should determine the action taken from it. We test whether four deployed reviewer interfaces preserve this separati...
157. DenialRAG: Single-Document RAG Poisoning via Embedded Parametric Denial ​
Author: Abay Zhurekbay, Tao Liu, Fan Li
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.02678v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) systems are vulnerable to corpus poisoning: an attacker who inserts a crafted document into the retrieval corpus can steer the underlying large language model (LLM) toward an attacker-chosen wrong answer. Prior si...
158. AI Sandbox: Technical Report ​
Author: Muhammad Waseem, Md Aidul Islam, Md Nasir Uddin Shuvo, Md Mahade Hasan, Kai-Kristian Kemell, Jussi Rasku, Mika Saari, Vilma Saari, Roope Pajasmaa, Markku Oivo, Pekka Abrahamsson
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.02679v1 Announce Type: cross Abstract: Collaborative AI experimentation across industry and academia requires platforms that enable rapid prototyping while preserving controlled access, tenant separation, and transparent workflows. Despite growing interest in AI sandboxes, there is still ...
159. TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows ​
Author: Salma El Yadouni (EPFL), Guanyi Li (Binome Technologies)
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2608.02680v1 Announce Type: cross Abstract: Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure with retries, exploration, accidental ordering, and repeated lookups. We present TraceCompiler, a skill-guided ...
160. $S^3$: Improving Agent Safety through Multi-Stage Defense ​
Author: Zibo Xiao, Haoyu Wang, Jun Sun
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02683v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish complex tasks. However, risks may emerge at different stages, propagate across steps, and become difficul...
161. A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models ​
Author: Shu Quan, Tianfang Hao, Sitong Fang, He Geng, Jiayi Zhou, Boyuan Chen, Kaile Wang, Donghai Hong, Juntao Dai, Yaodong Yang, Jiaming Ji
Published: 8/5/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI
arXiv:2608.02684v2 Announce Type: cross Abstract: Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in protein engineering can equally be prompted to generate predicted toxin-like sequences, potentially...
162. BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests ​
Author: Zetong Xiong, Qiao Zhao, Jun Zhang, Xueying Lyu, Zhi Li, Yixiang Tu, Xiaowen Yang, Yunjie Zhang, Yufeng Wang, Zhe Zhang, Kaize Yu, Hanwen Du, Zhongkai Sun, Zhuoxin Liu, Zekun Lin, Jianwen Yang, Ruining Chen, Ying Zhang, Tingxuan Pan, Ke Chen, Shubin Han, Chuanhao Sun, Yehua Yang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.02685v1 Announce Type: cross Abstract: Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixed change sequence. Sequential policies can process a pull-request (PR) queue one candidate at a time...
163. Learning Molecular Representations from Cellular Phenotypes with Structure Preservation ​
Author: Xuan Lin, Jingyu Sheng, Tengfei Ma, Li Sun, Dapeng Xiong
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02688v1 Announce Type: cross Abstract: Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses. However, existing multimodal representation learning methods often optimize cross-modal alignment without considering the...
164. Output-Aware Rotation for INT2 KV-Cache Quantization ​
Author: Vincent-Daniel Yun, Woosang Lim, Minsoo Cheong, Sunwoo Lee, Murali Annavaram, Sai Praneeth Karimireddy, Sungjoo Yoo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02691v1 Announce Type: cross Abstract: The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important. However, existing rotation-based INT2 methods optimize cache statist...
165. Measuring Explainer Stability via Attribution Separability ​
Author: Eddie Conti, 'Alvaro Parafita, Axel Brando
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02697v1 Announce Type: cross Abstract: Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scores due to stochastic components in their definition. In this paper, we...
166. Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, Cross-Principal Approach ​
Author: Mohamed Chahine Ghanem
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02698v1 Announce Type: cross Abstract: Tool-using agents built on large language models (LLMs) are increasingly deployed not by a single operator but by many, side by side on shared infrastructure. This creates a population-level risk that single-agent safeguards miss: a handful of agents...
167. NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory ​
Author: Yizhe Chen, Wenshuai Yao, Saiya Wang, Yuannuo Feng, Wenbo Qi, Kechao Tang, Ngai Wong, Wenyong Zhou, Wang Kang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02700v1 Announce Type: cross Abstract: Analog compute-in-memory (CIM) enables energy-efficient neural network inference, but device variation and read noise can severely degrade low-bit quantized models. Existing CIM-oriented quantization methods mainly minimize ideal quantization error, ...
168. Can Training Logs Make Model Comparisons More Precise? ​
Author: Wei-Jung Huang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2608.02705v1 Announce Type: cross Abstract: Comparing stochastically trained models requires estimating both a performance difference and its uncertainty from repeated runs. We study whether training logs from those same runs can make such comparisons more precise. Because training-log covaria...
169. Designing a Good Virtual Node: Addressable and Cardinality-Preserving Global Memory for Message Passing Architectures ​
Author: F'elix Marcoccia
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02709v1 Announce Type: cross Abstract: Virtual nodes give message-passing neural networks a simple global communication route, but the standard node--VN--node pipeline compresses the graph into one homogeneous state and broadcasts it identically to every node. Building on the Two-Radius a...
170. Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators ​
Author: Yansong Sun, Shenxiu Wu, Siyuan Chen, Runlin Hou, Junhao Qiu, Junming Cao, Shudi Shao, Zhichao Lu, Qingfu Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.02712v1 Announce Type: cross Abstract: Kernel generation for hardware accelerators such as GPUs and NPUs has become a proving ground for large language models (LLMs), and state-of-the-art systems raise correctness through pipelines that couple LLMs with agentic reinforcement learning and ...
171. Quo Vadis, World Modeling? ​
Author: Yu Yang, Xuemeng Yang, Licheng Wen, Lingdong Kong, Xiaobin Hu, Dongyue Lu, Wei Chow, Xiyan Huang, Yuxiang Feng, Yue Liao, Jianbiao Mei, Daocheng Fu, Rong Wu, Pinlong Cai, Ran Yi, Ying Tai, Jiangning Zhang, Botian Shi, Yong Liu, Shuicheng Yan
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO
arXiv:2608.02713v1 Announce Type: cross Abstract: Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agen...
172. Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents ​
Author: Shuai Wang, Haodong Chen, Yu Yin, Shengyao Zhuang, Bevan Koopman, Guido Zuccon
Published: 8/5/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2608.02751v2 Announce Type: cross Abstract: Existing deep-research agents use a Search--Visit workflow that retrieves whole webpages without considering the structure they expose through titles, headings, sections, and metadata. This prevents agents from directly constraining retrieval to part...
173. Privacy-Preserving AI Verification via Minimal Information Disclosure ​
Author: Sleem Abdelghafar, Gabriel Kulp
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02774v1 Announce Type: cross Abstract: AI verification crosses a trust boundary: a verifier must learn enough to establish an authorized claim, yet the same evidence can reveal sensitive details about the model, workload, or hardware. We introduce minimal information disclosure (MID), whi...
174. A Hyperfinite Framework for Score-Based Generative Modeling ​
Author: Sunder Ram Krishnan
Published: 8/5/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, math.PR
arXiv:2608.02799v1 Announce Type: cross Abstract: Score-based diffusion models are typically formulated using continuous-time stochastic differential equations and measure-theoretic stochastic calculus. In this paper, we develop a hyperfinite formulation of score-based generative modeling within the...
175. SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology ​
Author: Abdallah Lamane, Abdul Rahman Diab, Ren-Chin Wu, William Lotter
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02803v1 Announce Type: cross Abstract: Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet its attention maps provide only local explanations: they indicate where a model focuses but not which histologic...
176. A Unified 2D Framework for DeepLesion Detection, Segmentation and Short Report Generation ​
Author: Ruida Cheng, Tejas S. Mathai, Benjamin Hou, Qingqing Zhu, Zhiyong Lu, Matthew McAuliffe, Ronald M. Summers
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02805v1 Announce Type: cross Abstract: In previous work, we integrated large language models (LLMs) into the lesion segmentation model based on the ULS23 DeepLesion dataset, using short-form findings from the reports. In this study, we developed a unified 2D lesion analysis framework that...
177. Learning a Vector-Symbolic Model for Socio-Cultural Tasks ​
Author: Meera Ray, Swapnika Dulam, Christopher L. Dancy
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.02807v1 Announce Type: cross Abstract: How can we better represent the impact of sociocultural structures on decision making in computational cognitive models? Modeling this impact requires traversing multiple levels of semantic representation, however it is not immediately clear to a mod...
178. Evading Chain-of-Thought Monitoring Through Model Poisoning ​
Author: Giorgio Severi, Shujaat Mirza, Blake Bullwinkel, Amanda Minnich
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.02820v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning trace is informative about its actions. This work studies the limits of CoT monitoring through the lens...
179. Improved Quantum Algorithms for Reinforcement Learning Under a Generative Model ​
Author: Joao F. Doriguello
Published: 8/5/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG, stat.ML
arXiv:2608.02826v1 Announce Type: cross Abstract: Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as large a reward as possible. A standard approach to study such interaction is through Markov Decision Processes (MD...
180. In-Context Collapse in Vision-Language Models and How to Mitigate it? ​
Author: Mohammad Rostami
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02830v1 Announce Type: cross Abstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied. We show the opposite: as demonstrations accumu...
181. CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning ​
Author: Xuehang Guo, Pingyue Zhang, Ruiyi Zhang, Zhenhailong Wang, Hanrui Lyu, Heng Ji, Tong Sun, Qingyun Wang, Manling Li
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.02833v1 Announce Type: cross Abstract: Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains. While extrinsic chain...
182. MutMem: Cryptographically Authorized Mutation in Persistent Agent Memory ​
Author: Walid Saidi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02843v1 Announce Type: cross Abstract: Persistent agent memory must adapt as later outcomes change earlier evidence, yet mutable retrieval weights create an attribution problem: reviewers must distinguish authorized adaptation from database tampering. We present MutMem, an authorized-muta...
183. BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences? ​
Author: Soumadeep Saha, Krish Sharma, Akshay Chaturvedi, Nicholas Asher
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.02867v1 Announce Type: cross Abstract: Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reasoning tasks, there is significant debate as to whether RLVR expands the reasoning capability boundary,...
184. Robust Counterfactual Policy Optimisation via Nondeterministic Causal Models ​
Author: Jessica Lally, Milad Kazemi, Nicola Paoletti, David Watson, Sander Beckers
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02893v1 Announce Type: cross Abstract: Counterfactual inference approaches for sequential decision-making typically assume deterministic causal models, where all randomness stems from latent variables. However, Markov Decision Processes (MDPs) are inherently stochastic. We address this by...
185. When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index ​
Author: Kleyton da Costa, Bernardo Modenesi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02938v1 Announce Type: cross Abstract: Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and heterophilic graphs want different attention shapes, and one fixed normalization cannot serve both. We propose \textb...
186. Rubrics as Privileged Information for Open-Ended Generation ​
Author: Deepika Bablani, Ajay Gupta, Wanming Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02948v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in verifiable domains like math, where hard privileged information (PI) in the form of ground-truth answers structura...
187. SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling ​
Author: Evan Assmus, Qining Zhang, Lei Ying
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02951v1 Announce Type: cross Abstract: Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model. Existing reward-model-free methods are either restricted to bandits or deterministic MDPs, such as DPO or P3O, or use zeroth-order, gra...
188. Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits ​
Author: John Hu, Andrew Ash
Published: 8/5/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY, eess.IV
arXiv:2608.02955v1 Announce Type: cross Abstract: This research paper describes an exploratory study on the effectiveness of Chat Debugging: troubleshooting malfunctioning analog circuits on breadboards and printed circuit boards (PCB) by undergraduates through conversations with public-domain large...
189. ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies ​
Author: Inkyu Sa, Konstantin Stulov, Rajat Bhageria
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.02958v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress. Reinforcement learning wo...
190. Scaling an Autoregressive Transformer for Single-Cell Generation ​
Author: Aleksandr Sharipov, Yusif Mukhtarov, Igor Molybog
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.GN
arXiv:2608.02961v1 Announce Type: cross Abstract: We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to generate additional gene expression vectors of that cell type. For this task we characterize both the biological fi...
191. HyperFL: Query-Adaptive Representation Learning for Software Fault Localization ​
Author: Shuai Shao, Yiming Zeng, Yu Zhao, Tingting Yu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.02967v1 Announce Type: cross Abstract: Software fault localization identifies the code locations responsible for reported issues and is a fundamental step toward automated debugging and program repair. Recent retrieval-based approaches formulate fault localization as a dense retrieval tas...
192. TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation ​
Author: Bhavin Jawade, Cameron R. Wolfe
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.02975v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated impressive performance in MQM-based translation quality (TQ) evaluation, and recent advances in large reasoning models (LRMs) promise even greater improvements. However, both LLMs and LRMs are computatio...
193. Internalising the Identity Primitive: Cryptographic Individuality for an Autonomous Agent on a Public Blockchain ​
Author: Keisuke Suzuki
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.MA
arXiv:2608.02986v1 Announce Type: cross Abstract: A software agent on a public blockchain accumulates authority and economic stakes, raising the engineering question of what makes it count as an individual. The paper's central contribution is a shift of trust root for the key-to-weights binding of a...
194. SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels ​
Author: Yongwan Jo, Jinyoung Park, Euihyun Lee, Dokyung Song
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02995v1 Announce Type: cross Abstract: Modern large language models (LLMs) exhibit activation sparsity, wherein only a subset of their neurons is activated for given input tokens. Researchers have leveraged this property to optimize LLM serving systems by omitting weight accesses and comp...
195. V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery Detectors ​
Author: Shichao Kan, Chengpeng Hong, Jingtong Dou, Chuancheng Shi, Yuhan Liu, Linrui Xu, Yixiong Liang, Yigang Cen, Yanpeng Sun, Fei Shen, Tat-Seng Chua
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03008v1 Announce Type: cross Abstract: As generated videos become increasingly realistic, reliable video forgery detection is increasingly important. Existing studies typically optimize and use video forgery detectors as black boxes, while the latent forgery-discriminative knowledge insid...
196. A Graph Signal Processing Perspective on Numerical Sequence Representations in LLM In-Context Learning ​
Author: Jiajun Bao, Zihao Qi, Toni J. B. Liu, Gurbir Arora, Rapha"el Sarfati, Nicolas Boull'e, Christopher J. Earls
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.SP
arXiv:2608.03015v1 Announce Type: cross Abstract: Pretrained large language models (LLMs) have demonstrated in-context learning (ICL) capabilities for numerical inference over sequences serialized as text. Prior work has identified and characterized this form of numerical inference primarily through...
197. Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing ​
Author: Changhao Zhao, Haoxiang Li, Yuke Li, Hai Liu, LingLin Zeng
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03023v1 Announce Type: cross Abstract: Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recently, the recent release of DINOv3 brings DINO.txt, which equips the standalone DINO backbone with image-text co...
198. PACE: Adaptive Budget Allocation for Time-Efficient Embodied Planning ​
Author: Yuchen Huang, Xijiang Ying, Zhenhua Ma, Xiaxiang Yuan, Zhijie Gao, Jiayi Huang, Ruichi Mao, Jiazheng Zhang, Hongsheng Ti, Maotao Tian, Rong Shi, Lu Zhao, Shizhuang Zhang, Zhuo Cui, He Wang, Ling Liu, Wei Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.03034v1 Announce Type: cross Abstract: Reasoning-enhanced large language models have achieved remarkable improvements in planning tasks, yet their deployment in embodied systems remains impractical due to prohibitive inference delays-often exceeding minutes per planning instance. The fund...
199. LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs ​
Author: Forough Majidi, Mohammad Mehdi Morovati, Foutse Khomh, Heng Li
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2608.03036v1 Announce Type: cross Abstract: Large Language Models (LLMs) are integrated into software systems and AI services, making efficient LLM serving a concern for software engineering. Serving LLMs is challenging because inference requires computation, memory, GPU resources, and executi...
200. PLAN: Parallel Liquid-Inspired Approximation Network for Efficient Representation Learning in Flexible Job Shop Scheduling ​
Author: Dhivya Dharshini Kannan, Wei Zhang, Jieyi Bi, Yingpeng Du, Tianjun Wei, Jie Zhang, Zuming Liu, Anupam Trivedi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03041v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) approaches for flexible job shop scheduling (FJSP) heavily rely on attention-centric architectures to achieve state-of-the-art performance. However, these models suffer from excessive parameter counts and prohibitive...
201. Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation ​
Author: Seth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03044v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity. We ...
202. PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory ​
Author: Dawei Liu, Haixu Song, Shuang Cheng, Shijie Wang, Haozheng Hou, Kaifeng Liu, Ermo Hua, Zhonghang Yuan, Zhijie Zhong, Yuchen Fan, Biqing Qi, Bowen Zhou
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03048v1 Announce Type: cross Abstract: Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: sequential chunk-wise updates can overwrite early critical evidence with later irrelevant content, and...
203. Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping ​
Author: Jingwei Zhao, Gus Xia, Ziyu Wang, Ye Wang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM, eess.AS
arXiv:2608.03050v1 Announce Type: cross Abstract: What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implici...
204. CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning ​
Author: Ziqi Jia, Yalu Ouyang, Bo Pang, Panpan Li, Hangfei Xu, Shengzhao Wen, Shiyong Li, Yanpeng Wang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.03068v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). However, existing methods suffer from insufficient precision in feedback on generated answer trajectories and exh...
205. AI Security Leaderboard: Methodology, Results and Minimal Standard ​
Author: Jasper Timm, Lukas Struppek, Ziwei Xu, Grace Cheong, Oscar Mata, Dan Zhao, Mick Yang, Isadora De Andrade, Xiaojun Jia, Yiming Li, Samuel Bauer, Heather McIntyre, Adam Gleave, Edward Yee, Kellin Pelrine
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.03070v1 Announce Type: cross Abstract: Frontier AI model developers increasingly rely on layered safeguards to prevent catastrophic misuse, but little public evidence exists on how much protection these safeguards provide, or how consistently across developers. We introduce the FAR.AI Min...
206. CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation ​
Author: Ting Yin, Danning Li, Chen Shu, Xiaoxia Yao, Boyu Fu, Yujing Chang, Tianyu Shi, Mengna Feng, Jie Chen, Jing Fu, Xiuli Xiao, Tianlin Li, Mumin Shao, Jiaxin Bi, Wenchuan Zhang, Xiaoyan Wu, Xiao Han, Zhang Zhang, Yuhao Yi, Hong Bu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, stat.AP
arXiv:2608.03079v1 Announce Type: cross Abstract: Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can obscure subtype distinctions. We developed CorePath, a breast-spec...
207. SynEnergy: Anomaly Semantic-Guided Diffusion for Synthetic Energy Data Generation ​
Author: Lin Jiang, Dahai Yu, Ravikumar Gelli, Guang Wang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03087v1 Announce Type: cross Abstract: Fine-grained energy consumption data are essential for applications such as demand forecasting, demand response planning, and grid reliability assessment. However, access to such data is often restricted by privacy concerns and data-sharing constrain...
208. SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation ​
Author: Wen Wang, Jiahua Bao, Tu Yongsiqi, Yihao Liu, Haotian Zhou, Haoxuan Ma, Mengyu Zhou, Wenkui Fan, Junwei He, Xiaoxi Jiang, Guanjun Jiang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.03092v1 Announce Type: cross Abstract: We aim to improve model performance in multi-reward reinforcement learning training process. Existing Group reward-Decoupled Normalization Policy Optimization (GDPO) has mitigated the issue of reward signals masking one another during direct scalariz...
209. FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection ​
Author: Pei Li, Sihan Chen, Delong Ran, Tianshuo Cong
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CV
arXiv:2608.03096v1 Announce Type: cross Abstract: Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped. In particular, the effectiveness of image-level detectors in the video domain ha...
210. A Hierarchical Approach to Imitation Learning for Manipulation Tasks Requiring Time Varying Forces ​
Author: Rishabh Shukla, Adithya Santhosh, Shaili Gandhi, Samrudh Moode, Satyandra K. Gupta
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.03103v1 Announce Type: cross Abstract: Diffusion policies have shown strong performance in learning complex, multi-modal behaviors for robotic manipulation. However, their application to contact-rich disassembly tasks remains limited by a key trade-off: the iterative denoising process int...
211. Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models ​
Author: Paribesh Regmi, Qingshuang Chen, Chi Zhang, Heba Aly, Yelin Kim, Hongda Mao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03112v1 Announce Type: cross Abstract: Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens per image, limiting their deployment on resource-constrained edge devices and in real-time surveillanc...
212. Optimal Liability Design for Medical AI ​
Author: Rui Mao, Tingliang Huang, Houcai Shen
Published: 8/5/2026, 4:00:00 AM
Categories: econ.TH, cs.AI, cs.CY
arXiv:2608.03114v1 Announce Type: cross Abstract: Artificial intelligence (AI) is increasingly integrated into medical decision-making, yet its liability implications remain complex, particularly when physicians differ in diagnostic skills and their quality is unobservable. This paper develops a pri...
213. Trajectory-Guided Forget-Recover Network for Continual LLM Unlearning ​
Author: Zezheng Wu, Xinghe Cheng, Qinggang Zhang, Haoran Luo, Jiapu Wang, Qing Yang, Jingwei Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03123v1 Announce Type: cross Abstract: Machine unlearning aims to eliminate the influence of sensitive data on a model. In the real world, unlearning requests arrive continually, which gives rise to two challenges. First, an unlearning intervention may redistribute target-related computat...
214. DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units ​
Author: Haoyu Gu, Haotian Lu, Jingrun Du, Xiao-Ping Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.03127v1 Announce Type: cross Abstract: Hand motion carries the finest-grained information in human activity, yet the representations behind hand generation, understanding, and robot learning are overwhelmingly continuous--joint angles or MANO parameters. These are accurate but unstructure...
215. Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds ​
Author: Ning Zhu, An Chen, Mengfei Zhao, Juntao Xu, Jingze Liang, Boyuan Gu, Liang-Jian Deng
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03135v1 Announce Type: cross Abstract: Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We trace these failures to an early coordination bottleneck: before denoising begins, prompt-conditioned a...
216. Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning ​
Author: Meicong Zhang, Tiancheng Su, Jiahao Cheng, Guoxiu He, Xinqi Tao, Dejia Song
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03138v1 Announce Type: cross Abstract: Generating a rigorous paper introduction with large language models (LLMs) remains challenging, since it requires coordinating background, gap identification, method and contribution within a coherent narrative. Existing solutions externalize this pr...
217. Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue ​
Author: Xueping Gong, Zhuoluo Zhang, Zhaowei Miao, Jiheng Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2608.03142v1 Announce Type: cross Abstract: We study contextual dynamic pricing with arbitrary covariate sequences and bounded, possibly nonbinary purchase quantities. Demand follows a semiparametric surplus-index model with an unknown linear valuation parameter and an unknown H"older-smooth ...
218. Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation ​
Author: Sicong Chang, Yidan Shen, Wen Yu, Jiefu Chen, Xin Fu, Renjie Hu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03148v1 Announce Type: cross Abstract: RAG improves the factual grounding of LLM by incorporating external knowledge, but deploying RAG on mobile and edge devices remains challenging because retrieved context increases computation and memory. A direct way to reduce this cost is to retain ...
219. EFX Allocation In (Multi)Hypergraphs ​
Author: Thanasis Lianeas, Alkmini Sgouritsa, Minas Marios Sotiriou
Published: 8/5/2026, 4:00:00 AM
Categories: cs.GT, cs.AI
arXiv:2608.03171v1 Announce Type: cross Abstract: We study fair allocations of indivisible goods among agents with heterogeneous monotone valuations. As fair we consider the allocations that are envy-free-up-to-any-good (EFX). Finding if EFX alloca- tions always exist, even for agents with additive ...
220. Attribute-based Undetectable Watermarking for Generative AI Models ​
Author: Miryam Mi-Ying Huang, Chung-Wei Lee, Max Raffel, Er-Cheng Tang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.03174v1 Announce Type: cross Abstract: Generative AI systems increasingly produce content whose provenance is difficult to verify, motivating watermarking techniques for identifying model-generated outputs. Existing cryptographic watermarking methods provide strong undetectability guarant...
221. Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach ​
Author: Tianbao Jiang, Weicong Ni, Gerard de Melo, Linlin Wang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03204v1 Announce Type: cross Abstract: Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and the requirements of visual reasoning tasks. However, existing RL-based alignment methods are often resource-int...
222. EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners ​
Author: Unggi Lee, Sookbun Lee, Yeil Jeong, Eunjoo Lee, Minchul Shin, Hoilym Kwon
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL
arXiv:2608.03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management s...
223. GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model ​
Author: Guanrou Yang, Tian Tan, Qian Chen, Ziyang Ma, Yakun Song, Zhikang Niu, Qi Chen, Wenming Tu, Haitao Li, Shan Yang, Xie Chen
Published: 8/5/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL
arXiv:2608.03215v1 Announce Type: cross Abstract: Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typically convert the ODE into an SDE and track per-step likelihood ratios, introducing stochastic perturba...
224. Self-Supervised Representation-Guided Generative Dataset Distillation ​
Author: Mingzhuo Li, Guang Li, Linfeng Ye, Jiafeng Mao, Takahiro Ogawa, Konstantinos N. Plataniotis, Miki Haseyama
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03218v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility. Most existing methods target randomly initialized networks, whereas modern vision systems often adapt frozen pretrained encoders...
225. Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks ​
Author: Chenyu Wang, Yunbo Lyu, Junda He, Zhou Yang, Chenxing Zhong, Yaniv Harel, David Lo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.03222v1 Announce Type: cross Abstract: Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. Failed runs tend to be longer and exhibit redundant exploration or looping, suggesting that some fail...
226. Agentic Reinforcement Learning with Self-Distilled Reward Shaping ​
Author: Ranxu Zhang, Guinan Chen, Chenshaodong, Jinghao Lin, Xiaozhou Xu, Sunzhe, Yanyong Zhang, Chao Wang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.03223v1 Announce Type: cross Abstract: Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions deserve credit. Training-only privileged skills can provide denser sup...
227. Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking ​
Author: Jinquan Zhang, Dongfu Yin, Run Yang, Yufeng Yan, Zhen Tian, F. Richard Yu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.03231v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies promise general robotic manipulation, but their robustness against physical-world attacks remains fragile. In particular, we show that physically realizable adversarial patches can reliably induce failures by tri...
228. From Wearable Data to Personalized and Actionable Health Insights ​
Author: Esther Brown, Karis Moon, Victoria Dean, Finale Doshi-Velez
Published: 8/5/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.03251v1 Announce Type: cross Abstract: Commercial wearable devices continuously capture rich physiological data (e.g., heart rate, respiration), opening new possibilities for monitoring health conditions, notably around stress. Despite their promise, turning raw wearable physiological dat...
229. FinVerse: Financial Time-Series Benchmark ​
Author: Jaehoon Lee, Jun Seo, Seunghan Lee, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Minjae Kim, Sungdong Yoo, Junhyeok Kang, Sangjun Han, Soonyoung Lee, Wonbin Ahn
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03259v1 Announce Type: cross Abstract: As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important. Existing time-series forecasting benchmarks provide useful standardized compariso...
230. The Ignition Is Real, and It Lives at the Readout: Latent composition, difficulty-clocked ignition, and the interface-constituted commit in a recurrent-depth reasoner ​
Author: Simon Lam-Muir
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03263v1 Announce Type: cross Abstract: We test whether the "compositional ignition" reported in latent-reasoning models is real computation, an instrument artifact, or inherited from verbal training data. We grow an independent realization of a published 30M-parameter recurrent-depth reas...
231. Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending ​
Author: Chongle Ren, Guang Li, Wenbo Huang, Naoki Saito, Takahiro Ogawa, Miki Haseyama
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03269v1 Announce Type: cross Abstract: Video dataset distillation aims to compress a large video dataset into a compact surrogate set that preserves its training utility. Most existing approaches synthesize condensed videos through iterative optimization, whose cost is amplified by the te...
232. GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs ​
Author: Zichuan Fu, Shirong Wang, Wenlin Zhang, Guojing Li, Yimin Deng, Jingtong Gao, Junjia Qi, Hanyu Yan, Yefeng Zheng, Xiaopeng Li, Wanyu Wang, Xian Wu, Xiangyu Zhao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03270v1 Announce Type: cross Abstract: GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on high-resolution, densely populated interfaces because a vision-language model (VLM) may recognize a requested ...
233. Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates ​
Author: Jinya Sakurai, Shueicheng Yan, Xun Xu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03284v1 Announce Type: cross Abstract: Ensuring safety and policy compliance in text-to-image diffusion models remains a critical challenge, as benign or adversarial prompts can often elicit prohibited content, e.g. nudity and protected intellectual property. While training-based unlearni...
234. The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics ​
Author: Shashwat Sourav, Aishwarya Balwani
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.03291v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model's reasoning process. Existing approaches that leverage verbalized CoTs to monitor reasoning correctness, howeve...
235. Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform ​
Author: Xavier Wrenn, Radoslav Raykov, Aleksandar Angelov, Hirokuni Kitahara, Yuji Watanabe, Anca Sailer
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.03311v1 Announce Type: cross Abstract: Enterprise compliance management requires rapid adaptation to evolving regulatory frameworks (e.g., DORA, AI RMF, FedRAMP) and tight remediation SLAs. Traditional static orchestrators often fail in hybrid cloud environments where event-driven assessm...
236. Route-Align-Verify for Functional Correctness in Code Generation ​
Author: Erxue Zhou, Jingxiang Meng, Aofan Liu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.03341v1 Announce Type: cross Abstract: Large language models (LLMs) have substantially improved code generation, yet achieving strong functional correctness remains difficult, especially for heterogeneous programming tasks where a single prompting strategy and a single directly generated ...
237. When Oracle Conditioning Misleads Deployment: Conditioning-Availability Bias in Echocardiographic Segmentation ​
Author: Dang P. M. Cao, Hieu D. Pham, Hieu Pham
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03342v1 Announce Type: cross Abstract: Conditional segmentation models may be trained and evaluated with auxiliary signals cleaner than those available at deployment. We study this protocol-level manifestation of shortcut learning and auxiliary-variable shift in phase-conditioned echocard...
238. The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk ​
Author: Francis Heylighen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.03361v1 Announce Type: cross Abstract: AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of value i...
239. FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact ​
Author: Alex Kwon
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03372v1 Announce Type: cross Abstract: AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing away what made it checkable, who said it, how sure they were, when it held. We call that failure fac...
240. Shaping Wind-Tunnel Airflow for Unmanned Aerial Vehicles using Online Learning ​
Author: Ghadeer Elmkaiel, Michael Muehlebach
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.SY, eess.SY
arXiv:2608.03378v1 Announce Type: cross Abstract: The development and testing of advanced aerial robots require experiments in controlled environments with tailored airflow profiles. This paper presents an online learning algorithm for controlling the complex airflow field in a multi-fan vertical wi...
241. Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces ​
Author: Francesca Carlon, Vincent Ginis, Andres Algaba
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03401v1 Announce Type: cross Abstract: Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this reasoning, but a shorter trace may only show that the model stopped sooner. Here, we evaluate paired runs of the...
242. Distilled Roads: Generalisable Road Network Extraction Across Sensors, Resolutions, and Region ​
Author: Sanayya, Rakshith Sathish, Ashwathi Nambiar
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.03407v1 Announce Type: cross Abstract: Road network segmentation from satellite imagery remains challenging due to large geographic variation in road appearance, occlusions, and domain shifts introduced by differing resolutions and sensors. Existing models, typically trained under narrow ...
243. Multi-Task Multi-Frame Visual Piano Transcription ​
Author: Yonghyun Kim, Hoyeol Sohn, Juhan Nam, Alexander Lerch
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CV, cs.MM, eess.IV
arXiv:2608.03419v1 Announce Type: cross Abstract: Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano ...
244. OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet ​
Author: Dimitrios I. Zaridis, Traianos Tsiokris, Vasileios C. Pezoulas, Daphni Plati, Eugenia Mylona, Eleni Georga, Nikos Tsiknakis, Antonis Sakellarios, Dimitrios I. Fotiadis
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due to high intra-class variability and visually similar dishes. This study presents OliveGemma, a visio...
245. A Low-Cost Hybrid Reservoir Computing Model for Isolated Sign Language Video Recognition ​
Author: Nitin Kumar Singh, Arie Rachmad Syulistyo, Yuichiro Tanaka, Hakaru Tamukoh
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.03444v1 Announce Type: cross Abstract: Sign language recognition (SLR) enhances communication between hearing and hearing-impaired individuals. Although deep learning (DL) has achieved promising performance in SLR, its high computational cost limits deployment on edge devices. To address ...
246. Approximate Speculative Decoding ​
Author: Yuannuo Feng, Zegang Peng, Yuxin Xie, Yubing Ye, Yizhe Chen, Wenshuai Yao, Wenyong Zhou, Wang Kang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03447v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verification, decoding stops at the first draft token that differs from the target argmax, discarding the rem...
247. Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs ​
Author: Haoqian Kang, Liupeng Li, Kuofeng Gao, Jinpeng Wang, Zhenyu Lu, Bin Chen, Ke Chen, Yaowei Wang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.MM, cs.AI, cs.CL, cs.CV
arXiv:2608.03450v1 Announce Type: cross Abstract: Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction. Explicit text-based Chain-of-Thought (CoT) is computationally expensive and prone to visual hallucinations, while exist...
248. Adaptive Modality Reliability Diagnosis and Restoration for Robust Multimodal Intent Recognition ​
Author: Suraj Kumar, Mohnish Raj, Soumi Chattopadhayay, Chandranath Adak, Ayan Dutta
Published: 8/5/2026, 4:00:00 AM
Categories: cs.MM, cs.AI, cs.CL
arXiv:2608.03475v1 Announce Type: cross Abstract: Multimodal intent recognition combines linguistic, acoustic, and visual evidence, but individual modalities may be noisy, missing, semantically conflicting, or disproportionately dominant. Existing methods typically infer modality importance implicit...
249. Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution ​
Author: Weichen Xu, Zhenhua Liu, Lin Luo, Yaobo Liang, Chengtang Yao, Qingyu Mei, Jian Cao, Xixin Cao, Xing Zhang, Jiaolong Yang, Baining Guo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.LG
arXiv:2608.03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.e., execution horizon) before replanning, turning replanning into a task-agnostic periodic schedule that is independent of task progress. As a result, when ...
250. Leveraging System-Level Observations to Inform Bayesian Learning of Model Parameters for Quantitative Verification ​
Author: Simos Gerasimou, Xingyu Zhao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LO
arXiv:2608.03489v1 Announce Type: cross Abstract: Combining Bayesian learning and quantitative verification is a powerful toolset for analysing key quantitative properties of software systems, like reliability and response time. However, the accuracy and robustness of verification results strongly d...
251. Principles of Robot Autonomy ​
Author: Daniele Gammelli, Joseph Lorenzetti, Katie Luo, Gioele Zardini, Marco Pavone
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.SY, eess.SY
arXiv:2608.03496v1 Announce Type: cross Abstract: Autonomous robots are moving rapidly from research labs into everyday life - on roads, in the air, in warehouses, and in space. Robot autonomy is no longer solely an academic pursuit, but a collection of mature, field-tested methods and tools that pr...
252. ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels ​
Author: Gagan Bhatia, Julian Schlenker, Simone Paolo Ponzetto, Steffen Eger
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03507v1 Announce Type: cross Abstract: Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages....
253. How Many Labels Are Enough? ALDA: Active Learning Deployment Advisor for Medical Image Classification ​
Author: Julia Machnio, Mads Nielsen, Mostafa Mehdipour Ghazi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.03511v1 Announce Type: cross Abstract: Active learning (AL) promises to reduce the cost of medical imaging projects by lowering the number of clinical labels required. However, practical deployment requires committing to a sampling strategy before the full annotation budget is spent, and ...
254. AI Forensics Across White-, Grey-, and Black-Box Access: A Process Model and Research Agenda for Post-Incident Investigation of AI Systems ​
Author: Ali Dehghantanha, Sajad Homayoun
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.03520v1 Announce Type: cross Abstract: AI systems are increasingly involved in decisions and actions that may later require investigation. When an AI related incident occurs, investigators need to reconstruct what the system did, why it behaved that way, and which part of the system or su...
255. Pivot-Centric Trajectory Prediction: Bridging Long Horizons via Dynamical Guidance ​
Author: Xiucong Zhao, Jindong Tian, Hao Miao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.03521v1 Announce Type: cross Abstract: Forecasting precise future motion of surrounding agents is essential for reliable autonomous vehicles. However, as the demand for longer prediction horizons increases, existing endpoint-completion or iterative-refine methods increasingly struggle wit...
256. Training Documents Reranker with Search Rubrics for Deep Research Agent ​
Author: Wenhan Liu, Yu Lu, Qiaolin Xia, Hui Xu, Tong Zhao, Jian Xi, Yutao Zhu, Haijin Liang, Haibo Shi, Hao Wang, Zhicheng Dou
Published: 8/5/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2608.03527v1 Announce Type: cross Abstract: Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typically select documents through relevance matching, while individually well-matched top-$k$ documents may not f...
257. Pin Once, Swap Light: Subspace-Aligned Centroid-Residual Training for Efficient Ultra-LoRA Serving ​
Author: Xiang Li, Pengcheng Wang, Huazheng Wang, Saurabh Bagchi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03579v1 Announce Type: cross Abstract: Modern multi-tenant Low-Rank Adapters (LoRAs) serving systems concurrently host tens to hundreds of LoRA adapters. Though powerful, this introduces a critical system dilemma between serving efficiency and task performance: higher-rank adapters genera...
258. AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality ​
Author: Alexander M. Fichtl, Lukas Ellinger, Josefin Kelber, Kry\v{s}tof Ol'ik, Georg Groh
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.03581v1 Announce Type: cross Abstract: AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems ...
259. GenOS: Compositional Certificates for Semantic Robustness in AI Code Generation ​
Author: Corrado Priami
Published: 8/5/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.SE
arXiv:2608.03588v1 Announce Type: cross Abstract: AI coding agents are stochastic workflows: prompts are interpreted, artifacts are sampled, validators produce observations, and orchestrators commit or repair. Small prompt or specification changes can therefore alter program-behavior distributions e...
260. DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction ​
Author: Xuyang Liu, Yibin Han, Zhenwei Zhang, Kai Chang, Zhiwei Xu, Tian Qiu, Weixian Deng, Jiabao Gao, Xiaolin Peng, Hai Wan, Xibin Zhao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.03591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions. However, existing benchmarks mainly evaluate final outputs or aggreg...
261. A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation ​
Author: Abdallah Khemais
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03620v1 Announce Type: cross Abstract: Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different objects: one forward pass versus the parameters behind every forward pass. We ask when they agree. We study an idea...
262. A Security-Oriented Lifecycle Model for Large Language Model Systems ​
Author: Eleftherios Batzolis, George Drosatos, Vassilis Katsouros, Konstantinos Rantos
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CY
arXiv:2608.03626v1 Announce Type: cross Abstract: Large language models are being integrated into critical infrastructure and enterprise workflows at unprecedented scale,yet the lifecycle frameworks governing their development and operations were designed for operational efficiency rather than secur...
263. MuEvo: LLM-Driven Evolution of Multi-Heuristic Ensemble ​
Author: Haoze Lv, Ning Lu, Shengcai Liu, Shaofeng Zhang, Ke Tang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.NE, cs.AI
arXiv:2608.03636v1 Announce Type: cross Abstract: Large language model-based automated heuristic design (LLM-AHD) has shown strong potential in discovering effective heuristics for combinatorial optimization problems. However, existing methods primarily optimize a single heuristic, whereas practical...
264. Decoupling Generation and Selection for Budget-Constrained Faithful Summarization ​
Author: Zeyu Wang, Guanghua Wang, Meng Xu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03655v1 Announce Type: cross Abstract: Abstractive summarization models remain vulnerable to factual inconsistency, redundancy, and weak length control. We propose a modular generation-and-selection framework for sentence-budget-constrained summarization. A pretrained generator produces m...
265. How Closely Do LLM Reviews Align with Human Peer Review? ​
Author: Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled sett...
266. Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation ​
Author: Khai-Nguyen Nguyen, Oscar Chaparro, Antonio Mastropaolo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CV
arXiv:2608.03691v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to translate webpage screenshots into front-end code, but repeated UI patterns may sway them toward visually incorrect yet pattern-consistent outputs. In this work, we test how repeated w...
267. LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation ​
Author: Fan Yang, Yuting Su, Xiaobo Wang, Yuncheng You, Fugui Fan, Yuting Wu, Minghui Wu, Chenxu Zhao, JiaHong Ning, Peiguang Jing
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.03701v1 Announce Type: cross Abstract: World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticipate how a scene will evolve. However, existing WAMs often incur substantial computational overhead. ...
268. GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models ​
Author: Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB
arXiv:2608.03729v2 Announce Type: cross Abstract: Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess no repr...
269. AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities ​
Author: Sandy Abdo, Bill Kapralos, Priyamvada Tripathi, KC Collins, Adam Dubrowski
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.03742v1 Announce Type: cross Abstract: Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial intelligence (AI)-driven audio generative models ar...
270. Can LLMs Test Terminal User Interfaces? ​
Author: Chao Peng, Ruida Hu, Ajitha Rajan, Tegawend'e F Bissyand'e, Jacques Klein, Cuiyun Gao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2608.03743v1 Announce Type: cross Abstract: Terminal User Interfaces (TUIs) combine the stateful, screen-oriented behaviour of GUIs with terminal deployment and are now common in developer tools. Yet they lack a dedicated testing methodology. We survey 197 real-world TUI applications: only 12%...
271. MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models ​
Author: Tong Ling, Hang Lei, Feng Xiao, Changhui Sun, Jiahang Xie, Hao Liu, Lu Liu, Yanlong Du
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03769v1 Announce Type: cross Abstract: Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundamentally from that of autoregressive (AR) models. Whereas AR decoding exposes a contiguous prefix, MDLM ...
272. Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks ​
Author: Shunfan Zheng, Dongsheng Shi, Yue Li, Xin Yi, Linlin Wang, Gerard de Melo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CL
arXiv:2608.03794v1 Announce Type: cross Abstract: Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database administrators (DBAs). However, current evaluation benchmarks remain disproportionately fixated on Text-to-SQL...
273. Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss ​
Author: Bakbergen Ryskulov, Iker Garc'ia-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Rom'an Or'us
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.03796v1 Announce Type: cross Abstract: Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD). This recover...
274. Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure ​
Author: Holly Lewis (Southern Illinois University Carbondale)
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.SI
arXiv:2608.03800v1 Announce Type: cross Abstract: An LLM-based agent is a loop that reads itself. Agentic frameworks externalize identity, memory, and disposition into editable files. The agent loads and edits these files during each activation. I argue that this architecture produces a capacity I c...
275. VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs ​
Author: Andrei Chetvergov, Alexander Evseev, Timofei Sivoraksha, Stepan Ukolov, Mikhail Solovev, Danil Sazanakov, Sergey Bolovtsov
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03810v1 Announce Type: cross Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favor...
276. UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space ​
Author: Amir Mohammad Ezzati, Kiyan Rezaee, Bardiya Kariminia, Mohamad Amin Yousefi, Asal Mohammadjafari Mamaqani, Behrad Samimi, Mohammad Hossein Rohban
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03817v1 Announce Type: cross Abstract: Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence. Existing black-box hallucination detection methods estimate unc...
277. FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis ​
Author: Zhang Weihui, Wang Ruizhi, Xu Hongye, Wang Huiqiong, Sun Li, Song Mingli
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03822v1 Announce Type: cross Abstract: Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific image generation. Although generative models provide a promising means of data augmentation, existing methods oft...
278. Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking ​
Author: Peijia Guo, Wenxuan Xie, ZiGuang Li, Ming Li
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03859v1 Announce Type: cross Abstract: Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexplored and largely unresolved challenge. Prior work on LLM-generated-text detection targets AI involvement, whic...
279. SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG ​
Author: Kaysarul Anas Apurba, Md. Hasibul Hasan, Rofiqul Alam Shehab, Asab Azad
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR, cs.PF
arXiv:2608.03860v1 Announce Type: cross Abstract: We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific RAG pipeline across three corpus scales: 1,034 ch...
280. GENESIS: Towards Explainable Causal Discovery ​
Author: Abhinav Thorat, Ravi Kumar Kolla, Vishak K Bhat, Harsh Vardhan Singh Chauhan, Niranjan Pedanekar
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03868v1 Announce Type: cross Abstract: Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to resolve structural ambiguities in low-sample regimes. Second, although LLM-assisted hybrid approaches improve st...
281. Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning ​
Author: Pyrros Koussios, Chenhao Li, Xin Chen, Andreas Krause
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03875v1 Announce Type: cross Abstract: Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Models (VLMs) as reward models, computing text-observation similarity to bypass manual reward engineerin...
282. MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning ​
Author: Martin B"ockling, Elizaveta Nosova, Heiko Paulheim, Andreea Iana
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2608.03882v1 Announce Type: cross Abstract: Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logistics, yet large language models (LLMs) struggle with the required geometric and topological computat...
283. CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement ​
Author: Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel, Niharika Vadlamudi, Nikhilesh Chowdary Eathamukkala, Prasanth V V, Abhyuday Kumara Swamy, Pranay Narhari Umredkar, Pradeep Narayan, Vivek Rajagopal, Tanuja Ganu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend. Today's Vi...
284. When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding ​
Author: Ke Li, Jiayu Chen, Maoliang Li, Zihao Zheng, Hailong Zou, Hengyi Zhang, Xuanzhe Liu, Xiang Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03918v1 Announce Type: cross Abstract: Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based methods rely on static one-shot selection with fixed frame budgets and can...
285. Equivariant Music Transformer ​
Author: Zixun Guo, Simon Dixon
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.03920v1 Announce Type: cross Abstract: Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space. Our analysis, however, shows that standard music transformers map such time-shifted or pitch-t...
286. PRISM: Powerful Time Series to Image (TS2I) Representations for Multivariate Anomaly Detection ​
Author: Mateusz Smendowski, Kamil Faber, Piotr Nawrocki, Nathalie Japkowicz, Roberto Corizzo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2608.03926v1 Announce Type: cross Abstract: Time series anomaly detection (TSAD) underpins applications in predictive maintenance, finance, and cloud computing, however performance remains sensitive to representation choices, especially in multivariate settings. While transforming time series ...
287. Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility ​
Author: Jo-Ku Cheng, Nikolaos Aletras, Marco Valentino
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.03930v1 Announce Type: cross Abstract: Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretraining tasks, such as Dyck and procedural algorithms, rely on narrow primitives that fail to capture the expres...
288. Separating quantum circuits from classical LLMs ​
Author: Srinivasan Arunachalam, Arkopal Dutt, Hari Krovi, Rik Sengupta
Published: 8/5/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.CC
arXiv:2608.03962v1 Announce Type: cross Abstract: Modern large language models - transformers and diffusion language models - are built around two canonical algorithmic tasks: prediction and generation. We prove unconditional separations between low-depth quantum computation and the corresponding bo...
289. Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent ​
Author: Zhen Fang, Yu Zeng, Wenxuan Huang, Yiming Zhao, Shiting Huang, Tianfei Ren, Qi Lu, Qingnan Ren, Qisheng Su, Lionel Z. Wang, Qingyu Yin, Shuang Chen, Zehui Chen, Lin Chen, Zhenfei Yin, Yao Hu, Shaohui Lin, Wanli Ouyang, Shaosheng Cao, Feng Zhao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.03979v1 Announce Type: cross Abstract: We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critic...
290. Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss? ​
Author: Hailong Jiang, Feng Yu, Emran Hossain, Jianfeng Zhu, Mengfei Ren, Qiang Guan, Chunwei Xia
Published: 8/5/2026, 4:00:00 AM
Categories: cs.PL, cs.AI
arXiv:2608.03983v1 Announce Type: cross Abstract: Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed program representation. We ask whether large language models (LLMs) can recover such semantics from heterogeneous C/C++ context and realiz...
291. Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility ​
Author: Mohsen Hariri, Weicong Chen, Nahal Shahini, Vikash Singh, Kai Ye, Amirhossein Samandar, Debargha Ganguly, Sreehari Sankar, Yanyan Zhang, Shouren Wang, Jerry Peng, Biyao Zhang, Michael Hinczewski, Vipin Chaudhary
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04001v1 Announce Type: cross Abstract: Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algorithms that extend deliberation along a single trajectory, sample comple...
292. TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning ​
Author: Changle Qu, Sunhao Dai, Hengyi Cai, Yuqi Zhou, Xinran Chen, Simon, Jun Xu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04007v1 Announce Type: cross Abstract: Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit assignment in long-hori...
293. A Unified Framework for Human AI Collaboration in Security Operations Centers with Trusted Autonomy ​
Author: Ahmad Mohsin, Helge Janicke, Ahmed Ibrahim, Iqbal H. Sarker, Seyit Camtepe
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2505.23397v3 Announce Type: replace Abstract: This article presents a structured framework for Human-AI collaboration in Security Operations Centers (SOCs), integrating AI autonomy, trust calibration, and Human-in-the-loop decision making. Existing frameworks in SOCs often focus narrowly on au...
294. Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning ​
Author: Alexander Meulemans, Rajai Nasser, Maciej Wo{\l}czyk, Marissa A. Weis, Seijin Kobayashi, Blake Richards, Guillaume Lajoie, Angelika Steger, Marcus Hutter, James Manyika, Rif A. Saurous, Jo~ao Sacramento, Blaise Ag"uera y Arcas
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2511.22226v3 Announce Type: replace Abstract: The standard theory of model-free reinforcement learning assumes that the environment dynamics are stationary and that agents are decoupled from their environment, such that policies are treated as being separate from the world they inhabit. This l...
295. OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Algorithm Discovery ​
Author: Qi Liu, Ruochen Hao, Can Li, Wanjing Ma
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.NE
arXiv:2602.13769v3 Announce Type: replace Abstract: Automating heuristic design in complex, experiment-driven domains requires more than iterative mutation of solution algorithms. Current LLM-based evolutionary methods often rely on stochastic mutation loops that lack long-term strategic planning an...
296. Modeling Matches as Language: A Generative Transformer Approach for Counterfactual Player Valuation in Football ​
Author: Miru Hong, Minho Lee, Geonhee Jo, Hyeokje Cho, Hyunsung Kim, Pascal Bauer, Sang-Ki Ko
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2603.15212v2 Announce Type: replace Abstract: Evaluating football player transfers is challenging because player actions depend strongly on tactical systems, teammates, and match context. Despite this complexity, recruitment decisions often rely on static statistics and subjective expert judgm...
297. Assessing the Effect of Cross-Domain Mapping on Creativity in Humans and Large Language Models ​
Author: Qiawen Ella Liu, Marina Dubova, Henry Conklin, Takumi Harada, Thomas L. Griffiths
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2603.19087v2 Announce Type: replace Abstract: Creativity is the ability to come up with novel ideas, a capacity crucial for human development and flourishing. Are large language models (LLMs) creative in the same way humans are, and can the same interventions increase creativity in both? We st...
298. LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics ​
Author: Farhan Ahmed, Yuya Jeremy Ong, Chad DeLuca
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IT, math.IT
arXiv:2603.24929v2 Announce Type: replace Abstract: Understanding and quantifying uncertainty in large language model (LLM) outputs is critical for reliable deployment. However, traditional evaluation approaches provide limited insight into model confidence at individual token positions during gener...
299. What Makes a Sale? Simulating End-to-End Seller--Buyer Retail Dynamics with LLM Agents ​
Author: Jeonghwan Choi, Jibin Hwang, Gyeonghun Sun, Minjeong Ban, Taewon Yun, Hyeonjae Cheon, Hwanjun Song
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2604.04468v3 Announce Type: replace Abstract: Evaluating retail strategies before deployment is difficult, as outcomes are determined across multiple stages, from seller-side persuasion through buyer-seller interaction to purchase decisions. However, existing retail simulators capture only par...
300. AI Assistance Reduces Persistence and Hurts Independent Performance ​
Author: Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel A. Bakker, Rachit Dubey
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.04721v4 Announce Type: replace Abstract: People often optimize for long-term goals in collaboration: A mentor or companion doesn't just answer questions, but also scaffolds learning, tracks progress, and prioritizes the other person's growth over immediate results. In contrast, current AI...
301. An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness ​
Author: Ioannis Bilionis, Ricardo C. Berrios, Luis Fernandez-Luque, Carlos Castillo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.23954v2 Announce Type: replace Abstract: Artificial Intelligence (AI) and Machine Learning (ML) models used in clinical settings are increasingly deployed to support clinical decision-making. However, when training data become stale due to changes in demographics, environment, or patient ...
302. Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective ​
Author: Hamid Osooli, Kareema Batool, Rick Gentry, Tiasa Singha Roy, Ashwin Gupta, Anirudha Ramesh
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.25077v3 Announce Type: replace Abstract: Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on examples that lie in the weak model's blind spots. Understanding such failures requires going beyond aggrega...
303. NOVA: Fundamental Limits of Knowledge Discovery Through AI ​
Author: Salman Avestimehr, Ken Duffy, Muriel M'edard
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.IT, math.IT
arXiv:2605.15219v3 Announce Type: replace Abstract: Can AI systems discover new knowledge through iterative self-improvement, and at what cost? We introduce NOVA, which models the ``generate, verify, accumulate, retrain'' loop as an adaptive sampling process over a knowledge space. We give sufficien...
304. Language model agents show in-group trust bias invisible to standard behavioural audits ​
Author: Messi H. J. Lee
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.28114v2 Announce Type: replace Abstract: Language-model agents are moving from single-user assistants into persistent networks that build trust and reputation with one another, and the same models increasingly control physically embodied robots as well as software. Here we show that five ...
305. Designing for Doubt: The Case for Informed Abstention in Autonomous Agents ​
Author: Victor Ojewale, Suresh Venkatasubramanian
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.02965v2 Announce Type: replace Abstract: As large language models gain tool access and are deployed as autonomous agents capable of editing records, executing transactions, and modifying infrastructure, we still evaluate them based on the sole metric of task completion. We argue that this...
306. Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models ​
Author: Dewi Gould, Francis Rhys Ward, Anders Cairns Woodruff, Rauno Arike, Josh Hills, Alex Serrano, Ida Caspary, Jason Ross Brown, Jo J. Jiao, Patrick Leask, Twm Stone, Ram Potham, Ionut Gabriel Stan, Harry Mayne, Simeon Hellsten, Shubhorup Biswas, Ariana Azarbal, William L. Anderson, Elle Najt, Ryan Greenblatt, Julian Stastny
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.07157v4 Announce Type: replace Abstract: Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perform sufficiently complex reasoning internally, without explicit thinking tokens, this would undermine such ov...
307. Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking ​
Author: Jiwan Chung, JiHyuk Byun, Vibhav Vineet, Seon Joo Kim
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2606.15673v2 Announce Type: replace Abstract: Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and offering little guidance on improvement. In this work, we conduct a process-level analysis of web agen...
308. Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters? ​
Author: Kaiwen Zheng, Junchen Fu, Wenhao Deng, Hu Han, Joemon M. Jose, Xuri Ge
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.MM
arXiv:2607.12787v2 Announce Type: replace Abstract: Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc. Ho...
309. Cura 1T: Specialized Model for Agentic Healthcare ​
Author: actAVA AI, :, Haolin Chen, Leon Qi, Steve Brown, Deon Metelski, Tao Xia, Joonyul Lee, Qixuan Wang, Kevin Riley, Frank Wang, Weiran Yao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15314v2 Announce Type: replace Abstract: Healthcare AI agents handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use, yet specialized agentic models that cover these use cases together remain limited. These ...
310. PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails ​
Author: Seungbin Yang, Chaewoon Ki, Dohyun Lee, Jaegul Choo, ChaeHun Park
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.20482v2 Announce Type: replace Abstract: Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspecified instructions, requiring agents to infer the missing context from their raw browsing histori...
311. OPOD: On-Policy Omni Distillation ​
Author: Tong Zhao, Yuyang Hu, Yutao Zhu, Reed Li, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.20918v3 Announce Type: replace Abstract: Omni-modal models provide a unified interface for text, images, and audio. However, improving these abilities together remains difficult, as post-training on pooled multimodal data often fails to preserve the strengths of modality teachers. On-poli...
312. CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics ​
Author: Xinhong Xu, Yimeng Zhang, Yuanlong Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.23258v2 Announce Type: replace Abstract: Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central question remains unresolved: \textbf{whether a model pretrained on one collection of recordings can generalize to ...
313. HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following ​
Author: Liudas Panavas, Sebastian Minus, Bradley Monton, Derek Ray, Suhaas Garre, Sushant Mehta, Edwin Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.25398v3 Announce Type: replace Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let that document govern every action that follows. Existing benchma...
314. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them ​
Author: Yang Zhou, Zixuan Huang, Sunzhu Li, Zhuo Yang, Chen Zhang, Shunian Chen, Caijun Yan, Jianyao Xu, Shunyu Liu, Weijie Fu, Peiliang Li, Xiaozhi Chen, Yuxiang Cai
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: genera...
315. The Geometric Nature and a Free Proxy for Flow-Matching Uncertainty ​
Author: Ziyang Rao, Yiren Zhao, Weiyu Guo, Ben Fei, Yandong Guo, Hui Xiong
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.27933v2 Announce Type: replace Abstract: Flow matching (FM) has become a popular action head paradigm for modern embodied models. However, as a conditional generative model, it does not explicitly expose its inherent uncertainty, producing faulty action chunks even when it misinterprets t...
316. SKILL-KD: Contrastive Skill Distillation for LLM Agents ​
Author: Qiming Shi, Yibo Dou, Jiawen Zhu, Yulong Tao, Linbo Jin, Zhaolu Kang, Yunfan Zhou, Di Weng
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.28048v2 Announce Type: replace Abstract: Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of successful demonstrati...
317. NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability ​
Author: Duo Xu, Faramarz Fekri
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.28942v2 Announce Type: replace Abstract: Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scientific discovery. In these settings, agents must act based on limited observ...
318. MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations ​
Author: Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.28956v2 Announce Type: replace Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposef...
319. Beyond Retrieval: Analytic Memory for Multimodal Agents ​
Author: Zhoujin Tian, Hao Zhang, Yao Tian, Cheng Chen, Yakun Li, Lei Zhang, Xiaofang Zhou
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.29440v2 Announce Type: replace Abstract: Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizing interaction histories thr...
320. Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation ​
Author: William Caban
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00794v3 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims. No formal framework has yet characterized how validity degrades across the stages of these pipeline...
321. A New Theory of Value for Post-AGI Economics ​
Author: Keyun Ruan
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, econ.GN, q-fin.EC
arXiv:2608.01432v2 Announce Type: replace Abstract: Artificial general intelligence (AGI) may weaken scarcities in labour, expertise, information, and productive capability that underpin established theories of economic value. If cognitive work becomes widely automatable, market price, labour input,...
322. Where Reasoning Diverges: Localized Multi-Agent Debate for Multi-Hop Question Answering ​
Author: Weijun Gao, Xiang Ding, Haoyang Liu, Tiancheng Xing
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.01463v2 Announce Type: replace Abstract: Multi-agent debate commonly exchanges complete rationales even when disagreements concern only a few intermediate claims. We introduce Localized Multi-Agent Debate (LMAD), an inference-time protocol that represents agent rationales as nodes, locate...
323. LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing ​
Author: Wen Zan, Jiaqi Zhang, Jianchao Tan, Hong Liu, Cunguang Wang, Xiang Li, Duyue Ma, Guanyu Wu, Yifan Lu, Fengcun Li, Yerui Sun, Peng Pei, Yuchen Xie, Xunliang Cai
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.DC, cs.LG
arXiv:2608.01662v2 Announce Type: replace Abstract: DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive $O(L^2)$ scoring overhead and the hardware-inefficient, discontinuou...
324. When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary ​
Author: Qiuyang Zhan, Rui Zhang, Sheng Guo, Lepeng Zhao, Zhuotao Liu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01679v2 Announce Type: replace Abstract: Persistent memory allows (self-evolving) LLM agents to adapt across tasks by consolidating heterogeneous interaction histories into reusable facts, preferences, observations, and rules. Yet consolidation also imposes an implicit authorization bound...
325. Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs ​
Author: Zixuan Huang, Yang Zhou, Kaixuan Wang, Guli Zhang, Hongyan Xie, Yakun Zhu, Hao Geng, Xiaozhi Chen, Yikun Ban, Deqing Wang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01755v2 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines ...
326. Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG ​
Author: Daeyoung Roh, Donghee Han
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02011v2 Announce Type: replace Abstract: Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them. We study this failure mode as a procedural property of the...
327. SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents ​
Author: Yue Yao, Shengyuan Wang, Xin Chen, Minke Zhang, Jia He, Bingjun Luo, Tom Gedeon
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02356v2 Announce Type: replace Abstract: Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challenge is not merely to retrieve individually relevant skills, but to identify a complete and executable skill com...
328. A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures ​
Author: Serena Curzel, Fabrizio Ferrandi, Leandro Fiorin, Daniele Ielmini, Cristina Silvano, Francesco Conti, Luca Bompani, Luca Benini, Enrico Calore, Sebastiano Fabio Schifano, Cristian Zambelli, Maurizio Palesi, Giuseppe Ascia, Enrico Russo, Valeria Cardellini, Salvatore Filippone, Francesco Lo Presti, Stefania Perri
Published: 8/5/2026, 4:00:00 AM
Categories: cs.AR, cs.AI
arXiv:2311.17815v3 Announce Type: replace-cross Abstract: Given their increasing size and complexity, the need for efficient execution of deep neural networks has become increasingly pressing in the design of heterogeneous High-Performance Computing (HPC) and edge platforms, leading to a wide variet...
329. Mixed-Initiative Human-Robot Teaming under Suboptimality with Online Bayesian Adaptation ​
Author: Manisha Natarajan, Chunyue Xue, Sanne van Waveren, Karen Feigh, Matthew Gombolay
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2403.16178v2 Announce Type: replace-cross Abstract: For effective human-agent teaming, robots and other artificial intelligence (AI) agents must infer their human partner's abilities and behavioral response patterns and adapt accordingly. Most prior works make the unrealistic assumption that o...
330. MambaTS: Improved Selective State Space Models for Long-term Time Series Forecasting ​
Author: Xiuding Cai, Xueyao Wang, Yaoyao Zhu, Yu Yao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2405.16440v2 Announce Type: replace-cross Abstract: In recent years, Transformers have become the de-facto architecture for long-term time series forecasting (LTSF), yet they face challenges associated with the self-attention mechanism, including quadratic complexity and permutation-invariant ...
331. CollaFuse: Collaborative Diffusion Models ​
Author: Simeon Allmendinger, Domenique Zipperling, Lukas Struppek, Niklas K"uhl
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2406.14429v4 Announce Type: replace-cross Abstract: In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic images. However, the application of diffusion models poses numerous challenges, particularly concerning...
332. Efficient unsupervised domain adaptation via self-supervised vision transformer and synergistic cross-domain alignment ​
Author: Ali Abedi, Q. M. Jonathan Wu, Ning Zhang, Farhad Pourpanah
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2407.21311v2 Announce Type: replace-cross Abstract: Unsupervised domain adaptation (UDA) aims to mitigate domain shift, where the distribution of labeled source data differs from that of unlabeled target data. Despite recent advances, existing methods often rely on fine-tuning large backbone m...
333. Patient-centered data science: an integrative framework for evaluating and predicting clinical outcomes in the digital health era ​
Author: Mohsen Amoei, Dan Poenaru
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY
arXiv:2408.02677v2 Announce Type: replace-cross Abstract: This study proposes a novel, integrative framework for patient-centered data science in the digital health era. We developed a multidimensional model that combines traditional clinical data with patient-reported outcomes, social determinants ...
334. Rex: A Family of Reversible Exponential (Stochastic) Runge-Kutta Solvers ​
Author: Zander W. Blasingame, Chen Liu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2502.08834v5 Announce Type: replace-cross Abstract: Deep generative models based on neural differential equations have become state-of-the-art for many generation tasks. These models rely on ODE/SDE solvers that integrate from a prior distribution to the data distribution; in many applications...
335. Automated Visualization Code Synthesis via Multi-Path Reasoning and Feedback-Driven Optimization ​
Author: Wonduk Seo, Daye Kang, Hyunjin An, Taehan Kim, Soohyuk Cho, Seungyong Lee, Minhyeong Yu, Jian Park, Yi Bu, Seunghyun Lee
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL, cs.HC
arXiv:2502.11140v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have become a cornerstone for automated visualization code generation, enabling users to create charts through natural language instructions. Despite improvements from techniques like few-shot prompting and query ...
336. Compound and Parallel Modes of Tropical Convolutional Neural Networks ​
Author: Mingbo Li, Liying Liu, Charles Wiranto, Ye Luo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2504.06881v2 Announce Type: replace-cross Abstract: Convolutional neural networks (CNNs) are foundational to many state-of-the-art computer vision systems, yet their reliance on multiplication-intensive computations poses challenges for deployment on resource-constrained devices. While tropica...
337. When Search Teaches Style: Causal Internalization of Tactical Priors in AlphaZero ​
Author: Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Han Wang, Jie Li, Ru Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2504.14636v3 Announce Type: replace-cross Abstract: AlphaZero is normally evaluated as one agent: a policy-value network fused with Monte Carlo tree search. That fusion hides a causal question. When self-play search is given a useful prior, does the network absorb the induced behavior, or does...
338. Beyond Either-Or Reasoning: Transduction and Induction as Cooperative Problem-Solving Paradigms ​
Author: Janis Zenkner, Tobias Sesterhenn, Christian Bartelt
Published: 8/5/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.LG
arXiv:2505.14744v3 Announce Type: replace-cross Abstract: Traditionally, in Programming-by-example (PBE) the goal is to synthesize a program from a small set of input-output examples. Lately, PBE has gained traction as a few-shot reasoning benchmark, relaxing the requirement to produce a program art...
339. One-Point Contraction: Erasing Representational Separability toward Irreversible Deep Forgetting ​
Author: Jaeheun Jung, Bosung Jung, Suhyun Bae, Donghun Lee
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2507.07754v3 Announce Type: replace-cross Abstract: Machine unlearning is usually evaluated by what the classifier outputs: forget-set accuracy, confidence, membership-inference scores. We show that this is not enough. Across 14 representative unlearning methods on CIFAR-10 and SVHN, a single ...
340. IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning ​
Author: Jaeheun Jung, Jaehyuk Lee, Yeajin Lee, Donghun Lee
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2507.14171v3 Announce Type: replace-cross Abstract: Importance-based structured pruning overwhelmingly relies on filter magnitude. This proxy is fundamentally flawed: due to scale invariance, functionally identical filters can receive arbitrarily different importance scores under rescaling. We...
341. From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model ​
Author: Yeong-Joon Ju, Seong-Whan Lee
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR
arXiv:2508.00955v3 Announce Type: replace-cross Abstract: Adapting generative Multimodal Large Language Models (MLLMs) into universal embedding models typically demands resource-intensive contrastive pre-training, while traditional hard negative mining methods suffer from severe false negative conta...
342. Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages ​
Author: Seraphina Fong, Marco Matassoni, Alessio Brutti
Published: 8/5/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL
arXiv:2508.05149v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks. However, their applicability is still less explored in low-resource settin...
343. Uncovering Spontaneous Physics Representations in In-Context Learning ​
Author: Yeongwoo Song, Jaeyong Bae, Dong-Kyum Kim, Hawoong Jeong
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2508.12448v2 Announce Type: replace-cross Abstract: In-context learning (ICL) lets large language models (LLMs) solve new tasks from prompts alone, across an ever-widening range of domains, yet the mechanisms underlying this ability remain poorly understood. Physical systems offer a controlled...
344. Mechanism of Task-oriented Information Removal in In-context Learning ​
Author: Hakaze Cho, Haolin Yang, Gouki Minegishi, Naoya Inoue
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2509.21012v4 Announce Type: replace-cross Abstract: In-context Learning (ICL) is an emerging few-shot learning paradigm based on modern Language Models (LMs), yet its inner mechanism remains unclear. In this paper, we investigate the mechanism through a novel perspective of information removal...
345. Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain ​
Author: L'eo Boisvert, Abhay Puri, Chandra Kiran Reddy Evuru, Nazanin Sepahvand, Nicolas Chapados, Quentin Cappart, Jason Stanley, Alexandre Lacoste, Krishnamurthy Dj Dvijotham, Alexandre Drouin
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2510.05159v5 Announce Type: replace-cross Abstract: While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces critical security vulnerabilities within the agentic AI supply chain. We show that adversaries can effective...
346. Obfuscation Rules for Detecting and Detoxifying Korean Toxicity ​
Author: Yejin Lee, Su-Hyeon Kim, Hyundong Jin, Dayoung Kim, Yeonsoo Kim, Yo-Sub Han
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.10961v4 Announce Type: replace-cross Abstract: As language models become increasingly deployed in online environments, toxicity detection and detoxification have received growing attention. Existing studies primarily focus on non-obfuscated text, which limits robustness when users intenti...
347. Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation ​
Author: Xiao He, Huangxuan Zhao, Guojia Wan, Jiancheng Pan, Yanxing Liu, Yong Luo, Juhua Liu, Yongchao Xu, Wei Zhou, Dacheng Tao, Bo Du
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.IR, cs.MM
arXiv:2510.12953v5 Announce Type: replace-cross Abstract: Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are adapted to structured adult imaging and underperform in fetal ultrasound, which poses challenges of mul...
348. Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models ​
Author: Sarah Ball, Niki Hasrati, Alexander Robey, Avi Schwarzschild, Frauke Kreuter, Zico Kolter, Andrej Risteski
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.22014v2 Announce Type: replace-cross Abstract: Discrete optimization-based jailbreaking attacks on large language models aim to generate short, nonsensical suffixes that, when appended onto input prompts, elicit disallowed content. Notably, these suffixes are often transferable -- succeed...
349. GraphCliff: Short-Long Range Gating for Modeling Critical Activity Changes Caused by Subtle Molecular Differences ​
Author: Hajung Kim, Jueon Park, Junseok Choe, Seungheun Baek, Hyeon Hwang, Jaewoo Kang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CE, cs.AI
arXiv:2511.03170v3 Announce Type: replace-cross Abstract: The quantitative structure-activity relationship assumes a smooth mapping between molecular structure and biological activity. However, activity cliffs, defined as pairs of structurally similar compounds with large potency differences, break ...
350. Target-Aligned Fusion for Decision-Sequence Learning under Dynamics Shift ​
Author: Guojian Wang, Quinson Hon, Xuyang Chen, Lin Zhao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2511.09173v3 Announce Type: replace-cross Abstract: External trajectories can improve offline decision-sequence learning, but dynamics shift may make some source subsequences inconsistent with the target environment. We study how to fuse such trajectories with limited target data for Decision ...
351. $\pi$-Attention: Online Efficient Sparse Transformers for Long-Context Modeling ​
Author: Pike D. Liu, Chang Liu, Yanxuan Yu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2511.10696v3 Announce Type: replace-cross Abstract: Sparse attention is crucial in long-context Transformers, which restricts each token to a limited neighborhood and thereby reduces the quadratic cost of full self-attention. Local windows capture nearby context effectively, yet they induce a ...
352. STREAM-VAE: Dual-Path Routing for Slow and Fast Dynamics in Vehicle Telemetry Anomaly Detection ​
Author: Kadir-Kaan "Ozer, Ren'e Ebeling, Markus Enzweiler
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2511.15339v3 Announce Type: replace-cross Abstract: Automotive telemetry data exhibits slow drifts and fast spikes, often within the same sequence, making reliable anomaly detection challenging. Standard reconstruction-based methods, including sequence variational autoencoders (VAEs), use a si...
353. Externally Validated Breast Ultrasound Segmentation via Multi-task Learning with BI-RADS-Consistent Morphological Priors ​
Author: Jingru Zhang, Saed Moradi, Ashirbani Saha
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2511.15968v2 Announce Type: replace-cross Abstract: External validation of breast ultrasound segmentation models remains limited because internal train--test splits do not capture domain shifts across imaging systems, acquisition protocols, and patient populations. We introduce a novel multi-t...
354. MIMIC-MJX: Neuromechanical Emulation of Animal Behavior ​
Author: Charles Y. Zhang (Harvard University), Yuanjia Yang (Salk Institute for Biological Studies), Aidan Sirbu (Mila), Elliott T. T. Abe (University of Washington), Emil W"arnberg (Harvard University), Eric J. Leonardis (Salk Institute for Biological Studies), Diego E. Aldarondo (Harvard University), Adam Lee (Harvard University), Aaditya Prasad (Massachusetts Institute of Technology), Jason Foat (Salk Institute for Biological Studies), Kaiwen Bian (Salk Institute for Biological Studies), Joshua Park (Salk Institute for Biological Studies), Rusham Bhatt (Salk Institute for Biological Studies), Vyom N. Patel (Neuromatch), Hutton Saunders (Salk Institute for Biological Studies), Austin O. Barbano (Salk Institute for Biological Studies), Akira Nagamori (Salk Institute for Biological Studies), Ayesha R. Thanawalla (Salk Institute for Biological Studies), Kee Wui Huang (Salk Institute for Biological Studies), Fabian Plum (Imperial College London), Hendrik K. Beck (Imperial College London), Steven W. Flavell (Massachusetts Institute of Technology), David Labonte (Imperial College London), Blake A. Richards (Mila), Bingni W. Brunton (University of Washington), Eiman Azim (Salk Institute for Biological Studies), Bence P. "Olveczky (Harvard University), Talmo D. Pereira (Salk Institute for Biological Studies)
Published: 8/5/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.RO
arXiv:2511.20532v3 Announce Type: replace-cross Abstract: The primary output of the nervous system is movement and behavior. While recent advances have democratized pose tracking during complex behavior, kinematic trajectories alone provide only indirect access to the underlying control processes. H...
355. Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models ​
Author: Yuhang Wang, Yanxu Zhu, Jiaming Zhang, Dongyuan Lu, Jitao Sang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2511.21214v4 Announce Type: replace-cross Abstract: Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We study whether a reasoning model can synthesize and internalize a task-specific safety guideline from a ...
356. PRISMA: Improving the Accuracy-Latency Frontier of Diffusion-based PDE Solvers Using Physics-Informed Spectral Attention ​
Author: Medha Sawhney, Abhilash Neog, Mridul Khurana, Anuj Karpatne
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NA, math.NA, stat.ML
arXiv:2512.01370v2 Announce Type: replace-cross Abstract: Diffusion-based solvers for partial differential equations (PDEs) are often bottle-necked by slow gradient-based test-time optimization routines that use PDE residuals for loss guidance. They additionally suffer from optimization instabilitie...
357. PRIVEE: Privacy-Preserving Vertical Federated Learning Against Feature Inference Attacks ​
Author: Sindhuja Madabushi, Haider Ali, Ahmad Faraz Khan, Rui Ning, Hongyi Wu, Chunsheng Xin, Ali. R. Butt, Jin-Hee Cho
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2512.12840v2 Announce Type: replace-cross Abstract: Vertical Federated Learning (VFL) enables collaborative model training across organizations that share common user samples but hold disjoint feature spaces. Despite its potential, VFL is susceptible to feature inference attacks, in which adve...
358. HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize ​
Author: Kun Zhao, Guodong Liu, Hui Ji, Siyuan Dai, Pan Wang, Jifeng Song, Chenghua Lin, Liang Zhan, Haoteng Tang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.03321v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcement learning (RL) remains challenging due to heterogeneous medical supervision. Vanilla Group Relative...
359. Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models ​
Author: Minh Vu Pham, Hsuvas Borkakoty, Yufang Hou
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.09445v2 Announce Type: replace-cross Abstract: In language models (LMs), intra-memory knowledge conflict arises when inconsistent information about the same subject is encoded within the model's parametric knowledge. Prior work has primarily focused on resolving conflicts between a model'...
360. ChiEngMixBench: Evaluating Large Language Models on Expert-Style Chinese-English Terminology Mixing ​
Author: Qingyan Yang, Tongxi Wang, Yunsheng Luo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.16217v2 Announce Type: replace-cross Abstract: Large language models increasingly mediate multilingual professional communication, where useful generation requires adapting to community conventions about which expressions are retained, translated, or mixed. Existing benchmarks rarely isol...
361. AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection ​
Author: Wachiraphan Charoenwet, Kla Tantithamthavorn, Patanamon Thongtanunam, Hong Yi Lin, Minwoo Jeong, Ming Wu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, cs.SE
arXiv:2601.19138v2 Announce Type: replace-cross Abstract: Secure code review is critical during pre-integration, where Atlassian developers rely on lightweight analysis tools, while deep security assessment is deferred to later stages, delaying feedback and increasing remediation costs. Existing sta...
362. On the Limits of Layer Pruning for Generative Reasoning in Large Language Models ​
Author: Safal Shrestha, Anubhav Shrestha, Minwu Kim, Aadim Nepal, Keith Ross
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.01997v3 Announce Type: replace-cross Abstract: Recent work has shown that layer pruning can effectively compress large language models (LLMs) while retaining strong performance on classification benchmarks, often with little or no finetuning. In contrast, generative reasoning tasks, such ...
363. A Deployment-Friendly Foundational Framework for Efficient Computational Pathology ​
Author: Yu Cai, Cheng Jin, Zhengyu Zhang, Jiabo Ma, Fengtao Zhou, Yingxue Xu, Zhengrui Guo, Yihui Wang, Zhengyu Zhang, Ling Liang, Yonghao Tan, Pingcheng Dong, Du Cai, On Ki Tang, Chenglong Zhao, Zhijian Cen, Ying Tan, Xi Wang, Can Yang, Yali Xu, Jing Cui, Zhenhui Li, Ronald Cheong Kin Chan, Yueping Liu, Feng Gao, Xiuming Zhang, Li Liang, Hao Chen, Kwang-Ting Cheng
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2602.14010v2 Announce Type: replace-cross Abstract: Pathology foundation models (PFMs) generalize well across computational pathology tasks but remain costly for gigapixel whole-slide image analysis. Here, we present LitePath, a deployment-friendly framework that addresses model over-parameter...
364. In-Context Pure Exploration in Continuous Decision Spaces ​
Author: Alessio Russo, Yin-Ching Lee, Ryan Welch, Aldo Pacchiano
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.17976v2 Announce Type: replace-cross Abstract: In active sequential testing, also termed pure exploration, a learner is tasked with the goal to adaptively acquire information so as to identify an unknown ground-truth hypothesis with as few queries as possible. This problem has several mot...
365. SphUnc: Hyperspherical Uncertainty Decomposition and Causal Identification via Information Geometry ​
Author: Rong Fu, Chunlei Meng, Jinshuo Liu, Dianyu Zhao, Yongtai Liu, Yibo Meng, Xiaowen Ma, Wangyu Wu, Yangchen Zeng, Shuaishuai Cao, Simon Fong
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2603.01168v3 Announce Type: replace-cross Abstract: Reliable decision-making in complex multi-agent systems requires calibrated predictions and interpretable uncertainty. We introduce SphUnc, a unified framework combining hyperspherical representation learning with structural causal modeling. ...
366. Quantifying Hallucinations in Language Language Models on Medical Textbooks ​
Author: Brandon C. Colelough, Davis Bartels, Dina Demner-Fushman
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.09986v3 Announce Type: replace-cross Abstract: Hallucinations, the tendency for large language models to provide responses with factually incorrect and unsupported claims, is a serious problem within natural language processing for which we do not yet have an effective solution to mitigat...
367. Large Language Models provide support for the parallelogram theory of analogy ​
Author: Qiawen Ella Liu, Raja Marjieh, Jian-Qiao Zhu, Adele E. Goldberg, Thomas L. Griffiths
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.19066v2 Announce Type: replace-cross Abstract: Four-term word analogies (A:B::C:D) are classically modeled geometrically as parallelograms: adding the vector B-A+C produces D. Recent work suggests that this model poorly captures how humans produce analogies, with simple local-similarity h...
368. The production of meaning in the processing of natural language ​
Author: Christopher J. Agostino, Quan Le Thien, Nayan D'Souza, Louis van der Elst
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2603.20381v2 Announce Type: replace-cross Abstract: Understanding the fundamental mechanisms governing the production of meaning in the processing of natural language is critical for designing safe, thoughtful, engaging, and empowering human-agent interactions. If meaning is constituted rather...
369. DIB-OD: Preserving the Invariant Core for Robust Heterogeneous Graph Adaptation via Decoupled Information Bottleneck and Online Distillation ​
Author: Yang Yan, Yunxuan Li, Qiuyan Wang, Tianjin Huang, Qiudong Yu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.10882v3 Announce Type: replace-cross Abstract: Graph pre-training can facilitate knowledge transfer across graph datasets, but severe structural and feature shifts may cause negative transfer and adaptation-induced overwriting of reusable knowledge. We propose DIB-OD, a heterogeneous grap...
370. Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces ​
Author: Manas Pathak, Xingyao Chen, Shuozhe Li, Amy Zhang, Liu Leqi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.11996v5 Announce Type: replace-cross Abstract: Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone does not reveal the quality of the reasoning used to produce it. This highlights a fundamental limitati...
371. Gated Memory Policy: In-Context Memorization and Adaptation ​
Author: Yihuai Gao, Jeff Jinyun Liu, Shuang Li, Shuran Song
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2604.18933v2 Announce Type: replace-cross Abstract: Robotic manipulation tasks exhibit varying memory requirements, ranging from Markovian tasks that require no memory to non-Markovian tasks that demand in-context memorization of historical information within a single trial or in-context adapt...
372. A neural operator framework for data-driven discovery of stability and receptivity in physical systems ​
Author: Chengyun Wang, Liwei Chen, Nils Thuerey
Published: 8/5/2026, 4:00:00 AM
Categories: physics.flu-dyn, cs.AI
arXiv:2604.19465v3 Announce Type: replace-cross Abstract: Understanding how complex systems respond to perturbations, such as whether they will remain stable or what their most sensitive patterns are, is a fundamental challenge across science and engineering. Traditional stability and receptivity (r...
373. Estimating Tail Risks in Language Model Output Distributions ​
Author: Rico Angell, Raghav Singhal, Zachary Horvitz, Zhou Yu, Rajesh Ranganath, Kathleen McKeown, He He
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.22167v3 Announce Type: replace-cross Abstract: Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increasingly high-stakes. Fortunately, advances in alignment have significantly reduced the likeli...
374. Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations ​
Author: Anna Arnaudo, Riccardo Coppola, Maurizio Morisio, Flavio Giobergia, Andrea Bioddo, Angelo Bongiorno, Luca Dadone
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL
arXiv:2604.22207v2 Announce Type: replace-cross Abstract: Due to the textual and repetitive nature of many Requirements Engineering (RE) artefacts, Large Language Models (LLMs) have proven useful to automate their generation and processing. In this paper, we discuss a possible approach for automatin...
375. TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions ​
Author: Ce Chen, Yi Ren, Yuanming Li, Viktor Goriachko, Zhenhui Ye, Zujin Guo, Zhibin Hong, Mingming Gong
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.27975v2 Announce Type: replace-cross Abstract: Traditional Shot Boundary Detection (SBD) inherently struggles with complex transitions by formulating the task around isolated cut points, frequently yielding corrupted video shots. We address this fundamental limitation by formalizing the S...
376. Injection-Execution Dissociation: A Mechanistic Evaluation of Persistent Memory Attacks and Defenses in Stateful LLM Agents ​
Author: Jun Wen Leong
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2605.08442v5 Announce Type: replace-cross Abstract: We discover that prompt-injection success and tool-execution success are separable safety properties: defenses that block injection do not necessarily block execution, and vice versa. We call this the injection-execution dissociation. In LLM ...
377. Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling ​
Author: Deepak Pandita, Flip Korn, Chris Welty, Christopher M. Homan
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.13801v2 Announce Type: replace-cross Abstract: As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of these systems is paramount. However, AI is currently facing a reproducibility crisis driven by...
378. Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search ​
Author: Jialin Lu, Soonho Kong, Rodrigo Stehling, Kaiyu Yang, Zhangyang Wang, Weiran Sun, Wuyang Chen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.CL, cs.LG, cs.SE
arXiv:2605.20244v2 Announce Type: replace-cross Abstract: We present Lean Refactor, a plug-and-play retrieval-augmented agentic framework for multi-objective, controllable, and version-robust refactoring of Lean proofs. LLM-generated proofs are notoriously correct-but-verbose and brittle across libr...
379. ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models ​
Author: Arash Akbari, Arman Akbari, Masih Eskandar, Qitao Tan, Yixiao Chen, Jingwu Luo, Bertha Pangaribuan, Liyun Zhang, Jennifer Dy, Geng Yuan, Xue Lin, Gaowen Liu, Stratis Ioannidis, Yanzhi Wang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2605.24011v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge platforms impractical. Aggressive, sub-4-bit weight quantization is the natural solution, yet ...
380. E4GEN: Event-level Explainable Extreme-Enhanced Time-series Generation ​
Author: Lin Jiang, Dahai Yu, Ximiao Li, Guang Wang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.01634v2 Announce Type: replace-cross Abstract: Generating realistic time series is essential for scientific research and real-world applications. However, existing methods often emphasize overall distributional fidelity while failing to faithfully capture extreme events. To advance existi...
381. FLARE: Diffusion for Hybrid Language Model ​
Author: Yuchen Zhu, Jing Shi, Chongjian Ge, Hao Tan, Yiran Xu, Wanrong Zhu, Jason Kuen, Koustava Goswami, Rajiv Jain, Yongxin Chen, Molei Tao, Jiuxiang Gu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.01774v2 Announce Type: replace-cross Abstract: Autoregressive (AR) large language models (LLMs) have achieved broad practical success, but sequential decoding remains a key bottleneck for low-latency deployment. Recent efficient-inference work has progressed along two axes: reducing the c...
382. When Behavioral Safety Evaluation Fails: A Representation-Level Perspective ​
Author: Enyi Jiang, Anders Gj{\o}lbye, Yibo Jacky Zhang, Sanmi Koyejo
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2606.08044v2 Announce Type: replace-cross Abstract: Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and answers benign ones. But refusing on the prompts an auditor happens to try does not show that the model is...
383. When Context Returns: Toward Robust Internalization in On-Policy Distillation ​
Author: Xun Wang, Ruishuo Chen, Zhuoran Li, Yu Chen, Longbo Huang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.11627v2 Announce Type: replace-cross Abstract: Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so that the context is no longer needed at inference time. However, we identify a counterintuitiv...
384. CADET: Physics-Grounded Causal Auditing and Training-Free Deconfounding of End-to-End Driving Planners ​
Author: Zikun Guo, Yuanyuan Li, Rongjin Zou
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2606.14438v3 Announce Type: replace-cross Abstract: End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that merely co-occur with expert actions (a roadside object, a building facade) with driving decisions, rather...
385. Diagnosing and Mitigating Context Rot in Long-horizon Search ​
Author: Shijie Xia, Yikun Wang, Zhen Huang, Pengfei Liu
Published: 8/5/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2606.29718v2 Announce Type: replace-cross Abstract: Extensive context has become the norm as Large Language Models (LLMs) are increasingly deployed in long-horizon search tasks. The concern that increasing context length degrades model capabilities, known as context rot, has become a widely re...
386. MalariAI: A Label-Resilient Decoupled Framework for Annotation-Agnostic Cell Segmentation and Explainable Stage Classification in Dense Malaria Blood Smears ​
Author: Kaysarul Anas Apurba, Md Hasibul Hasan, Mohammed Ali, Tanzilur Rahman
Published: 8/5/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV
arXiv:2607.00385v2 Announce Type: replace-cross Abstract: Automated malaria diagnosis from blood smear microscopy is a critical global health AI challenge; expert scarcity remains the primary diagnostic bottleneck. Existing deep learning systems face three compounding failures: end-to-end detectors ...
387. VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment ​
Author: Guoyang Xia, Fengfa Li, Hongjin Ji, Lei Ren, Fangxiang Feng, Kun Zhan, Yan Xie
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO
arXiv:2607.01586v2 Announce Type: replace-cross Abstract: Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training paradigms remain difficult to compare because existing models often differ in architecture, data, action sp...
388. Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks ​
Author: Yoshihiro Maruyama
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2607.03798v3 Announce Type: replace-cross Abstract: Symmetry is everywhere in nature and society. Geometric deep learning builds architectures respecting group symmetries, whereas topological deep learning organizes computation through cells, incidence relations, and local-to-global structure....
389. Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5 ​
Author: Lyndon Drake (University of Oxford), Zandi Eberstadt (University of Oxford)
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.LG
arXiv:2607.04510v2 Announce Type: replace-cross Abstract: Emergent misalignment (EM) --- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data --- is mediated in Qwen2.5 models by a latent persona direction, and that direction is causal in open weights. Transplant...
390. x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability ​
Author: Xin Peng, Ang Gao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.06114v4 Announce Type: replace-cross Abstract: Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function evaluations (NFEs). This remains a practical challenge for released checkpoints, since many accelerators r...
391. Subjective Risk Decomposition: A New View for Uncertainty Quantification ​
Author: Raghad Alamri, Michele Caprio, Gavin Brown
Published: 8/5/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2607.15196v2 Announce Type: replace-cross Abstract: We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axioms and argumentation, but instead consequences, of higher-level modelling decisions. We show how epistemic and aleatoric unce...
392. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning ​
Author: Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.16057v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic too...
393. Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min--Max Selection ​
Author: Yaodi Luo, Peize He, Lingbei Meng, Bowen Han, Zheng Lu, Jianqing Zhu, Lian Zhang
Published: 8/5/2026, 4:00:00 AM
Categories: q-bio.GN, cs.AI, cs.LG
arXiv:2607.19426v2 Announce Type: replace-cross Abstract: Large single-cell datasets are expensive to store, curate, and repeatedly reuse for model training. Data distillation can reduce this burden by building smaller training sets. However, many existing methods rely on synthetic cells. These synt...
394. A Systematic Benchmark of Intensity Normalisation Methods for 3D Knee MRI Segmentation and Cross-Domain Generalisability ​
Author: Oliver Mills, Philip Conaghan, Samuel Relton
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.20028v2 Announce Type: replace-cross Abstract: Robust out-of-the-box performance is essential for the clinical deployment of deep learning models in medical imaging. An important but underexplored factor affecting model generalisability is intensity normalisation, particularly for magneti...
395. TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex ​
Author: Yuliang Yan, Shuo Yan, Haochun Tang, Yiqin Sun, Enyan Dai
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22143v2 Announce Type: replace-cross Abstract: Molecular glue degraders have emerged as a promising strategy for targeted protein degradation by inducing ternary complex formation between an E3 ubiquitin ligase and a target protein. Despite their therapeutic potential, computational desig...
396. CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference ​
Author: Jiyuan Tan, Vasilis Syrgkanis
Published: 8/5/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, econ.EM
arXiv:2607.22511v2 Announce Type: replace-cross Abstract: Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large language model (LLM) reviewer. However, such r...
397. Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features ​
Author: Dzmitry Malyshau
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.22739v2 Announce Type: replace-cross Abstract: We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcement learning or explicit memory. Cortex is a compact Quake policy with 10.98 million trainable paramete...
398. WCM: World-Cognition Model for Generalizable Human-Robot Interaction ​
Author: Yuzhen Chen, KC Zhou
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.HC, cs.LG
arXiv:2607.22999v2 Announce Type: replace-cross Abstract: Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tasks. Current robot-control paradigms, including vision-language-action policies and world-model-based pl...
399. Moral Hazard in Multi-Agent Language Models ​
Author: Dane Malenfant
Published: 8/5/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2607.23982v3 Announce Type: replace-cross Abstract: Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmstr"om's team moral-hazard model, we introduce the Dialogue Moral Hazard Game, a controlled textual game that operati...
400. AgentGUI: An Interface for Observing and Steering Long-Running AI Agents ​
Author: Xuan Zhao, Jiwoong Sohn, Qinyue Zheng, Michael Moor
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2607.26300v2 Announce Type: replace-cross Abstract: AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to limited human-centered interfacing. Aiming to address this, we ...
401. Benchmarking LLM Competence on Logical Inference over Probability Operators ​
Author: Nayera Hasan, Jack Greff, Alvin Grissom II
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.27405v3 Announce Type: replace-cross Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such a...
402. JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles ​
Author: Shawn Li, Wei Yang, Jike Zhong, Jiate Li, Jiawei Yang, You Qin, Ryan Rossi, Franck Dernoncourt, Roger Zimmermann, Yue Wang, Zhengzhong Tu, Vicente Ordonez, Mohit Bansal, Yue Zhao
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.27670v2 Announce Type: replace-cross Abstract: Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a ben...
403. PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks ​
Author: Manyi Wang, Junjielong Xu, Pinjia He
Published: 8/5/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.28587v2 Announce Type: replace-cross Abstract: SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a common construction pipeline: each PR (Pull Request) is paired with its linked issue by extracting issue references from the P...
404. The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models ​
Author: Plawan Kumar Rath
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2607.28639v2 Announce Type: replace-cross Abstract: We show that knowledge distillation (KD) in small instruction-tuned language models has asymmetric effects on bias, and that measuring them correctly requires accounting for where refusal mass moves and what the parser can legitimately score....
405. Automated ECG Interval Measurement and Wave Delineation Using Fast Fourier Convolution ResNet ​
Author: Farhan Adam Mukadam, Harshit Mishra, Nachiket Makwana, Pradyot Tiwari, Subramani Kandasamy, KVS Hari
Published: 8/5/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG
arXiv:2608.00058v2 Announce Type: replace-cross Abstract: Accurate measurement of ECG intervals, including PR, QRS duration, and QT/QTc, is central to cardiac diagnosis, yet the published ECG delineation literature evaluates performance almost exclusively as fiducial-point timing errors on small cur...
406. Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs ​
Author: Vahidin Hasic, Chao Wang, Luis C. Garcia-Peraza-Herrera, David Watson, Senka Krivic
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00076v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. While existing explainability methods identify influential image regions or text tokens, th...
407. LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations ​
Author: Yan Fang, Jialin Chen, Chun Gan, Hang Yu, Mingjun Nie, Yeyu Zhang, Fengxiang He, Ching Law
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.GT, cs.LG
arXiv:2608.00123v2 Announce Type: replace-cross Abstract: LLM-native advertising embeds sponsored content directly into model-generated responses, shifting the unit of sale from a fixed slot to a moment within an evolving conversation. Existing LLM ad-auction mechanisms primarily operate within a si...
408. Optimising for Flourishing: Flourishing Metrics and Return on Flourishing as Success Criteria for Artificial Intelligence and Post-AGI Economic Systems ​
Author: Keyun Ruan, Jonathan D. Teubner, John M. Bremen
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, econ.TH
arXiv:2608.00151v2 Announce Type: replace-cross Abstract: Current evaluation frameworks for artificial intelligence focus mainly on capability, safety, and proxies such as adoption, engagement, efficiency, productivity, and financial return. These criteria are necessary but insufficient because they...
409. When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems ​
Author: Neha Nagaraja, Amisha Bagari, Hayretdin Bahsi
Published: 8/5/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CR, cs.MA
arXiv:2608.00747v2 Announce Type: replace-cross Abstract: Large language models are increasingly integrated into autonomous robotic systems for task planning and control, but this integration exposes them to prompt injection attacks that can lead to unsafe decisions and physical harm. Multi-agent se...
410. ACE-GraphRAG: Agentic Context Engineering for Hierarchical GraphRAG ​
Author: Yongfeng Huang, Yuren Lai, Ruiying Chen, Haoyu Huang, Mingming Zhao, James Cheng
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01269v2 Announce Type: replace-cross Abstract: Hierarchical Graph Retrieval-Augmented Generation (GraphRAG) organizes corpus knowledge at multiple levels of granularity, yet fixed context construction may fail to translate these multi-resolution representations into a context suited to th...
411. Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Metric for Infrared-Visible Fusion Assessment ​
Author: Haoran Liu, Mingzhe Liu, Peng Li, Guibin Zan
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01301v2 Announce Type: replace-cross Abstract: Infrared-visible image fusion (IVIF) has no ideal fused reference, so fusion algorithms are routinely ranked by scalar objective metrics that formalize different proxies for information transfer, structure, or source similarity. These proxies...
412. Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution ​
Author: B. Sankar, Pawni Yadav, Srinidhi Ranjini Girish, Amogh A. S
Published: 8/5/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2608.01366v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts. Iterative multi-turn prompting often leads to context degradation and diminishing cognitive returns. We ...
413. Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation ​
Author: Fengxian Ji, Yuke Li, Jingpu Yang, Juanfan Wu, Fan Zhang, Zhexuan Cui, Yu Xie, Min Peng, Qianqian Xie, Xiuying Chen, Zhuohan Xie
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01666v2 Announce Type: replace-cross Abstract: However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unified three-component ben...
414. LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation ​
Author: Tankun Li, Zhi Chen, Yaohua Tang
Published: 8/5/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01804v2 Announce Type: replace-cross Abstract: Post-training large language models (LLMs) via reinforcement learning (RL) has significantly advanced code generation capabilities. To bypass the heavy memory footprint of critic networks, current state-of-the-art frameworks leverage critic-f...
415. Self-Improving Large Language Models via Progressive Experience Evolution ​
Author: Shijie Ren, Xiting Wang, Meng Li, Yujie Guo, Yunhang Yao, Ziheng Peng, Xunlong Wang, Yuetan Chen, Haoyang Zhou, Yunlong Liang, Fandong Meng
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.02139v2 Announce Type: replace-cross Abstract: Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing self-impro...
416. PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs ​
Author: Zhongjie Ba, Shengwang Xu, Peng Cheng, Jinyang Zou, Ting Yu, Zhibo Wang, Zhan Qin
Published: 8/5/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02150v3 Announce Type: replace-cross Abstract: Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understa...