arXiv cs.AI - 2026-08-12 ​
344 items collected.
1. Closed-Loop LLM Co-Pilots for Digital Agriculture ​
Author: Serge Kernbach
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, physics.bio-ph
arXiv:2608.09949v1 Announce Type: new Abstract: This study evaluates the application of Large Language Models (LLMs) in complex biological systems, evolving from data analysis to autonomous, AI-guided experimentation. The framework is driven by data from a 49-channel phytosensor network, encompassin...
2. SPOTting the Future: Lookahead Explanations for Deep Reinforcement Learning ​
Author: Tamar Gozlan, Claudia V. Goldman
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.09967v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) agents achieve strong performance in complex environments, yet their decision-making processes remain difficult to interpret. We introduce SPOT (Sampling Policy Observation Tree), a novel model-agnostic, sampling-based...
3. MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis ​
Author: Yuhua Wen, Yingying Zhou, Qifei Li, Yingming Gao, Zhengqi Wen, Jianhua Tao, Ya Li
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.09986v1 Announce Type: new Abstract: Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted modalities, posing a critical challenge. Although several methods have bee...
4. Towards Sustainable Artificial Intelligence: A Comprehensive Review and Comparative Analysis of Deep Learning Models' Carbon Footprint ​
Author: Samar Garrab, Sarra Boughriou, Manel BenSassi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG, cs.SE
arXiv:2608.09998v1 Announce Type: new Abstract: Artificial Intelligence (AI) and Machine Learning (ML) have become powerful tools for supporting and automating complex human tasks. Despite their benefits, growing attention has been directed toward their environmental implications, primarily due to t...
5. ReCBM: Uncertainty-Gated Relational Reasoning for Concept Bottleneck Models ​
Author: An Sui, Yuzhu Li, Fuping Wu, Xiahai Zhuang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.10004v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) provide an interpretable framework by grounding predictions in human-understandable concepts, enabling semantic inspection and test-time intervention. Recent variants have improved CBMs through richer concept representa...
6. Automating and Scaling Behavioral Scientific Research on AI Agents ​
Author: Soo Yong Lee, Jongha Lee, Jaewan Chun, Hyunjin Hwang, Fanchen Bu, Ziv Ben-Zion, Taekwan Kim, Denny Borsboom, Jaemin Yoo, Kijung Shin
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.10030v1 Announce Type: new Abstract: As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI agents remains manual and labor-intensive. We introduce AEROBAT, the first multi-agent system to au...
7. CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation ​
Author: Hejia Zhang, Sheng Lu, Zhongming Yu, Chia-Tung Ho, Brucek Khailany, Jishen Zhao
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation alone. Hardware verification is an important application of code generation and accounts for a substan...
8. MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory ​
Author: Beidi Zhao, Yaoqi Chen, Yuru Feng, Menghao Li, Qianxi Zhang, Baotong Lu, Jianan Lu, Zhirui Wang, Xinjiang Wang, Shusen Xu, Zengzhong Li, Xiaoxiao Li, Qi Chen
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10108v1 Announce Type: new Abstract: Long-horizon agents accumulate trajectories spanning hundreds of interleaved reasoning, action, and observation steps, where answering a query may depend on evidence buried far back in the history. External memory stores such trajectories as structured...
9. The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI ​
Author: Srinivas Telukunta, Georgios Nektarios Lilis, Lucio Baron
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10153v1 Announce Type: new Abstract: Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSecOps built for deterministic automation, across every scale of agency. We argue that agentic AI gover...
10. SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents ​
Author: Vivek Kulkarni, Sudipta Paul, Aounon Kumar, Nicholas Tzou, Srinivas Chappidi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10157v1 Announce Type: new Abstract: Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over time. Recently, methods like the Darwin G"odel Machine and the Huxley G"odel Machine have been prop...
11. Generating Attacks for LLMs with GFlowNets ​
Author: Berkay Ozcam, Irem Onen, Mehmet Fatih Amasyali, Emin Islam Tatli
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2608.10171v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespread adoption. However, this escalating trend has introduced significant security vulnerabilities, necessitating t...
12. TRACE: Trustworthy Retrieval-Augmented Conversational Engine ​
Author: Touseef Hasan, Laila Cure, Souvika Sarkar
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10176v1 Announce Type: new Abstract: Public service chatbots are expected to deliver recommendations from an underlying public service directory, while also making sure that the recommendations respect explicit user constraints. In practice, public service directories are noisy and incons...
13. Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model Agents ​
Author: Di Wu, Xiaohui Zhu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10198v1 Announce Type: new Abstract: Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing visual and reasoning states into text. Vision Wormhole realizes this approach by translating visual features into a ...
14. Edge Phoneme Recognition for Children's Speech through Age-Aware Training ​
Author: Matthew Arboleda, Ryan Arboleda, Sophie Haak, Sam Hjelmeset, Andrew Franck, Bingrui Yang, Jose Bustamante Ortiz, Yuanrong Shen, Joel Walsh
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SD
arXiv:2608.10206v1 Announce Type: new Abstract: Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a phoneme detection competition, we found that training a lightweight model to predi...
15. Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop Embedding ​
Author: Xin Dong, Vikash V. Gayah
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10207v1 Announce Type: new Abstract: Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit. Existing reinforcement-learning-based holding controllers primarily rely on instantaneous operational variables or route-specific stop identifiers, whic...
16. Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes ​
Author: Alec Harris, Kasey Corra, Archie Chaudhury, Yixiong Hao
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10209v1 Announce Type: new Abstract: Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives. However, a key limitation of current post-training methods is the inabili...
17. Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models ​
Author: Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10214v1 Announce Type: new Abstract: Do large language models contain domain-specific parametric shells: concentrated, causally necessary neuron populations whose removal selectively degrades a target domain while sparing others? We apply a uniform causal methodology across two domain gra...
18. Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems ​
Author: Vassilis Papadopoulos, McNair Shah, Sam Zimmerman, Jack Lindsey
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.10218v1 Announce Type: new Abstract: AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of mind viruses: ideas or goals that propagate through multi-agent systems ...
19. Self-evolving Agentic Customer Support System at LinkedIn ​
Author: Chih Hui Wang, Mengdie Tu, Qianyun Zhang, Wei Wu, Lili Zhou, Mingqi Shen, Changshuai Wei
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10224v1 Announce Type: new Abstract: Enterprise support agents operate in rapidly changing environments where policies, product capabilities, and knowledge bases evolve continuously, making static assistants brittle and costly to maintain. We present LinkedIn's self-evolving agentic suppo...
20. Beyond Decision Boundaries: Relational Geometry Attacks on Contrastive Embedding Manifolds ​
Author: Fei Zhao, Peiyuan Zhang, Xi Li, Chengcui Zhang, Nitesh Saxena
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.10237v1 Announce Type: new Abstract: Contrastive learning and Siamese embedding models have become the foundation of modern verification systems, where decisions are governed not by discrete classification boundaries, but by relational geometry in embedding space. However, existing advers...
21. Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction ​
Author: Yuqiao Xu, Osama Zafar, Alexander Nemecek, Erman Ayday
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10239v1 Announce Type: new Abstract: Generative AI makes social-engineering attacks more fluent, adaptive, and scalable, increasing the need for LLM-based de- fenders that can protect users during ongoing interactions. We ask whether such defenders identify the structural source of risk o...
22. Interpreting Language Model Hidden States at Scale ​
Author: Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson, Daniel McKenzie, Kyle Chard, Ian Foster
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10260v1 Announce Type: new Abstract: Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network. Trained lenses remain expensive: affine-translator parameters grow quadr...
23. Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability ​
Author: Alvin Spivey, Yu Huang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.LG
arXiv:2608.10300v2 Announce Type: new Abstract: Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human reviewers may each expose rich internal states, while operational exchange requires a narrow shared in...
24. Neuroevolution Arena: Nested Ecological Evaluation of Update-and-Inheritance Regimes across Neural Architectures ​
Author: Yuxu Ge, Yifei Cheng
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2608.10323v1 Announce Type: new Abstract: Competitive artificial-life systems can rank trained controllers differently under training and ecological evaluation. We present Neuroevolution Arena, a GPU-accelerated spatial ecology of independently parameterized neural-network cells, and an audit-...
25. Toward a Theory of Value in AI Alignment ​
Author: Andrew Smart, Shazeda Ahmed, Jackie Kay, Jimmy Tobin, Kris Shrishak, Abeba Birhane
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from toxic speech and hallucinations to AI agents executing unauthorized actions. Within t...
26. Hierarchical Compositionality for An Assistive AI Agent ​
Author: Tianyi Fu, Mohan Sridharan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10330v1 Announce Type: new Abstract: AI agents are increasingly being developed to assist humans in various applications, and Large Language Models and other deep network architectures are considered to be state of the art for such agents. These methods are impressive stochastic predictor...
27. Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research ​
Author: Lin Liao, Peng Li
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10363v1 Announce Type: new Abstract: AI agents can accelerate nutrition research, but their analyses inherit the identity, semantic, and release ambiguities of the underlying data. We present Nutrition Data Service (NDS), source-preserving infrastructure that operationalizes FAIR for auto...
28. DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? ​
Author: Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.10366v1 Announce Type: new Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operati...
29. Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models ​
Author: Jiahui Han, Yuhui Yao, Xin Wang, Jiafei Cao, Mingxuan Zhang, Danfeng Shan, Huiqi Deng, Guanchu Wang, Xia Hu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.10393v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial robustness remains largely underexplored, and exploiting this weakness can lead to physical-world har...
30. Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning ​
Author: Xincong Hu (Nanjing University), Lei Ou (Nanjing University), Maosen Li (Yinwang Intelligent Technology Co., Ltd), Jingtao Zhang (Yinwang Intelligent Technology Co., Ltd), Liguo Hou (Yinwang Intelligent Technology Co., Ltd), Zongzhang Zhang (Nanjing University)
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10403v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world tr...
31. Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects ​
Author: Xin Xu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10420v1 Announce Type: new Abstract: Reasoning shortcuts are solutions of a neurosymbolic system's rules that produce correct predictions through unintended concepts. A recent framework of Takemura, Inoue, and Nishino analyzes them through an automorphism group of value relabelings and as...
32. Recovering Wasted Compute in Autoresearch Agents ​
Author: Au Kwok Chun, Abhigyan Acherjee, Amrutha Rao, Zaiqian Chen, Kazem Meidani, C. Bayan Bruss, Micah Goldblum
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.10424v1 Announce Type: new Abstract: A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. Such agents have inspired large industry investment, motivated by their potential to automate time-consuming human labo...
33. Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance ​
Author: Cong Chi Nguyen, Trang Mai Xuan, Vu-Duc Ngo, Kim-Ngan Thi Nguyen, Trong-Nghia Nguyen, Thien Van Luong
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10434v1 Announce Type: new Abstract: Machine learning-based Intrusion Detection Systems (IDS) have demonstrated superior performance in securing Unmanned Aerial Vehicle (UAV) networks. However, the 'black-box' nature of these models, combined with the high dimensionality of multimodal cyb...
34. Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning ​
Author: Yuhang Cao
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10438v1 Announce Type: new Abstract: Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world. For autoregressive models, tool use naturally fits the generation process: the model emits a tool call...
35. Rationale-Guided Learning for Multimodal Emotion Recognition ​
Author: Sujung Oh, Jung Uk Kim, Sangmin Lee
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10448v1 Announce Type: new Abstract: Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. However, most existing approaches fundamentally treat this as a direct input-output (multimodal cues-emotion labels) m...
36. Quantum Incremental Learning with Mixed State Prototypes ​
Author: Yu Wu, Qianli Zhou, Xinyang Deng, Wen Jiang, Kang Hao Cheong, Witold Pedrycz
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.10464v1 Announce Type: new Abstract: Incremental learning models are required to learn new classes sequentially without catastrophic forgetting, while operating under parameter and memory constraints. In the Noisy Intermediate-Scale Quantum (NISQ) era, although quantum neural networks off...
37. RLMOpt: Adaptive Prompt Optimization via Recursive Language Models ​
Author: Subhash Bangalore Satheesha, Nirvik Pande, Deepthi Duddempudi, Bharath Dandala
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10471v1 Announce Type: new Abstract: Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization procedure: the algorithm determines which candidates to explore and how the search progresses, while the l...
38. Evaluating Rational Contracting in Natural Language ​
Author: Bhavyesh Sajja, Max Kleiman-Weiner, Roger Zimmermann, Tan Zhi-Xuan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.GT
arXiv:2608.10475v1 Announce Type: new Abstract: The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or following hard-coded protocols, such agents can be used to negotiate and execute agreements in open-ended natural ...
39. Multi-Granular Rationale-Guided Molecular LLM for Property Prediction ​
Author: Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.10480v1 Announce Type: new Abstract: Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a molecule through several modalities, notably a 1D SMILES sequence or a 2D molecular...
40. Predicting Space Groups of Double Perovskites by LLM with Dynamic Few-Shot Learning ​
Author: Jongwon Park, Inhyo Lee, Junhyeong Lee, Seunghwa Ryu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci
arXiv:2608.10483v1 Announce Type: new Abstract: Double perovskites (DPs) offer broad compositional tunability, but predicting the space groups (SGs) of stable structures remains difficult because available datasets are often strongly imbalanced toward dominant SG classes. We refer to dominant SG cla...
41. INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators ​
Author: Rose Niousha, Minwoo Kang, Narges Norouzi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.10492v1 Announce Type: new Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring sy...
42. GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning ​
Author: Xin Xiao, Jiang Zhong, Junnan Zhu, Yingchao Feng, Peijin Wang, Yidan Zhang, Kaiwen Wei
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.10494v1 Announce Type: new Abstract: Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challenging because EO workflows are constrained by sensing semantics, product dependencies, spatial and te...
43. From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents ​
Author: Caili Yu, Yiqi Wang, Jiaqi Zhang, Yiqun Duan, Mingkai Zheng, Zhangkai Wu, Kaize Shi, Taotao Cai
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10502v1 Announce Type: new Abstract: Persistent memory lets language-model agents reuse information across sessions, but it also makes errors durable: a poisoned, stale, or misattributed record can alter reasoning, tool use, answers, and subsequent memory writes. Existing defenses mainly ...
44. MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph ​
Author: Jung Hwan Lee, Kyu Ho Lee, Gwang Hoon Yoo
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10504v1 Announce Type: new Abstract: As coding agents increasingly handle implementation, the central challenge shifts from building individual agents to building an infrastructure that systematically improves them. Current approaches optimize agent systems without accumulating transferab...
45. RadFusion: Towards Threshold-Controllable Radiology Report Generation ​
Author: Ying Jin, Noel C. F. Codella, John Corring, Mu Wei, Dinei Florencio, Eric Horvitz
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV
arXiv:2608.10505v1 Announce Type: new Abstract: Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content. S...
46. MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows ​
Author: Yiqi Wang, Zihao Yan, Jiaqi Zhang, Zhangkai Wu, Mingkai Zheng, Zequn Sun, Yanming Zhu, Taotao Cai
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.10509v1 Announce Type: new Abstract: Shared memory helps language-model agents reuse information across long workflows, yet relevant evidence may not be admissible for a particular agent or action. Because restrictions propagate through derivations, summaries can conceal private, poisoned...
47. Measuring Semantic Abstractness of SAE Features via Nonlocality ​
Author: Chuqiao Lin, Shivaji Sondhi, Xiao-Liang Qi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.10537v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanation...
48. SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models ​
Author: Chenhao Dang, Siyuan Xiong, Conghui He, Weijia Li
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10538v1 Announce Type: new Abstract: Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-quality...
49. Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization ​
Author: Khanh Quan Pham, Majid Kundroo, Geunwoo Ban, Seongho Bae, Taehong Kim
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10549v1 Announce Type: new Abstract: Achieving high accuracy in laser-based cutting of optical films requires careful tuning of parameters such as focal length and laser power beam, adjusted according to the specific properties of each film type. Trial-and-error based traditional methods ...
50. DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation ​
Author: Xiaotong Wang, Dazhen Deng
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10567v1 Announce Type: new Abstract: Analytic dashboards combine coordinated views and interactions for data exploration and decision-making. Recent models can generate them from data and natural-language goals, but evaluating their usefulness remains difficult. Dashboard generation is op...
51. Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent ​
Author: Fanqi Zhou, Qiaosheng Chen, Zixian Huang, Gong Cheng
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10579v1 Announce Type: new Abstract: Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real-world datasets makes it impractical for any single metric to generalize across all scenarios. Developers are thus often forced to manu...
52. HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment ​
Author: Xiaokang Qu, Yiting Lin
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.DL
arXiv:2608.10584v2 Announce Type: new Abstract: Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery. Existing scholar assessment methods predominantly rely on bibliometric indicators and reputation proxies, while recent lar...
53. Curate Before You Connect: Identity and Ontology Tagging in a Production Knowledge Graph ​
Author: Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10644v1 Announce Type: new Abstract: Extraction produces candidate entities and relationships; writing them into a graph is where identity is decided, and identity decisions are destructive in a way extraction errors are not. A wrong type can be corrected later, but two records merged und...
54. Decision-Aware Approximation of Belief Functions for Evidential Combinatorial Optimization ​
Author: Sohaib Afifi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, math.OC
arXiv:2608.10650v1 Announce Type: new Abstract: Reducing the number of focal elements of a mass function is classically driven by an intrinsic distance, such as Jaccard or Jousselme, that keeps the approximation close to the original as a body of evidence. We consider instead the case where the mass...
55. Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome ​
Author: Fabrizio Russo, Mark Somers
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10664v1 Announce Type: new Abstract: The Relativity of Causal Knowledge (RCK) explains how a network of agents with different structural causal models can exchange causal knowledge through a shared interventionally consistent abstraction, or backbone. We ask the prior identification quest...
56. VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus ​
Author: Rohit Sinha, Kunal Tilaganji, Tanuja Ganu, Nagarajan Natarajan, Amit Sharma, Vineeth Balasubramanian
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.GT
arXiv:2608.10665v1 Announce Type: new Abstract: Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approaches have notable limitations. Existing approaches either require expensive labelled supervision with i...
57. FITTER: Vocabulary-Agnostic Cross-Domain Inference on Temporal Knowledge Graphs ​
Author: Jiaxin Pan, Mojtaba Nayyeri, Osama Mohammed, Daniel Hernandez, Rongchuan Zhang, Cheng Cheng, Steffen Staab
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10668v1 Announce Type: new Abstract: Temporal knowledge graphs are central to many uses of the Semantic Web, but existing completion methods assume the entities, relation names, and timestamps to be reasoned about are already known at training time, restricting each model to a single grap...
58. REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems ​
Author: Zixing Chen, Xingyuan Liu, Jie Zhu, Huaixia Dou, Shuo Jiang, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10669v1 Announce Type: new Abstract: Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety policies during exec...
59. Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory ​
Author: Aijun Yang, Qianxue Guo, Ziyi Huang, Yuxuan Chen, Shiyou Qian, Jian Cao
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10676v1 Announce Type: new Abstract: Large language model (LLM)-based search agents answer questions through multi-step interactions with external environments. However, providing complete execution trajectories to the LLM causes unbounded context growth and introduces noise. Existing com...
60. Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence ​
Author: Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu, Yiwen Guo
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV
arXiv:2608.10720v1 Announce Type: new Abstract: Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce \textbf{Ex-Omni-2D}, an omni-modal dialogue framework that generates a coordinated response compris...
61. Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution ​
Author: Xun Li, Yiying Yang, Pengtao Li, Xiao Yao, Suyu Liu, Xiaoyang Ye, Ziyu Lu, Yuan Yao, Yangning Li, Yinghui Li, Wenhao Jiang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10740v1 Announce Type: new Abstract: Effective research ideation requires moving beyond a static understanding of prior work to trace how research problems and solutions evolve across the literature. Existing methods either treat papers as unstructured context or model scholarly evolution...
62. Compositional Benchmark Synthesis for Hierarchical Human Action Recognition ​
Author: Farnaz Soleimani (LISSI), Abdelghani Chibani (LISSI), Yacine Amirat (LISSI), Ghazaleh Khodabandelou (LISSI)
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.10765v1 Announce Type: new Abstract: Recognizing human behavior across levels of abstraction, from atomic actions to long-horizon intentions, requires data annotated along a semantic hierarchy. Large corpora provide isolated, atomically labeled clips without temporal composition, whereas ...
63. Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information ​
Author: Kaivalya Rawal, Daria Onitiu, Brent Mittelstadt, Sandra Wachter, Chris Russell
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ML
arXiv:2608.10766v1 Announce Type: new Abstract: Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rule of Thumb'' (RoT) explanations, a new approach to XAI based upon a novel formulation that identifie...
64. SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation ​
Author: Zhou Liu, Ligang Huang, Zeli Su, Zewei Pan, Zhaoyang Han, Xing Chen, Yuanfeng Song, Wentao Zhang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10775v1 Announce Type: new Abstract: Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize individual controls without identifying which familiar workflow is active, which control matters next, or what evid...
65. ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent Experimentation ​
Author: Jiangjie Qiu, Yijun Li, Xiaonan Wang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.10792v1 Announce Type: new Abstract: Autonomous chemistry increasingly depends on environments in which agents can repeatedly act, observe, and adapt.Physical laboratories provide essential real-material evidence but are costly to repeat and difficult to use for tightly matched interventi...
66. EvoMem: Memory-Augmented Evolution for Code Optimization ​
Author: Viktor Volkov, Valentin Khrulkov, Andrey V. Galichin, Danil Sivtsov, Nikita Glazkov, Olga Volkova, Konstantin Pchelin, Iaroslav Bespalov, Dmitry V. Dylov, Petr Anokhin, Ivan Oseledets
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2608.10795v1 Announce Type: new Abstract: Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some cases may transfer across related tasks and domains. However, existing LLM-driven evolutionary frameworks largely ...
67. Hypothesis Frontier: Verifier Guided LLM and Symbolic Search for First-Order Induction ​
Author: Serafim Batzoglou
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10843v1 Announce Type: new Abstract: First-order concept synthesis asks a system to infer one formula that classifies labeled objects consistently across several finite relational structures. Every candidate can be evaluated exactly, but quantified first-order formulas form a vast search ...
68. Enhanced Filtering Algorithms for the Euclidean Traveling Salesperson Problem and its variants in Constraint Logic Programming ​
Author: Alessandro Bertagnon, Marco Gavanelli
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LO
arXiv:2608.10881v1 Announce Type: new Abstract: The Traveling Salesperson Problem (TSP) is one of the best-known problems in computer science and arises in many engineering applications, such as smart vehicles and intelligent transportation systems. In the "Euclidean" case, each node is defined by i...
69. ComBodied Agents: a New Paradigm of Human-Centric Agentic AI ​
Author: Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Yao, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10915v2 Announce Type: new Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what s...
70. IO Factory: Simulating AI-Enabled Influence Campaigns at Scale ​
Author: Lukasz Olejnik, Wenchao Dong, Jonas R. Kunst, Signe Riemer-S{\o}rensen, Tobias Herb, Meeyoung Cha, Daniel Thilo Schroeder
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10920v1 Announce Type: new Abstract: We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat of digital manipulation now extends beyond persuasive text from individual language models to AI swa...
71. ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling ​
Author: Vaibhav Singh, Soumya Suvra Ghosal, Sarvesh Gharat, Soumyabrata Pal, Ramasuri Narayanam, Dinesh Manocha
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10928v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by allocating additional inference-time compute to generate extended chain-of-thought reasoning. However, recent studies reveal that sequential test-time scaling often yields diminishing or even negativ...
72. FedCGR: Federated Cross-Domain Generative Recommendation ​
Author: Zhuodong Liu, Hugen Lv, Xiangyu Li, Bohan Guo, Peiyu Hu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10929v1 Announce Type: new Abstract: Cross-domain recommendation (CDR) transfers preference knowledge across related domains, but federated deployment makes cross-domain alignment difficult because the behavioral anchors that align item spaces, such as overlapping users and shared interac...
73. XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving ​
Author: Foundation Model Team, XPeng Inc
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10976v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verbose natural-language Chain-of-Thought (CoT) is poorly suited to real-time control because it is open...
74. V-FiLLM: Verified Financial LLM Reasoning Benchmark ​
Author: Alicia Larsen, Victoire Laurent, Aulia Kharis Rakhamsari, Lara Turgut, Nino Antulov-Fantulin
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.LG
arXiv:2608.11047v1 Announce Type: new Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored. We introduce V-FiLLM, a framework that generates financial reasoning benchmar...
75. SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure ​
Author: Xiaofan Bai, Hongqiang Lin, Chao Liu, Yantao Zhang, Xuan Jin, Xipeng Cao, Yuhong Li
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11079v1 Announce Type: new Abstract: Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are copied rather than re...
76. RTSKG: Building a Rail Transit Station Knowledge Graph Dataset ​
Author: Shutong Zhu, Tianxing Wu, Runfeng Liu, Yuang Gu, Xuan He, Yuan Zhu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11080v1 Announce Type: new Abstract: Rail transit systems play a vital role in urban mobility and economic development. As key components of such systems, rail transit stations function as critical transport hubs that enhance urban accessibility and stimulate development in surrounding ar...
77. Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding ​
Author: Kushal Chakrabarti
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SE
arXiv:2608.11095v1 Announce Type: new Abstract: Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. We trace this to imperfect recall: appending an instruction is always cheap, but once an in...
78. sLTN: Structural Logic Tensor Networks ​
Author: Davide Rinaldi, Luciano Serafini
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11136v1 Announce Type: new Abstract: Logic Tensor Networks (LTN) provide a neurosymbolic framework in which first-order logic is interpreted through tensor operations, enabling logical constraints to be integrated with differentiable learning. However, the original formulation of LTN is p...
79. Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration ​
Author: Alan Li, Rahul Saha, Anton Xue, Swarat Chaudhuri, Adam Klivans, Pravesh K Kothari, Raghu Meka
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CC, cs.HC, math.FA
arXiv:2608.11195v1 Announce Type: new Abstract: AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures the h...
80. The Gaussian-Multinoulli Restricted Boltzmann Machine: A Potts Model Extension of the GRBM ​
Author: Nikhil Kapasi, Mohamed Elfouly, William Whitehead, Luke Theogarajan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2505.11635v2 Announce Type: cross Abstract: Many real-world tasks, from associative memory to symbolic reasoning, benefit from discrete, structured representations that standard continuous latent models can struggle to express. We introduce the Gaussian-Multinoulli Restricted Boltzmann Machine...
81. "YES! YES! I absolutely love this insight!" Affirmative Narration as Interactional Strategy in Dialogues with LLM Chatbots ​
Author: Hanna-Riikka Roine, Anne Sigrid Refsum, Jill Walker Rettberg
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2607.28646v2 Announce Type: cross Abstract: This article analyses narrative mechanisms that are common in dialogues with LLM chatbots. In combination, these mechanisms produce an interactional strategy for maximising user engagement, which we call affirmative narration. Affirmative narration s...
82. LLM Agents Factory: Retrieval of Domain-Specific LLM Agents ​
Author: Vitalii Belov, Artyom Sosedka, Andrey Sakhovskiy, Elizaveta Kovtun, Artyom Boyarskikh, Semen Budennyy
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the on-the-fly agent de...
83. How to Dogfood Your AI Chat Agent: A Three-Layer Evaluation Framework with Goal-Directed NPC Simulation ​
Author: Alexandre Cristov~ao Maiorano
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.09939v1 Announce Type: cross Abstract: Production teams deploying LLM chat agents face a specific quality assurance gap: existing evaluation tools test individual responses or simulate social interactions, but none systematically verify whether real users can achieve their goals through m...
84. When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning ​
Author: Tughanbulut Kurtulush
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds only asymptotically (...
85. Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents ​
Author: Santosh Patapati
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure behind visual layout. Recent UI agents can act on such interfaces; however, for assistive agents to be...
86. HoosierHelp: Benchmarking LLM Agents for Social Service Navigation ​
Author: Yiyang Li, Weixiang Sun, Tianyi Ma, Kaiwen Shi, Zheyuan Zhang, Yanfang Ye
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.09946v1 Announce Type: cross Abstract: Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising interface for conversational resource navigation, existing benchmarks do not ...
87. Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems ​
Author: Thales Sales Almeida, Giovana Kerche Bon'as, Thiago Laitz, Jo~ao Guilherme Alves Santos, Hugo Abonizio, Roseval Malaquias Junior, Marcos Piau, Celio Larcher, Ramon Pires, Rodrigo Nogueira
Published: 8/12/2026, 4:00:00 AM
Categories: cs.DL, cs.AI
arXiv:2608.09964v1 Announce Type: cross Abstract: The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Brazilian Computer Society since 2012 and publishing work from institutions across the country. Across e...
88. Evidence-Based Scientific Question Discovery: A Framework with Historical Backtesting ​
Author: Hui Mao
Published: 8/12/2026, 4:00:00 AM
Categories: cs.DL, cs.AI
arXiv:2608.09968v1 Announce Type: cross Abstract: Current AI systems are optimized for answering questions; the scientific enterprise is bottlenecked earlier, at discovering the questions worth investigating. We present a framework that turns a traceable, reproducible, scope controlled research corp...
89. Rescene: band-limited stochastic forcing turns a frozen neural weather operator into a climate emulator ​
Author: Minjong Cheon
Published: 8/12/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.AI, cs.CV
arXiv:2608.09971v1 Announce Type: cross Abstract: Over the past few years, the rapid development of machine learning (ML) models for weather forecasting has produced deterministic models whose medium-range skill matches or exceeds that of the European Centre for Medium-Range Weather Forecasts (ECMWF...
90. Do AI weather models miss extremes? ​
Author: Marvin Vincent Gabler, Roberto Molinaro, Niall Siegenheim, Henry Martin, Mark Frey, Niels Poulsen, Philipp Seitz, Olivier Lam
Published: 8/12/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.AI, cs.LG
arXiv:2608.09972v1 Announce Type: cross Abstract: First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression systems. We verify eleven physical and AI forecast systems against European synoptic, solar, and rai...
91. Knowledge-Guided 3D CT Generation: A Conditioning-Centric Taxonomy ​
Author: Francesca Pia Panaccione, Eugenio Lomurno, Matteo Matteucci
Published: 8/12/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG
arXiv:2608.09992v1 Announce Type: cross Abstract: Controllable generation guided by external knowledge is a key requirement in modern generative deep learning applications, enabling the synthesis of samples with explicit constraints on semantic content, structural properties, and variability. In 3D ...
92. Energy and Performance Benchmarking of Deep Learning Models for Breast Cancer Detection ​
Author: Samar Garrab, Ghada Achour
Published: 8/12/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG
arXiv:2608.09996v1 Announce Type: cross Abstract: Recent advances in machine learning have greatly improved breast cancer detection, enabling more accurate and timely diagnosis. Deep learning (DL) models show strong potential for medical image analysis; however, as their architectural complexity inc...
93. Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification ​
Author: M. Sajid, A. Quadir, A. Rahaman, P. N. Suganthan, M. Tanveer
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10007v1 Announce Type: cross Abstract: The current state-of-the-art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep RVFL (edRVFL), treat all training samples uniformly, which limits their robustness and effectiveness when applie...
94. Sheaf-Based Federated Representation Learning ​
Author: Gabriele D'Acunto, Enrico Grimaldi, Valeria Avino, Mario Edoardo Pandolfo, Leonardo Di Nino, Sergio Barbarossa, Paolo Di Lorenzo
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA, eess.SP
arXiv:2608.10016v1 Announce Type: cross Abstract: Heterogeneous federated systems require agents to learn and exchange informative representations despite differences in data distributions, sensing modalities, model architectures, latent dimensionalities, and local learning objectives. To address th...
95. DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents ​
Author: You Lu, Kun Zhang, Bihuan Chen, Xin Peng
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10037v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. Existing studies mainly focus on improving the tool-use capabilities of L...
96. UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs ​
Author: Xuexiong Yin, Zechuan Chen, Yongsen Zheng, Yuxiang Zhang, Jingyuan Yang, Bin Wang, Yubin Wang, Keze Wang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10042v1 Announce Type: cross Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark for personalized dec...
97. Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons ​
Author: Kaustubh Shivshankar Shejole, Tanish Agarwal, Arpit Agarwal, Avishek Ghosh
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10045v1 Announce Type: cross Abstract: The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice, and more recently, fine-tuning large language models. In this problem, the goal is to learn item rewards base...
98. Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review ​
Author: Christopher Braun, Julian Raible, Marco F. Huber
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10047v1 Announce Type: cross Abstract: In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and Health Management (PHM). Machine Learning (ML) has largely driven advancements in diagnostics and prognostics, yet purely data-driven models face inhe...
99. Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds ​
Author: Shiting Gong, Jianpeng Yao, Jinfeng Wang, Marco Pavone, Jiachen Li
Published: 8/12/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG, cs.SY, eess.SY
arXiv:2608.10056v1 Announce Type: cross Abstract: Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely among surrounding pedestrians and obstacles. This conflict becomes more severe in dense scenarios, where aggressi...
100. Status Association Does Not Reliably Predict Decision Leakage ​
Author: Abdullah X
Published: 8/12/2026, 4:00:00 AM
Categories: stat.AP, cs.AI, cs.CY, cs.LG
arXiv:2608.10089v1 Announce Type: cross Abstract: Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled s...
101. Exploring Semantic Stability Across Reviews in the Linux Kernel ​
Author: Lucas Ciziks, Paulo Meirelles, Marco Aur'elio Gerosa
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.SY, eess.SY
arXiv:2608.10101v1 Announce Type: cross Abstract: Code review is credited with substantially changing a patch's code between its first submission and the version that eventually lands. However, prior work typically studied only the final merged patch without comparing it to the first submission. We ...
102. Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4 ​
Author: Matt J. Borowski, Blazej Osinski
Published: 8/12/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2608.10103v1 Announce Type: cross Abstract: High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.async, warp-level matrix loads with ldmatrix, and matrix multiply-accumulate operations with mma.sync. However, most application code ...
103. Procedural Fairness Failures in RLHF from Preference Averaging ​
Author: M P V S Gopinadh, Karthik Kamuju, Kummari Avinash, John Joshua, Srinivasa Raju Rudraraju
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.10126v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. When preferences are heterogeneous, this aggregation induces a procedural fairness failure where major...
104. Multimodal Item Parameter Estimation using Simulated Response Probabilitie ​
Author: Christopher Ormerod, YoungKoung Kim
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10154v1 Announce Type: cross Abstract: We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to replicate choice pro...
105. MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation ​
Author: Jie Cao, Qi Li, Zelin Zhang, Xiaodong Wu, Lingshuang Liu, Xiangman Li, Jianbing Ni
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.10166v1 Announce Type: cross Abstract: Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustness against realistic, model-agnostic removal attacks remains poorly explored. Existing attacks either succeed o...
106. From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation ​
Author: Changshuai Wei, John Bencina, Phuc Nguyen, Andre Assuncao Silva T Ribeiro, Benjamin Zelditch
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10182v1 Announce Type: cross Abstract: Large-scale targeting and recommendation systems are typically built around predictive scores fed into heuristic or local allocation. When the business goal is incremental impact, as in marketing campaigns, incentives, and notifications, this paradig...
107. The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse ​
Author: Maurice Flechtner
Published: 8/12/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CY
arXiv:2608.10186v1 Announce Type: cross Abstract: LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of the...
108. ELMER: Evolutionary Language Model that Explores and Refines ​
Author: Matthew Siper, Ahmed Khalifa, Julian Togelius
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is an unreliable proxy: a small code change can alter nearly every action, while a larger rewrite can pr...
109. Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems ​
Author: Scott E. Frias
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10216v1 Announce Type: cross Abstract: Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication filters, semantic caches, drift guards, and answer grader gates deploy to answer the question: "Does this text sti...
110. FACT: Failure-Aware Causal Training for World-Action Models ​
Author: Quanquan Peng, Yutong Liang, Rui Yan, Nicklas Hansen, Xiaolong Wang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2608.10232v1 Announce Type: cross Abstract: Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation. Building on the future-prediction ability of video models, many WAMs generate future videos and recover actions...
111. Unsupervised Detection of Groundwater Storage Anomalies in Ghana Using GRACE Satellite Data ​
Author: George Yamoah Afrifa, Theophilus Ansah-Narh, Marcellin Atemkeng
Published: 8/12/2026, 4:00:00 AM
Categories: cs.ET, cs.AI, physics.app-ph, physics.geo-ph
arXiv:2608.10233v1 Announce Type: cross Abstract: Groundwater variability in Ghana remains poorly characterized due to limited long-term in-situ observations. This study investigates groundwater storage anomalies using GRACE-derived data from 2004-2024 combined with statistical analysis and unsuperv...
112. TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent ​
Author: Waleed Jamil, Raphael Schmitt
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10258v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after explicit self-trea...
113. Toward Human Rights Benchmarking for LLMs: A Pilot Methodology ​
Author: Savannah Thais, Wm. Matthew Kennedy, Abhigyan Acherjee, Matilda Wysocki, Malcolm Langford, Caitlin Kraft Buchman
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY
arXiv:2608.10268v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this end, we report our ef...
114. Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies ​
Author: Qingfeng Zhang, Yuanxiong Guo, Yanmin Gong
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10273v1 Announce Type: cross Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic evaluation of fine-tuning st...
115. Withholding the Completing Chunk: Deterministic Pair-Completion Guardrails for Streaming LLM Output ​
Author: Christopher M. Frost
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.10279v1 Announce Type: cross Abstract: Streaming language-model output creates a release-timing problem: complete-response moderation acts after streamed text has escaped, whereas repeated semantic classification of partial text can be costly and unstable. We study a narrow deterministic ...
116. Comprendia: AI-Augmented Code Comprehension ​
Author: Costain Nachuma, Minhaz F. Zibran
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.HC, cs.PL
arXiv:2608.10290v1 Announce Type: cross Abstract: Comprendia is an Eclipse plugin that integrates structural dependency visualization with LLM-powered code explanation on a shared interactive graph for Java program comprehension. The tool rests on four pillars: (1) a multi-edge-type dependency graph...
117. MRIComp4Flow: Compression of 3D Brain MRI for Training Multi-Modal Generative Models ​
Author: Lisa K. Fischer, Mykhailo Riabets, Daniel Rueckert, Benedikt Wiestler, Anke Meyer-Baese, Sandeep Nagar
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.10291v1 Announce Type: cross Abstract: Large-scale multi-modal MRI datasets impose substantial storage and I/O costs, limiting the training of 3D generative models on commodity infrastructure. While lossy compression is known to preserve accuracy for discriminative segmentation networks, ...
118. Frozen Brain-MRI Foundation Models Are Site Fingerprints ​
Author: Saman Rahbar
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10295v1 Announce Type: cross Abstract: Frozen foundation-model (FM) embeddings are increasingly used as off-the-shelf brain-MRI representations, on the assumption that they capture anatomy. We audit what they actually encode and find that acquisition site is a large, intrinsic component o...
119. Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility ​
Author: Siyang Wu, Yibo Jiang, Bryon Aragam
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10315v1 Announce Type: cross Abstract: Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. We identify cross-contextual consistency as an underutilized behavioral...
120. Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories ​
Author: Shuyan Huang, Kai Du, Andrew Lan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.10319v1 Announce Type: cross Abstract: Large language model (LLM)-powered agents have rapidly evolved from code-completion tools into solvers of complex software engineering tasks. As developers collaborate with coding agents over time, their preferences emerge through repeated interactio...
121. Narrative Keyframing for Generative Creative Writing ​
Author: Chao Zhang, Abe Davis
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL
arXiv:2608.10337v1 Announce Type: cross Abstract: We introduce narrative keyframing, an interaction technique for AI-assisted creative writing that lets writers specify different types of narrative constraints at selected moments in a story, then use AI to generate intervening prose. Inspired by the...
122. Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement ​
Author: Patrick Vossler, Jialin Ouyang, F. Richard Guo, Anran Huang, Ali Shojaie, Lucas Zier, Fan Xia, Jean Feng
Published: 8/12/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, stat.AP
arXiv:2608.10339v1 Announce Type: cross Abstract: Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to estimate and rank the causal effects of such interventions. This work focuses on one of the most st...
123. Towards Unified Dynamic Face Landmark Detection ​
Author: Sebastian Regalado, Varshanth R. Rao, Ruowei Jiang, Parham Aarabi, Igor Gilitschenski
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10346v1 Announce Type: cross Abstract: Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional limitations: (1) different network parameters need to be trained independently for each ``$N$-point'' benchmark...
124. Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks ​
Author: Zelei Cheng, Amritansh Mishra, Sambit Sahu, William Campbell
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10357v1 Announce Type: cross Abstract: Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcement learning (RL) is a natural fit for this setting, but multi-turn on-policy rollouts create long con...
125. MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model ​
Author: Jiaxin Du, Boulbaba Abdeljaouad, Yong Zhuang, Haoyu Li
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, eess.AS
arXiv:2608.10360v1 Announce Type: cross Abstract: Arabic maqam music microtonal, modal, and built on ornamented call and response is among the traditions most underserved by generative music models, whose training frameworks remain predominantly Western and equaltempered. Real time accompaniment sha...
126. MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices ​
Author: Eunjeong Kim, Yeong Jun Jeon, Myeonggyun Han
Published: 8/12/2026, 4:00:00 AM
Categories: cs.OS, cs.AI
arXiv:2608.10362v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to speculate multiple tokens, reducing expensive target model decoding steps. Its effectiveness depends heavily on draft selection...
127. Beyond Forecasting: Recasting Volatility Control as a Routing Problem ​
Author: Hongji Pu, Leyang Zhou
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CE, cs.AI
arXiv:2608.10375v1 Announce Type: cross Abstract: Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a pre-defined control rule that may not adapt to changing market conditions. We propose VolRouter, a modular fra...
128. A Single Atom in Front of a Mirror is a Universal Reservoir Computer ​
Author: Peter J. Ehlers, Phi Hung Nguyen, Kanu Sinha, Noelle Daigle, Travis W. Sawyer, Hendra I. Nurdin, Daniel Soh
Published: 8/12/2026, 4:00:00 AM
Categories: quant-ph, cs.AI
arXiv:2608.10382v1 Announce Type: cross Abstract: Universal approximation in reservoir computing is typically associated with a class of reservoirs. We show that universality can be associated with a single reservoir, considering a minimal setup of a single atom in front of a mirror. In its linear-t...
129. Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation ​
Author: Samaneh Mohtadi, Pietro Bernardelle, Joel Mackenzie, Gianluca Demartini
Published: 8/12/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.10385v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison. We study persona condition...
130. ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation ​
Author: Ge Wang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10398v1 Announce Type: cross Abstract: Variational autoencoders generate samples from probabilistic latent representations but do not distinguish uncertainty about the latent location from variability around it. We formulate ELVAE, an evidential learning-based VAE in which each latent coo...
131. Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models ​
Author: Shuozhe Cheng, Kunlan Xiang, Mingxuan Li, Ji Zhang, Dongxiao Liu, Wenbo Jiang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.10405v1 Announce Type: cross Abstract: Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in significant computational overhead and resource consumption. While most existing denial-of-service (DoS) ...
132. Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry ​
Author: Liangchen Ge
Published: 8/12/2026, 4:00:00 AM
Categories: cs.DS, cs.AI, cs.CL, cs.LG
arXiv:2608.10416v1 Announce Type: cross Abstract: We present a theoretical foundation for inverse-distance attention, from its Euclidean prototype (Resolver) to its non-Euclidean realization (Riemann GeoResolver). The Euclidean part establishes three core theorems: (1) circuit separation---IDA achie...
133. Causality Sum Rules in Conventional Scattering Matrices ​
Author: Ning Han, Rui Zhao, Shuxing Yang, Mingzhu Li, Hongsheng Chen, Yihao Yang
Published: 8/12/2026, 4:00:00 AM
Categories: physics.optics, cs.AI
arXiv:2608.10427v1 Announce Type: cross Abstract: Scattering matrices are the standard experimental and computational description of photonic and electromagnetic devices. Passivity is explicit in the conventional incoming-outgoing matrix, whereas causality sum rules are usually formulated only after...
134. Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique ​
Author: Sanidhya Vijayvargiya, Rahul Lokesh
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions to force a resolution rather than expressing uncertainty. Existing detection methods fail to provid...
135. What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research ​
Author: Wesley Hanwen Deng, Agathe Balayn, Andrew Selbst, Jason I. Hong, Motahhare Eslami, Kenneth Holstein, Hanna Wallach, Jennifer Wortman Vaughan, Solon Barocas
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.10431v1 Announce Type: cross Abstract: Responsible AI (RAI) has become a central concern for technology companies, regulators, and the public. How industry practitioners interpret, implement, and sustain RAI work directly shapes the design and deployment of AI systems. As empirical schola...
136. FUSE: Frame-Unified Stress Estimation from Facial Video ​
Author: Stefanos Gkikas, Thomas Kassiotis, Yang Guo, Guangliang Li, Giorgos Giannakakis
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10442v1 Announce Type: cross Abstract: Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approaches commonly decompose full recordings into short temporal windows before classification. This design introduces ...
137. From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models ​
Author: Si'an Xie, Jiaxun Liu, Biao Yang, Wei Yuan, Fan Yang, Tingting Gao, Ming Wu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10444v2 Announce Type: cross Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively unexamined capability is...
138. Towards Efficient Reasoning in LLM-Based Recommender Systems via Model Merging ​
Author: Linh Dieu Le, Tong Chen, Shazia Sadiq, Hongzhi Yin, Ming Jin, Junliang Yu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.10447v1 Announce Type: cross Abstract: Large language model-based recommender systems are increasingly adopting slow-thinking models that generate step-by-step reasoning before making predictions, often achieving higher accuracy than fast-thinking models that predict directly. However, th...
139. Persistent Recursive Worlds Enable Autonomous Software Evolution ​
Author: Beichen Huang, Zhenyu Liang, Bowen Zheng, Ran Cheng
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.MA, cs.NE
arXiv:2608.10450v2 Announce Type: cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity through persistent sessions, memories, managers or shared context. We introduce EvoX Genesis (...
140. MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection ​
Author: Jinmo Han, Jimin Hong, Chanyeong Moon, Ju Yeon Kang, Seonuk Kim, Nam Soo Kim
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10459v1 Announce Type: cross Abstract: As LLM-generated content becomes more sophisticated, detection systems for distinguishing those texts from human-written text must operate at scale while handling diverse writing styles, domains, languages, and generator models. Input-only encoder de...
141. Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning ​
Author: Daoyi Li, Yixian Zhang, Chao Yu, Wenbo Ding, Yu Wang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10473v1 Announce Type: cross Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction. However, directly reusing an offline-trained critic can hinder online fine-tuning: as the policy a...
142. Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models ​
Author: Li Wenjie, Yash Jangir, Ignacy Stepka, Yash Agarwal, Marion Kipsang, Yonatan Bisk
Published: 8/12/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CL
arXiv:2608.10484v1 Announce Type: cross Abstract: Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are typically optimized for reconstruction under L1/L2 losses in raw action ...
143. Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation ​
Author: Md Rafid Islam, Rafsan Jany, Zahid Hasan, Ratun Rahman
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10499v1 Announce Type: cross Abstract: Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's policy. Many current meth...
144. SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning ​
Author: Caoyuan Ma, Wenpu Liu, Weichu Xie, Tian Gu, Shilei Zhao, Lingxi Min, Shuai Dong, Yuqi Xu, Ji Zhao, Ziyue Wang, Wenzheng Chang, Taiqiang Wu, Yongfu Zhu, Wenqi Shao, Yinqiang Zheng
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10513v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs throug...
145. Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training ​
Author: Yingsheng Liu, Haiming Li, Jingmin Zhu, Jiajun Sun, Victoria Mar, Monika Janda, H. Peter Soyer, Zongyuan Ge, Zhen Yu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10522v1 Announce Type: cross Abstract: While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent in structured clinical tables. However, existing multimodal pre-training methods underutilize this p...
146. Improving TensorSketch Using Complex Random Variables ​
Author: Amit Sharma, Mohammad Azhar Khan, Rameshwar Pratap, Keegan Kang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.DS, cs.AI, stat.ML
arXiv:2608.10523v1 Announce Type: cross Abstract: \texttt{TensorSketch} by~\cite{pham2013fast,kar2012random} provides efficient sketching algorithms for high-dimensional polynomial kernels $\vec{x}^{\otimes p} \in \R^{d^p}$. \cite{kar2012random} uses dense Johnson-Lindenstrauss (JL)-type projections...
147. Rethinking Text-Based Image Retrieval in Specific Domain ​
Author: Jingyang Tan, Sheng Yang, Yuanpeng Chen, Jian Wang, Nianjin Ye, Chen Xing, Lanpeng Jia
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10524v1 Announce Type: cross Abstract: Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. However, existing benchmarks are predominantly constructed on an exclusive single-match assumption between query ...
148. Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models ​
Author: Yuhang Song, Bor-Jiun Lin, Jiaxu Liu, Te-Chuan Chiu, Anh Nguyen, Chun-Yi Lee
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10525v1 Announce Type: cross Abstract: Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Current VLMs process visual inputs independently, which creates critical limitations for downstream applications th...
149. Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits ​
Author: Ricardo Parada, Chenzhang Zhao, William Chang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10526v1 Announce Type: cross Abstract: Motivated by decentralized applications, we study cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschitz constant is unknown. We consider three information structures: (A)~unobserved actions with common rewards, (B...
150. Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry ​
Author: Daphne Feng, Ricardo Parada, Lily Jiang, Sophia Yi, William Chang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10529v1 Announce Type: cross Abstract: The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumptions. However, real-world applications often involve heavy-tailed reward distributions and decentralized, inform...
151. On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models ​
Author: Md Jafrin Hossain, Mohammad Arif Hossain, Nirwan Ansari
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.10530v1 Announce Type: cross Abstract: Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step planning, tool invocation, code execution, and maintaining persistent memory. When these agents operate with real-...
152. Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration ​
Author: Sangwoo Jo, Donggeun Ko, Jayeon Kang, Youngsang Kwak, Jaehwa Kwak, Sungjoon Choi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.10544v1 Announce Type: cross Abstract: Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel-wise error yields over-smoothed results, whereas optimizing for perceptual realism often introduces structural deviations. Recent appro...
153. ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover ​
Author: Minwoo Kim, Soochang Song, Namyoon Lee, Bang Chul Jung, Yongjune Kim
Published: 8/12/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.DC
arXiv:2608.10545v1 Announce Type: cross Abstract: Edge LLMs must preserve inference continuity when a user hands over between edge nodes, requiring key-value (KV) cache transfer to the target node. However, simultaneous handovers saturate the backhaul, preventing full cache delivery within the mobil...
154. Retrieval-Corrected Conformal Prediction for Time Series ​
Author: Sangjin Jin, Kangmin Kim, Junhyeong Lee, Yongjae Lee
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10553v1 Announce Type: cross Abstract: Conformal prediction (CP) provides distribution-free prediction intervals for fixed forecasters, but its standard calibration procedure is often inefficient for time series data, where forecast errors are temporally dependent and change across time a...
155. A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language ​
Author: Ushnish Sarkar, Suvajit Patra, Bhaswar Chattopadhyay, Pranab Singha Roy, Tapas Samanta
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10588v1 Announce Type: cross Abstract: Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited. This work introduces a bench...
156. $\pi$-SUB: A Physics-Informed Synthetic Underwater Benchmark Dataset for Underwater Image Enhancement ​
Author: Namritha Lasyapriya Maddali, Rajini Makam, Suresh Sundaram, Narasimhan Sundararajan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10589v1 Announce Type: cross Abstract: This paper presents $\pi$-SUB, a physics-informed framework for generating synthetic underwater benchmark datasets that bridges the synthetic-to-real gap for Underwater Image Enhancement (UIE). The proposed framework extends the classical underwater ...
157. DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction ​
Author: Dong Xu, Zhangfan Yang, Jiantao Wu, Zexuan Zhu, Jianqiang Li, Junkai Ji
Published: 8/12/2026, 4:00:00 AM
Categories: q-bio.BM, cs.AI
arXiv:2608.10595v1 Announce Type: cross Abstract: Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degradation a joint outcome of the degrader molecule and its biological context. Although public databases contain th...
158. Inferential Capability Does Not Determine Legal Scope ​
Author: Nicola Fabiano
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.10601v2 Announce Type: cross Abstract: Two instruments of EU digital law place inference at their centre and mean different things by it. Article 3(1) of the AI Act uses the capability to infer constitutively: it is the central feature separating the regulated category from conventional s...
159. Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts ​
Author: Soumajyoti Sarkar, Yuxin Tang, Sheng Zha
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10605v1 Announce Type: cross Abstract: In large-scale pretraining, the algorithm, architecture, and systems decisions are conventionally made in disconnected stages. A scaling law stage selects an architecture and training recipe, optimizing loss under compute constraints, and a separate ...
160. MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models ​
Author: Yuan Wang, Hualiang Wang, Yixin Chen, Songtao Jiang, Shujian Gao, Jiaming Lin, Siming Fu, Jian Wu, Zuozhu Liu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10635v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches either verbalize regions as coordinate strings or rely on external modul...
161. Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving ​
Author: Jiaping Wang, Shaobo Li, Zhen Wang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10660v1 Announce Type: cross Abstract: Continuous and reliable localization is essential for autonomous driving. Cross-view visual localization matches ground images with satellite maps, providing complementary localization cues for pipelines that depend on Global Navigation Satellite Sys...
162. Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement ​
Author: Lisa M"uhl, Jessica M. Szczuka
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL
arXiv:2608.10672v1 Announce Type: cross Abstract: Social interaction has become one of the most common uses of LLMs, yet research on emotional bonds with AI has focused largely on how users experience these systems, leaving the systems' role in relationship formation poorly understood. Empirically e...
163. Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics ​
Author: Qingjie Zhang, Ziqi Tang, Jie Zhang, Gelei Deng, Jinfeng Li, YueFeng Chen, Yitong Yang, Hui Xue, Tianwei Zhang, Han Qiu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10678v1 Announce Type: cross Abstract: Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three challenges: (1) their web-scale size makes full scan costly; (2) prior analyses are often too coarse to expose token...
164. ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering ​
Author: Akrin Zheng, Alexander Wu, Alaia Liu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2608.10679v2 Announce Type: cross Abstract: Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are work by-products in which required organizational relations remain implicit across heterogeneous source...
165. SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information ​
Author: Junjie Ye, Zhuohui Sheng, Shaofan Liu, Yulun Zhu, Wenjie Fu, Dingwei Zhu, Ming Zhang, Yujiong Shen, Weichao Wang, Xin Zhao, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang, Pluto Zhou
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10692v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated b...
166. Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization ​
Author: Tal Oved, Roi Pony, Oshri Naparstek, Udi barzelay
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.NE
arXiv:2608.10694v2 Announce Type: cross Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs an answering LLM over a validation set, so the evaluator's price tier dictates total search cost. We restructur...
167. ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes ​
Author: Ziyan Wang, Liwen Wu, Cheng Xie, Song Gao, Zhenli He, Xin Jin
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10699v1 Announce Type: cross Abstract: Text-Attributed Graphs (TAGs), endowed with abundant textual content along with topological structures, have emerged as a versatile backbone for real-world anomaly detection spanning large language model security, social network moderation, and cyber...
168. Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control ​
Author: Haoze Liu, Run Liu, Haiying Xu, Jiahui Han, Siyuan Fang, Siyu Yan, Huiqi Deng, Guanchu Wang, Na Zou
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.HC
arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report questionnaires administer...
169. Conversational Orchestration for Organic 6G ​
Author: Masoud Shokrnezhad, Tarik Taleb
Published: 8/12/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.DC, cs.ET, cs.MA
arXiv:2608.10714v1 Announce Type: cross Abstract: The Organic 6G vision of a network of networks spanning an edge-cloud continuum complemented by non-terrestrial resources requires, to realize its promise, service provisioning that is simple to operate, scalable across independently administered dom...
170. Most biomedical publications show signs of LLM-assisted writing ​
Author: Lena Holzwarth, Rita Gonz'alez-M'arquez, Dmitry Kobak
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.DL, cs.SI
arXiv:2608.10715v1 Announce Type: cross Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. T...
171. DuplexWorld: Can voice agents help you get through the day? ​
Author: Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli, Asif Shaik, Abhishek Mukherji, Dinesh Manocha
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL
arXiv:2608.10716v1 Announce Type: cross Abstract: Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistica...
172. Optimal Stopping of Self-Refining Foundation Models ​
Author: Kim Hammar, Tansu Alpcan, Emil C. Lupu
Published: 8/12/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY
arXiv:2608.10729v1 Announce Type: cross Abstract: Foundation models can improve their outputs through a self-refinement process driven by external feedback. In this process, the model is embedded in an iterative loop where it generates outputs, receives feedback from verifiers, and refines its respo...
173. Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGI ​
Author: David Klotz
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.10730v1 Announce Type: cross Abstract: The pursuit of artificial general intelligence (AGI) rests on a seemingly self-evident premise: that general intelligence, the kind of flexible, domain-general cognitive capacity exemplified by Homo sapiens, is extraordinarily valuable. This paper su...
174. A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem ​
Author: Suraj Kumar, Amy Wang, Srinivasan Manoharan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.10760v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has become the de-facto interface for connecting LLM agents to enterprise tools, and adoption has been explosive: within a year, large organizations went from zero to dozens of internally built MCP servers. That speed...
175. The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian Newsrooms During the 2025 Parliamentary Election ​
Author: Mari Reisj{\aa}, Anders Sundnes L{\o}vlie
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.10773v1 Announce Type: cross Abstract: The increasing use of Generative Artificial Intelligence (GenAI) in journalism raises concerns about possible detrimental effects both on journalism and its democratic function. We explore these risks through a case study of GenAI in Norwegian Newsro...
176. MVTrack: Ultrafast Appearance-Free Moving Object Tracking from Compressed Bitstreams ​
Author: I~naki Erregue, Kamal Nasrollahi, Sergio Escalera
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10790v1 Announce Type: cross Abstract: Deploying modern video trackers at scale is bottlenecked by the computational cost of RGB-based object detectors. To this end, we present MVTrack, an ultrafast tracker for moving objects that operates directly on H.264 bitstreams. MVTrack combines MV...
177. Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization ​
Author: Swarnim Maheshwari, Syed Imam Ali, Vineeth N. Balasubramanian
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.10798v2 Announce Type: cross Abstract: Most image colorization systems operate in $Lab$ space by predicting chroma ($ab$) while preserving an input-derived luminance channel ($L$). While effective on standard benchmarks, this fixed-luminance design restricts brightness changes and becomes...
178. BPG: Balancing Plasticity and Generalization for Domain Incremental Learning ​
Author: Qiang Wang, Songlin Dong, Shaokun Wang, Jizhou Han, Xiang Song, Chenhao Ding, Yuhang He, Yihong Gong
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.10804v1 Announce Type: cross Abstract: Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performance degradation under domain shifts. Domain incremental learning (DIL) addresses this challenge by enabling mode...
179. Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation ​
Author: Amit Aflalo, Shahaf E. Finder, Roy Amoyal, Eran Treister, Oren Freifeld
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10805v1 Announce Type: cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive field exponentially with the number of decomposition levels while keeping the parameter count linear. H...
180. Modelling Geographic Atrophy Progression using Implicit Neural Representations ​
Author: Simone Sarrocco, Paul Friedrich, Florentin Bieder, Christina Bornberg, Philippe Valmaggia, Peter Maloca, Philippe Cattin
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10807v1 Announce Type: cross Abstract: Age-related Macular Degeneration (AMD) is the major cause of blindness in the Western world. Its late dry phase is characterised by irreversible atrophic areas, namely Geographic Atrophy (GA). Longitudinal Fundus Autofluorescence (FAF) image acquisit...
181. Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse ​
Author: Zhenyan Zheng, Yunyao Zhang, Junxi Sheng, Junqing Yu, Zikai Song
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10810v2 Announce Type: cross Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks mainly annotate surfac...
182. Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation ​
Author: Chris Han, Pengzhi Gao, Pei Fu, Jian Luan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10812v2 Announce Type: cross Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages t...
183. MIRA: Medical Image Reflection for Agentic Diagnosis ​
Author: Shengzhi Wang, Jun Yang, Kai Wu, Xiaozhong Ji, Yiwen Ye, Ziyang Chen, Mingliang Xiong, Wen Fang, Mingqing Liu, Mengyuan Xu, Miaoxuan Shan, Caiyan Liu, Bin He, Qingwen Liu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10827v1 Announce Type: cross Abstract: Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations, but also ...
184. Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition ​
Author: Gaopeng Xu, Zhenyu Wang, Zheng Xue, Yinfeng Xia, Haitao Yao
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.10836v1 Announce Type: cross Abstract: The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise. This paper introduces the Whisper-Aware LLM, a framework that teaches an Audio...
185. TACTICL: Task-Aware Compression of Tabular ICL Models ​
Author: Mykhailo Koshil, Matthias Feurer, Katharina Eggensperger
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10837v1 Announce Type: cross Abstract: The strong performance of foundation models for tabular tasks comes at substantial inference costs. Distilling models into task-specific architectures reduces model size and computational demands but also sacrifices in-context adaptability. Here we i...
186. VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? ​
Author: Xiaohongshu Inc
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10875v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather t...
187. GitSkills: A Dataset of Agent Skills on GitHub ​
Author: Giuseppe Destefanis, Daniel Graziotin, Matteo Vaccargiu, Marco Ortu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.10906v1 Announce Type: cross Abstract: An agent skill is a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and reference files. The agent loads the skill when it judges that a task matches the skill description. Anthropic i...
188. FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation ​
Author: Rob Cornish, Iacopo Ghinassi, Po-Hung Yeh, Shuqi Liu, Qiyuan Xu, Haoxuan Yin, Dominik Wagner, Wenda Li, Yee Whye Teh, Luke Ong
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LO
arXiv:2608.10916v1 Announce Type: cross Abstract: Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean. We consider how to assess the faithfulness of these systems. Existing approaches require expensive human-annotated ground tr...
189. Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation ​
Author: Dazhao Du, Shiyan Du, Jian Liu, Yongjian Yu, Bohai Gu, Tao Han, Hualuo Liu, Eric Liu, Yujia Zhang, Xi Chen, Song Guo
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10932v1 Announce Type: cross Abstract: Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs) provide a natural interface for this task, but existing work typical...
190. A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models ​
Author: Wajdi Ben Saad, Safa Madiouni
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10939v1 Announce Type: cross Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages....
191. Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes ​
Author: Zhaoyang Wei, Bowen Jiang, Xumeng Han, Jiashu Li, Xuehui Yu, Yuling Liu, Guorong Li, Zhenjun Han, Jianbin Jiao
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10954v1 Announce Type: cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on implicit...
192. CARE: Confidence-Aware Reasoning for Reliable Medical VQA ​
Author: Yuetian Du, Yucheng Wang, Zhenyuan Chen, Luyuan Chen, Rongyu Zhang, Jinjian Zhang, Wei Zhou, Zhijie Xu, Ming Kong, Zhan Zhou, Jie Liu, Qiang Zhu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from $\textit{confidence miscalibration}$---a systematic g...
193. ReLTEx: Reliable LLM-based Taxonomy Expansion ​
Author: Zeinab Ghamlouch, Mehwish Alam
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10970v1 Announce Type: cross Abstract: Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and relations, making them promising tools for taxonomy enrichment. However, directly relying on LLM-generated expansio...
194. TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation ​
Author: Pengyu Zhang, Yangqin Jiang, Klim Zaporojets, Congfeng Cao, Paul Groth
Published: 8/12/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.10983v1 Announce Type: cross Abstract: Multi-modal recommenders fuse collaborative signals with item modalities such as text, images, and audio, but the usefulness of each drifts over time and at different rates. For example, chocolate purchases typically guided by textual ingredient cues...
195. Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers ​
Author: Hongsen Cao, Mona Jaber, Shanxin Yuan, Ahmed Sayed
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.10989v1 Announce Type: cross Abstract: Token-pruning policies are usually designed for a single recognition pipeline, but pretrained Vision Transformers are reused across tasks with different spatial demands. We ask which parts of a pruning policy transfer across image classification, sem...
196. On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation ​
Author: Sicheng Zhang, Zhonghao Yan, Binzhu Xie, Shi Qiu, Muzammal Naseer, Naveed Akhtar, Mubarak Shah
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11002v1 Announce Type: cross Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects insufficiently explored. To...
197. Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis ​
Author: Benjamin Faveri (CEIMIA, Carleton University), Brie Bhasin (University of Ottawa)
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.11006v2 Announce Type: cross Abstract: Governments worldwide have responded to the rapid expansion of AI by publishing national and regional AI strategies. Comparing national and regional AI strategies to identify their convergences and divergences can uncover their common practices, unde...
198. R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video ​
Author: Ke Ma, Yamin Mao, Weiming Li, Shuai Tan, Yijie Zhong, Hao Chen, Haofen Wang, Meng Wang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.HC, cs.MM
arXiv:2608.11017v1 Announce Type: cross Abstract: Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption- and transcript-based mem...
199. Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data ​
Author: Nicola Giuseppe Marchioro, Gabriele Padovani, Amal Gueroudji, Rafael Ferreira da Silva, Wesley Brewer, Valentine Anantharaj, Sandro Fiore, Renan Souza
Published: 8/12/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2608.11022v1 Announce Type: cross Abstract: Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their context, parameters, limitations, and intended use. However, these practices remain focused on static a...
200. Multiclass Sentiment Analysis for Identifying Political Viewpoints ​
Author: Girma Yohannis Bade, Olga Kolesnikova, Jose Luis Oropeza, Grigori Sidorov
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11049v1 Announce Type: cross Abstract: The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives. Sentiment Analysis (SA) is a core task in Natural Langua...
201. 3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment ​
Author: Alam Noor, Luis Almeida, Mohamed Daoudi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.11050v1 Announce Type: cross Abstract: Deep learning systems perform mainly within the 2D for a single image domain and take the face as a single-dimension representation, losing sight of the 3D anatomy of sheep and cross-landmark spatial relationships that are intrinsic to the clinically...
202. A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa ​
Author: Ismail Ismail Tijjani, Sunusi Muhammad Ibrahim, Amina Ibrahim Khaleel, Lanre Olusegun Akinola, Fatima Isa Jibrin, Muhammad Bashir Aliyu, Abdullahi Abdussalam Dalhat, Abdullahi Suiudeen
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.11053v1 Announce Type: cross Abstract: The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming. However, many existing approaches rely on controlled datasets that do not adequately represent realworld farming co...
203. Entropy-Centric Explainable AI for Remote Sensing Image Segmentation ​
Author: Ali Saleh, Abdul Karim Gizzini, Mohamad Ghassany, Ali J. Ghandour
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.11064v1 Announce Type: cross Abstract: Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decision-making process of its models, mainly due to deep neural networks outperforming their peers at the ...
204. Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory ​
Author: Ming Yang
Published: 8/12/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.CC
arXiv:2608.11066v1 Announce Type: cross Abstract: We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses semantic history into a future-accessible boundary state and later answers a query. We count communication $B$, persistent instance-dep...
205. Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting ​
Author: Kiran Madhusudhanan, Christian Kl"otergens, Lars Schmidt-Thieme, Vijaya Krishna Yalavarthi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11114v1 Announce Type: cross Abstract: Probabilistic forecasting plays an essential role in risk-sensitive decision-making, particularly in long-horizon settings. However, existing approaches often face a fundamental trade-off between distributional flexibility and accurate mean predictio...
206. Attention-Path Fragility as an Uncertainty Signal in Large Language Models ​
Author: Minsoo Kim, Sungyoung Ji, Kisung Moon, Ilyong Yoon
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11138v1 Announce Type: cross Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways. We instantiate this as ASMI...
207. From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop ​
Author: Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle, Anil Ramakrishna, Anubrata Das, Apurv Verma, Jwala Dhamala, Ninareh Mehrabi, Tharindu Kumarage, Yada Pruksachatkun, Yang Trista Cao, Kai-Wei Chang, Aram Galstyan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.11171v1 Announce Type: cross Abstract: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of s...
208. How to Verify Consistency of Probabilistic Claims ​
Author: Orr Paradise, Oliver Richardson, Yoshua Bengio, Shafi Goldwasser
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CC, cs.AI, cs.LG
arXiv:2608.11181v1 Announce Type: cross Abstract: When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem is of interest for AI safety, where safety is derived from honesty about probabili...
209. Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation ​
Author: Shiyu Xuan, Zechao Li
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.11191v1 Announce Type: cross Abstract: GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time r...
210. ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls ​
Author: Chen Lyu, Xingwei Tan, Simon Cullen, Shelley Wilson, Lois Arthurs, Arshad Jhumka, Gabriele Pergola
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.11200v1 Announce Type: cross Abstract: Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly...
211. Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning ​
Author: Wenrui Bao, Tianyun Jiang, Zhiben Chen, Ser-Nam Lim, Peter D. Peng, Yuzhang Shang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.11204v1 Announce Type: cross Abstract: Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g., dVRK) trajectories with synchronized kinematics are costly to collect, while surgical tasks demand p...
212. Representation and Invariance in Reinforcement Learning ​
Author: Samuel Alexander, Arthur Paul Pedersen
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, cs.LG
arXiv:2112.07752v5 Announce Type: replace Abstract: Researchers have formalized reinforcement learning (RL) in different ways. If an agent in one RL framework is to run within another RL framework's environments, the agent must first be converted, or mapped, into that other framework. In this paper,...
213. GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning ​
Author: Jusheng Zhang, Yijia Fan, Wenjun Lin, Ruiqi Chen, Haoyi Jiang, Wenhao Chai, Jian Wang, Keze Wang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2505.23399v2 Announce Type: replace Abstract: We propose GAM-Agent, a game-theoretic multi-agent framework for enhancing vision-language reasoning. Unlike prior single-agent or monolithic models, GAM-Agent formulates the reasoning process as a non-zero-sum game between base agents--each specia...
214. Closing a 17-Year Gap: Algorithmic Detection and Empirical Prevalence of Rank Reversal in Multi-Criteria Decision Analysis ​
Author: Juan Bautista Cabral, Gonzalo Giarda, Diego Nicol'as Gimenez Irusta, Paula Pacheco, Alvaro Roy Schachner, Agust'in Borda
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, math.OC
arXiv:2508.00129v2 Announce Type: replace Abstract: Rank Reversal, where the relative order of alternatives changes in ways that violate axioms of rational decision-making, is a well-documented threat to the reliability of Multi-Criteria Decision Analysis (MCDA) methods. Wang and Triantaphyllou (200...
215. Multiplayer Nash Preference Optimization ​
Author: Fang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang, Yijia Xiao, Guancheng Wan, Xiaomin Li, Bing Hu, Peng Xia, Jure Leskovec, Yejin Choi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2509.23102v4 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However, reward-based methods grounded in the Bradley-Terry assumption struggle to capture the nontran...
216. On The Statistical Limits of Self-Improving Agents ​
Author: Charles L. Wang, Keir Dorchen, Peter Jin
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2510.04399v3 Announce Type: replace Abstract: We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework, we prove a sharp boundary: under standard i.i.d. assumptions, distribution-free PAC learnability i...
217. Situation Graph Prediction for User Perspective Modeling ​
Author: Jisung Shin, Daniel Platnick, Marjan Alirezaie, Hossein Rahnama
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data bottleneck: digital footprints are privacy-sensitive and perspective states are rarely labeled. We p...
218. Leveraging Large Language Models for Causal Discovery: a Constraint-based, Argumentation-driven Approach ​
Author: Zihao Li, Fabrizio Russo
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2602.16481v2 Announce Type: replace Abstract: Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the effects of interventions. While expert knowledge is required to construct principled causal graphs, many stati...
219. JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures ​
Author: Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Marissa Wirth, Alexander W. Charney, Nati Daniel, Yoli Shavit
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, q-bio.GN
arXiv:2602.17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature". While effective at capturing local syntax, these generative paradigms prioritize token-level reconstruct...
220. CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification ​
Author: Hanrong Zhang (Steve), Shicheng Fan (Steve), Henry Peng Zou (Steve), Yankai Chen (Steve), Zhenting Wang (Steve), Jiayu Zhou (Steve), Chengze Li (Steve), Wei-Chieh Huang (Steve), Yifei Yao (Steve), Kening Zheng (Steve), Xue (Steve), Liu, Xiaoxiao Li, Philip S. Yu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.01687v3 Announce Type: replace Abstract: Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool is a single, self-contained function, whereas a skill is a structured bundle of interdependent multi...
221. Planning Task Shielding: Detecting and Repairing Flaws in Planning Tasks through Turning them Unsolvable ​
Author: Alberto Pozanco, Pietro Totis, Marianela Morales, Daniel Borrajo
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.07042v3 Announce Type: replace Abstract: Most research in planning focuses on generating a plan to achieve a desired set of goals. However, a goal specification can also be used to encode a property that should never hold, allowing a planner to identify a trace that would reach a flawed s...
222. Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents ​
Author: Bhaskar Gurram
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA
arXiv:2604.16706v2 Announce Type: replace Abstract: Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, yet this assumption is rarely validated against human annotation. We present AgentProp-Bench, a diagnostic benchmark of 14,750 execution traces f...
223. When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning ​
Author: Vasilis Niarchos, Constantinos Papageorgakis, Alexander G. Stapleton, Sokratis Trifinopoulos
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, hep-ph, hep-th
arXiv:2605.06772v2 Announce Type: replace Abstract: As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical question emerges: How does the interaction between researchers and agents affect the results? We study...
224. CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG ​
Author: Jianghan Shen, Siqi Luo, Xinyu Cheng, Jing Xiong, Yue Li, Jiyao Liu, Jiashi Lin, Yirong Chen, Junjun He
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.11611v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG) systems from outcome-only supervision. Most existing methods optimize policies from uniformly sample...
225. Memory-Augmented Reinforcement Learning Agent for CAD Generation ​
Author: Yin Xiaolong, Liu Yu, Shen Jiahang, Lu Xingyu, Ni Jingzhe, Fan Fengxiao, Sang Fan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2605.19748v2 Announce Type: replace Abstract: Automatic generation of computer-aided design (CAD) models is a core technology for enabling intelligence in advanced manufacturing. Existing generation methods based on large language models (LLMs) often fall short when handling complex CAD models...
226. A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents ​
Author: Vasundra Srinivasan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2605.20173v2 Announce Type: replace Abstract: Production LLM agents combine stochastic model outputs with deterministic software systems, yet the boundary between the two is rarely treated as a first-class architectural object. This paper names that boundary the stochastic-deterministic bounda...
227. From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection ​
Author: Ke Liu, Jiwei Wei, Wenyu Zhang, Shuchang Zhou, Ruikun Chai, Yutao Dai, Chaoning Zhang, Yang Yang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.MM, cs.SD
arXiv:2605.27944v2 Announce Type: replace Abstract: With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In singing, rhythmic vocalization w...
228. Moxia: A Trust-First Neuro-Symbolic Execution Architecture for Self-Explaining Mathematical Reasoning ​
Author: Alessio Bruno
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2606.00671v3 Announce Type: replace Abstract: We present Moxia (formerly AXIOM), a trust-first neuro-symbolic architecture for self-explaining mathematical reasoning over natural-language input. Its language model is strictly a canonicalizer: it rewrites informal problem text into a narrow sch...
229. A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline ​
Author: Kai A. Horstmann, Ethan Lin, Alice A. Robie, Jennifer J. Sun, Kristin Branson
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2606.07718v2 Announce Type: replace Abstract: Agentic AI offers a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that take domain experts days to months to build and where correctness and robustness matter more than imple...
230. When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting ​
Author: Binyan Xu, Xilin Dai, Fan Yang, Kehuan Zhang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CE
arXiv:2606.16465v2 Announce Type: replace Abstract: AI agents can now take irreversible actions in operational systems, but agent-caused losses are still not clearly assigned, priced, or transferred. Providers often disclaim consequential damages, users are left with uncompensated losses, and defaul...
231. ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection ​
Author: Chenhao Dang, Dantong Zhu, Jun Yang, Conghui He, Weijia Li
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.24112v2 Announce Type: replace Abstract: Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed provenance, and subtle cross-modal framing errors. Existing benchmarks and methods remain poorly match...
232. Coachable agents for interactive gameplay ​
Author: Roberto Capobianco (Sony AI, Zurich, Switzerland), Harm van Seijen (Sony AI, North America, various locations), Nolan D. Bard (Sony AI, North America, various locations), Neil Burch (Sony AI, North America, various locations), Fatima Davelouis (Sony AI, North America, various locations), Josh Davidson (Sony AI, North America, various locations), Alisa Devlic (Sony AI, Zurich, Switzerland), Yunshu Du (Sony AI, North America, various locations), Ishan Durugkar (Sony AI, North America, various locations), Siddhant Gangapurwala (Sony AI, North America, various locations), Daniel Hernandez (Sony AI, North America, various locations), G. Zacharias Holland (Sony AI, North America, various locations), Sahil Jain (Sony AI, North America, various locations), Kenta Kawamoto (Sony AI, Tokyo, Japan), Raksha Kumaraswamy (Sony AI, North America, various locations), Patrick MacAlpine (Sony AI, North America, various locations), Dustin R. Morrill (Sony AI, North America, various locations), Declan Oller (Sony AI, North America, various locations), Francesco Riccio (Sony AI, Zurich, Switzerland), Akanksha Saran (Sony AI, North America, various locations), Craig Sherstan (Sony AI, Tokyo, Japan), Kaushik Subramanian (Sony AI, Zurich, Switzerland), Thomas J. Walsh (Sony AI, North America, various locations), Samuel Barrett (Sony AI, North America, various locations), Kizza N. Frisbee (Sony AI, North America, various locations), Mady Govil (Sony AI, North America, various locations), Johannes G"unther (Sony AI, North America, various locations), Varun R. Kompella (Sony AI, North America, various locations), James A. MacGlashan (Sony AI, North America, various locations), Maxwell Svetlik (Sony AI, North America, various locations), Michael D. Thomure (Sony AI, North America, various locations), Jaden B. Travnik (Sony AI, North America, various locations), Kevin Waugh (Sony AI, North America, various locations), Elahe Aghapour (Sony AI, North America, various locations), Florian Fuchs (Sony AI, Zurich, Switzerland), Andreanne Lemay (Sony AI, North America, various locations), Shruti Mishra (Sony AI, Zurich, Switzerland), Takuma Seno (Sony AI, Tokyo, Japan), Peter Stone (Sony AI, North America, various locations), Michael Spranger (Sony AI, Tokyo, Japan), Peter R. Wurman (Sony AI, North America, various locations)
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.00642v2 Announce Type: replace Abstract: Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models. Through trial-and-error, these AI systems typically learn one...
233. Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems ​
Author: Rongping Zhou, Omid Tavallaie, Shuaijun Chen, Albert Y. Zomaya
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.10309v2 Announce Type: replace Abstract: Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT). However, the trial-and-error nature of RL, when conducted in real-world environments, is costly a...
234. UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention ​
Author: Cheng Yan, Zhijun Fan, Guangyang Ye, Fan Xu, Xiang Xia, Yawei Wang, Wuyang Zhang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.17188v2 Announce Type: replace Abstract: While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can also exacerbate overthinking and underthinking, which we formulate as reasoning state--action mismat...
235. SAE-StatSteer: Statistical Consensus Feature Selection for Optimization-Free Activation Steering of Large Language Models ​
Author: Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.19364v2 Announce Type: replace Abstract: Activation steering adds a residual-stream direction at inference time, providing lightweight behavioral control without fine-tuning. Sparse autoencoders (SAEs) can make such interventions auditable by decomposing dense activations into an approxim...
236. AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction ​
Author: Qinfeng Li, Yuntai Bao, Xinyan Yu, Hongze Chen, Yanming Liu, Huifeng Zhu, Yier Jin, Jintao Chen, Wenqi Zhang, Xuhong Zhang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21106v3 Announce Type: replace Abstract: Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory meth...
237. Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog ​
Author: Bartolomeo Bogliolo
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SE
arXiv:2607.21412v2 Announce Type: replace Abstract: Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning, especially in safety-critical or compliance-sensitive domains. Recent neuro-symbolic approaches address this...
238. Learning and Structurally Validating Simulation Scenario Continuations in Dynamic Graph Systems ​
Author: Michael Romei de Socio, Gian Luca Pozzato, Alessio Merlo
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21421v2 Announce Type: replace Abstract: Data-driven generative models can extend partially observed simulation trajectories into ensembles of alternative future scenarios. However, consistency with a learned trajectory distribution does not ensure that generated continuations satisfy the...
239. EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff ​
Author: Xiao Ma, Zhiquan Hu, Yi Wei, Chenchen Zhao, Yijun Chen, Jicheng Zhao, Yuming Li, Chuang Dai
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.23955v3 Announce Type: replace Abstract: Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups provide no comparative signal and may hide useful search behavior. We present EviBack, an evidence- constrai...
240. TrAC: Trace-Conditioned Answer Consistency for Efficient Uncertainty Quantification in LLMs ​
Author: Dahai Yu, Lin Jiang, Rongchao Xu, Guang Wang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00422v2 Announce Type: replace Abstract: Large language models (LLMs) can generate fluent reasoning traces that nevertheless lead to incorrect answers, making response-level uncertainty estimation important for abstention, human review, and adaptive compute allocation. Existing approaches...
241. DiffImaginE: Imagine to Verify Entity Types with Diffusio ​
Author: Feng Zhang, Feiyu Han, Rongxin Yang, Yang Liu, Yancheng Chen, Rui Wang, Yingguang Yang, Tian Xueyun, Chongyang Zhang, Hao Zheng, Xu Kefu, Congjing Ran, Fuhai Chen, Bin Chong
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03025v2 Announce Type: replace Abstract: Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visua...
242. CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting ​
Author: Xiaoyu Tao, Mingyue Cheng, Bokai Pan, Chuang Jiang, Huanjian Zhang, Tian Gao, Yaguo Liu, Qi Liu, Enhong Chen
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03031v2 Announce Type: replace Abstract: Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features. Recent advances in large language models (LLMs) have ex...
243. TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management in Long-Term Agents ​
Author: Han Xiao, Hongjun Xu, Xin Zhang, Yidong Chen, Xiaodong Shi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03699v2 Announce Type: replace Abstract: Persistent memory helps long-term agents retain knowledge, yet a single update error can repeatedly distort future retrieval and reasoning. Most existing systems reduce memory updating to a binary Write/Hold decision, which cannot distinguish wheth...
244. Small Foundation Models of Human Cognition and Behaviour ​
Author: Nick Oh, Fernand Gobet
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.05224v3 Announce Type: replace Abstract: Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We t...
245. ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment ​
Author: Abdulkadir K"ul\c{c}e, Alihan Esen, \c{C}a\u{g}la Fikir, Berke Kurt, Kuzey Arar, G"okhan Ercan, Faik Boray Tek
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.06110v2 Announce Type: replace Abstract: This paper presents ECHO (Enhanced Care & Health Observer), a locally-deployable conversational health assistant for long-term chronic care management. ECHO integrates three complementary software modules developed under shared supervision as a uni...
246. Contextual Information Policy Optimization for Search Agents ​
Author: Xingyu Guo, Wei Chen, Linlin Yang, Baochang Zhang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06128v3 Announce Type: replace Abstract: Search agents extend large language models beyond static parametric memory by enabling them to acquire and use external evidence during multi-step reasoning. For knowledge-intensive tasks involving complex or evolving information, their reliability...
247. Blast Radius ​
Author: MY Pitsane, Hope Mogale
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07440v2 Announce Type: replace Abstract: Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates an incoming prompt's reach through coupled context and code channels. NECROPHORESIS enables rev...
248. TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents ​
Author: Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07917v2 Announce Type: replace Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis. Optical character recognition (OCR) can bri...
249. Thought-Level Beam Search for Reasoning ​
Author: Lijie Yang, Hongyin Luo, Jiawei Zhao, Tri Dao, Ravi Netravali
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.08020v2 Announce Type: replace Abstract: Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from \emph{how much} compute to spend, to \emph{where} to allocate it....
250. StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning ​
Author: Yifan Li, Ruxin Sun, Tongzhou Zhao
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.08326v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, most existing methods evaluate an entire response using a binary reward based only on final-answer correctness, ...
251. SDDBMs: Soft Denoising Diffusion Bridge Models ​
Author: Shiyi Qi, Kun He, Mingmou Liu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.08594v2 Announce Type: replace Abstract: Diffusion bridge models leverage Doob's (h)-transform to construct stochastic transports between arbitrary endpoint distributions, and have shown strong potential in image-to-image translation and restoration. However, most existing bridge models...
252. ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration ​
Author: Guo Chen, Ziwen Li, Reed Li, Yu Lu, Haibo Shi, Bingbing Xu, Junjie Huang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.08605v2 Announce Type: replace Abstract: Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common basis for evaluation across methods. Outcome-only benchmarks discard collaborations, whereas LLM-a...
253. Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution ​
Author: Bohan Lin, Hejia Geng, Xinyi Xie, Heng Zhou, Qinghua Xing, Bo Liu, Chen Zhang, Yudong Zhang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.09248v2 Announce Type: replace Abstract: Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirely on text-level signals such as task descriptions, verbal reflections, and experience-derived rules, while th...
254. Entropy-based Code Adversarial Translation for Real-world Repository Migration ​
Author: Yushun Tang, Yisen Cao, Zhicheng Chen, Lin Peng, Junkang Mao, Fengyi Song, Yantao Jia
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.09273v2 Announce Type: replace Abstract: LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produces a runnable application because long-horizon translation challenges LLM-based agents' ability to maintain ...
255. Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models ​
Author: Kevin Murphy
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.09696v2 Announce Type: replace Abstract: Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, causal model, not a curve fit; and learning such a model requires \emph{experiments}, because passive data lea...
256. CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation ​
Author: Yaoning Yu, Kai-Min Chang, Ye Yu, Yi-Chia Wang, Haojing Luo, Haohan Wang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SI
arXiv:2608.09790v2 Announce Type: replace Abstract: Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these discussions requires more than just generating individual comments, the generated threads should also match h...
257. Emergent Neural Network Mechanisms for Generalization to Objects in Novel Orientations ​
Author: Avi Cooper, Xavier Boix, Daniel Harari, Spandan Madan, Hanspeter Pfister, Tomotake Sasaki, Pawan Sinha
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, q-bio.NC, stat.ML
arXiv:2109.13445v3 Announce Type: replace-cross Abstract: The capability of Deep Neural Networks (DNNs) to recognize objects in orientations outside the distribution of the training data is not well understood. We present evidence that DNNs are capable of generalizing to objects in novel orientation...
258. Graphical Models of False Information and Fact Checking Ecosystems ​
Author: Haiyue Yuan, Enes Altuncu, Shujun Li, Can Baskent, Jason R. C. Nurse
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SI, cs.AI, cs.CR
arXiv:2208.11582v2 Announce Type: replace-cross Abstract: The wide spread of false information online, including misinformation and disinformation, has become a major problem for our highly digitised and globalised society. A lot of research has been done to better understand different aspects of fa...
259. Pretrained Optimization Model for Zero-Shot Black Box Optimization ​
Author: Xiaobin Li, Kai Wu, Yujian Betterest Li, Xiaoyu Zhang, Handing Wang, Jing Liu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.NE, cs.AI
arXiv:2405.03728v3 Announce Type: replace-cross Abstract: Zero-shot optimization involves optimizing a target task that was not seen during training, aiming to provide the optimal solution without or with minimal adjustments to the optimizer. It is crucial to ensure reliable and robust performance i...
260. Regression and Classification with Single-Qubit Quantum Neural Networks ​
Author: Leandro C. Souza, Bruno C. Guingo, Gilson Giraldi, Renato Portugal
Published: 8/12/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG
arXiv:2412.09486v2 Announce Type: replace-cross Abstract: The literature reflects a mutually beneficial relationship between machine learning and quantum computing, where progress in one field frequently drives improvements in the other. Motivated by the rich connection between these areas, we use a...
261. Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking ​
Author: Ziwei Zhang, Juan Wen, Wanli Peng, Zhengxian Wu, Yinghan Zhou, Yiming Xue
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2504.00035v4 Announce Type: replace-cross Abstract: Large language models (LLMs) enable powerful knowledge injection through approaches such as in-context learning and fine-tuning, but they also introduce new risks of unauthorized imitation of high-value creative works. Existing copyright prot...
262. TransitReID: Transit OD Data Collection with Occlusion-Resistant Dynamic Passenger Re-Identification ​
Author: Kaicong Huang, Talha Azfar, Jack Reilly, Ruimin Ke
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, eess.IV
arXiv:2504.11500v3 Announce Type: replace-cross Abstract: Transit Origin-Destination (OD) data are fundamental for optimizing public transit services, yet current collection methods, such as manual surveys, Bluetooth/WiFi tracking, and Automated Passenger Counters, are often costly, device-dependent...
263. X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation ​
Author: Peizhen Li, Longbing Cao, Xiao-Ming Wu, Runze Yang, Xiaohan Yu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.HC
arXiv:2505.11146v4 Announce Type: replace-cross Abstract: Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition challenge due to the significant domain gap between biological facial dynamics and mechanical control spaces. While visual synthesis ...
264. Demystifying Adversarial Robustness in Diffusion Models: Compression, Randomness, and Geometry ​
Author: Liu Yuezhang, Xue-Xin Wei
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2505.22839v2 Announce Type: replace-cross Abstract: Recent studies suggest that diffusion models significantly improve the empirical adversarial robustness of deep neural network models. While intuitive explanations have been proposed, the mechanisms underlying diffusion-based robustness remai...
265. Music Interpretation and Emotion Perception: A Computational and Neurophysiological Investigation ​
Author: Vassilis Lyberatos, Spyridon Kantarelis, Ioanna Zioga, Christina Anagnostopoulou, Giorgos Stamou, Anastasia Georgaki
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2506.01982v5 Announce Type: replace-cross Abstract: This study investigates emotional expression and perception in music performance using computational and neurophysiological methods. The influence of different performance settings, such as repertoire, diatonic modal etudes, and improvisation...
266. HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models ​
Author: Zhaolu Kang, Junhao Gong, Jiaxu Yan, Wanke Xia, Yian Wang, Ziwen Wang, Huaxuan Ding, Zhuo Cheng, Wenhao Cao, Zhiyuan Feng, Siqi He, Shannan Yan, Junzhe Chen, Xiaomin He, Chaoya Jiang, Wei Ye, Kaidong Yu, Xuelong Li
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV
arXiv:2506.03922v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical step-by-step reasoning t...
267. SynBoost: A Synergistic Framework for Fast Sampling of Diffusion Models ​
Author: Hu Yu, Hao Luo, Xueyang Fu, Jie Huang, Fan Wang, Feng Zhao
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2506.13058v2 Announce Type: replace-cross Abstract: Diffusion probabilistic models (DPMs) have demonstrated remarkable success in visual generation. However, their iterative sampling mechanism results in slow inference speeds. While reducing sampling steps offers an intuitive acceleration stra...
268. OpenDPDv2: A Unified Learning and Optimization Framework for Neural Network Digital Predistortion ​
Author: Yizhuo Wu, Ang Li, Chang Gao
Published: 8/12/2026, 4:00:00 AM
Categories: eess.SP, cs.AI
arXiv:2507.06849v3 Announce Type: replace-cross Abstract: Neural network (NN)-based Digital Predistortion (DPD) improves linearization for wideband radio frequency (RF) power amplifiers (PAs) but often increases the complexity of the digital back-end. This paper presents OpenDPDv2, an open-source en...
269. Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling ​
Author: Wei Da, Evangelia Kalyvianaki
Published: 8/12/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2508.03611v3 Announce Type: replace-cross Abstract: This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language model (LLM) serving. Astrolabe improves load balancing without relying on migration-based rebalancing, whos...
270. Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False Negatives ​
Author: Sarah Jabbour, David Fouhey, Nikola Banovic, Stephanie D. Shepard, Ella Kazerooni, Michael W. Sjoding, Jenna Wiens
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2508.07617v2 Announce Type: replace-cross Abstract: AI has the potential to augment human decision making. However, even high-performing models can produce inaccurate predictions when deployed. These inaccuracies, combined with automation bias, where humans overrely on AI predictions, can resu...
271. Reconfiguration of pivoting cube ensembles under local sensing constraints using geometric deep learning ​
Author: Nadezhda Dobreva, Emmanuel Blazquez, Jai Grover, Dario Izzo, Yuzhen Qin, Dominik Dold
Published: 8/12/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.RO
arXiv:2509.03140v2 Announce Type: replace-cross Abstract: We demonstrate that local sensing is sufficient for effective global reconfiguration of homogeneous pivoting cube modular robots in two dimensions. While cube selection (i.e., which cube executes a movement) is assumed to be globally coordina...
272. Token-Based Detection of Spurious Correlations in Vision Transformers ​
Author: Solha Kang, Esla Timothy Anzaku, Wesley De Neve, Arnout Van Messem, Joris Vankerschaver, Francois Rameau, Utku Ozbulak
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2509.04009v2 Announce Type: replace-cross Abstract: Due to their powerful feature association capabilities, neural network-based computer vision models have the ability to detect and exploit unintended patterns within the data, potentially leading to correct predictions based on incorrect or u...
273. Faster Results from a Smarter Schedule: Reframing Collegiate Cross Country through Analysis of the National Running Club Database ​
Author: Jonathan A. Karr Jr, Ryan M. Fryer, Nitesh V. Chawla
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG
arXiv:2509.10600v5 Announce Type: replace-cross Abstract: Collegiate cross country teams often build their season schedules on intuition rather than evidence, partly because large-scale performance datasets were not publicly accessible prior to the National Running Club Database (NRCD). We analyze t...
274. Diffusion-Based Impedance Learning for Contact-Rich Manipulation Tasks ​
Author: Noah Geiger, Tamim Asfour, Neville Hogan, Johannes Lachner
Published: 8/12/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2509.19696v4 Announce Type: replace-cross Abstract: Learning-based methods excel at robot motion generation but remain limited in contact-rich physical interaction. Impedance control provides stable and safe contact behavior but requires task-specific tuning of stiffness and damping parameters...
275. Pricing Access to Dynamic Information Services ​
Author: Weijie Zhong
Published: 8/12/2026, 4:00:00 AM
Categories: econ.TH, cs.AI
arXiv:2510.09859v5 Announce Type: replace-cross Abstract: A provider sells a \emph{dynamic information service}---a real-time, capacity-constrained process that resolves a customer's uncertainty---to customers who differ privately in urgency. I characterize the revenue-optimal mechanism: deploy a \e...
276. HyWA: Architecture-Preserving Personalized Voice Activity Detection for Full-Duplex Voice Assistants ​
Author: Hamed Jafarzadeh Asl, Amin Edraki, Mahsa Ghazvini Nejad, Masoud Asgharian, Mohammadreza Sadeghi, Yuanhao Yu, Vahid Partovi Nia
Published: 8/12/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.LG, cs.SD
arXiv:2510.12947v3 Announce Type: replace-cross Abstract: Voice activity detection (VAD) serves as an early gate in voice-assistant pipelines for smart devices. Because conventional VADs respond to speech from any speaker, nearby conversations and residual assistant playback lead to unwanted trigger...
277. VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection ​
Author: Qiang Wang, Xinyuan Gao, Yuhang He, Jizhou Han, Jiangyang Li, SongLin Dong, Zhiheng Ma, Yihong Gong
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.MM
arXiv:2511.19436v2 Announce Type: replace-cross Abstract: Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary models, creating a dependency on external supervision. In this paper, we propose VDC-Agent, an autonomou...
278. On the Condition Number Dependency in Bilevel Optimization ​
Author: Lesi Chen, Kaiyi Ji, Jingzhao Zhang
Published: 8/12/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.LG
arXiv:2511.22331v4 Announce Type: replace-cross Abstract: Bilevel optimization minimizes an objective function, defined by an upper-level problem whose feasible region is the solution of a lower-level problem. We study the oracle complexity of finding an $\epsilon$-stationary point with first-order ...
279. Auto-exploration for online reinforcement learning ​
Author: Caleb Ju, Guanghui Lan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC
arXiv:2512.06244v4 Announce Type: replace-cross Abstract: The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. Existing algorithms for finite state and action discounted RL problems address this by assuming sufficient exploration ...
280. Hybrid Token Compression for Vision-Language Models ​
Author: Jusheng Zhang, Xiaoyang Guo, Tongyu Mo, Qinhan Lv, Wenhao Chai, Jian Wang, Keze Wang, Liang Lin
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2512.08240v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) rely on hundreds of visual tokens, leading to high computational and memory costs. Existing compression methods face a trade-off: continuous compression can weaken high-level semantics, while discrete quantizatio...
281. IndexTTS 2.5 Technical Report ​
Author: Yunpei Li, Xun Zhou, Jinchao Wang, Lu Wang, Yong Wu, Siyi Zhou, Yiquan Zhou, Yining Wang, Yaogen Yang, Zhetao Hu, Shiyao Duan, Jiacheng Xu, Jingchen Shu, Bin Xia
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2601.03888v5 Announce Type: replace-cross Abstract: In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-based Text-to-Semantic (T2S) module and a non-autoregressive Semantic-to-Mel (S2M) module, which togethe...
282. On Solomonoff Induction in Large Language Models and the Limits of Self-Improving: The Singularity Is Not Near Without Symbolic Model Synthesis ​
Author: Hector Zenil
Published: 8/12/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, cs.LG, math.IT
arXiv:2601.05280v3 Announce Type: replace-cross Abstract: On the one hand, the question of whether large language models (LLMs) are Solomonoff induction estimators has become an explicit question at the intersection of Algorithmic Information Theory (AIT) and Machine Learning (ML) of great interest....
283. LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning ​
Author: Linquan Wu, Tianxiang Jiang, Yifei Dong, Haoyu Yang, Fengji Zhang, Shichaang Meng, Ai Xuan, Linqi Song, Jacky Keung
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2601.10129v2 Announce Type: replace-cross Abstract: Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dynamics. In this work, we identify a critical Perception Gap in distillation: student models frequently mi...
284. GraFine: Retrieval-Time Refinement for Efficient Graph RAG over Corpus Graphs ​
Author: Seonho An, Chaejeong Hyun, Min-Soo Kim
Published: 8/12/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2601.18579v2 Announce Type: replace-cross Abstract: Graph RAG on corpus graphs enhances retrieval by leveraging intermediate node content as contextual clues to uncover unretrieved oracle nodes. However, existing methods suffer from two critical blind spots, namely semantically blind graph exp...
285. Bandwidth-Efficient Multi-Agent Communication through Information Bottleneck and Vector Quantization ​
Author: Ahmad Farooq, Kamran Iqbal
Published: 8/12/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.IT, cs.LG, cs.MA, math.IT
arXiv:2602.02035v2 Announce Type: replace-cross Abstract: Multi-agent reinforcement learning systems deployed in real-world robotics applications face severe communication constraints that significantly impact coordination effectiveness. We present a framework that combines information bottleneck th...
286. LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations ​
Author: William Lugoloobi, Thomas Foster, William Bankes, Chris Russell
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2602.09924v4 Announce Type: replace-cross Abstract: Running LLMs with extended reasoning on every problem is expensive, but determining which inputs actually require additional compute remains challenging. We investigate whether their own likelihood of success is recoverable from their interna...
287. Do LLMs Benefit From Their Own Words? ​
Author: Jenny Y. Huang, Leshem Choshen, Wei Sun, Omar Khattab, Ram'on Fernandez Astudillo, Mehul Damani, Tamara Broderick, Jacob Andreas
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2602.24287v2 Announce Type: replace-cross Abstract: In multi-turn conversations, large language models typically condition on the full conversation history: both past user prompts and assistant responses. We revisit this design choice by comparing full-context prompting to four alternative, su...
288. Exact and Asymptotically Complete Robust Verifications of Neural Networks via Ising Solvers ​
Author: Wenxin Li, Wenchao Liu, Weihao Li, Chuan Wang, Qi Gao, Yin Ma, Hai Wei, Kai Wen
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.optics, quant-ph
arXiv:2603.00408v4 Announce Type: replace-cross Abstract: We present an Ising-compatible framework for formal neural-network robustness verification under bounded input perturbations. For piecewise-linear activations, the Exact Logarithmic PWL Model (Log-PWL) provides an exact, sound, and complete f...
289. Can Computational Reducibility Lead to Transferable Models for Graph Combinatorial Optimization? ​
Author: Semih Cant"urk, Thomas Sabourin, Frederik Wenkel, Michael Perlmutter, Guy Wolf
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2603.02462v2 Announce Type: replace-cross Abstract: A key challenge in developing unified neural solvers for combinatorial optimization (CO) is the efficient generalization of models from a given set of tasks to new tasks unseen during initial training. To address this, we first establish a ne...
290. $\mathrm{ECI}_{\mathrm{sem}}$: Semantic Residual Effective Contrastive Information for Evaluating Hard Negatives ​
Author: Aarush Sinha, Rahul Seetharaman, Aman Bansal
Published: 8/12/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2603.20990v4 Announce Type: replace-cross Abstract: Hard-negative source selection for dense retrieval is usually decided only after fine-tuning and downstream evaluation. We propose ECIsem, a validity-weighted diagnostic that ranks candidate hard-negative sources using frozen target-encoder e...
291. Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding ​
Author: Gregor Baer, Chao Zhang, Isel Grau, Pieter Van Gorp
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.LG
arXiv:2603.25251v2 Announce Type: replace-cross Abstract: Explainable AI (XAI) methods are commonly evaluated using functional correctness metrics, sometimes termed faithfulness or fidelity, which estimate how closely an explanation reflects the model's reasoning. Higher correctness is assumed to pr...
292. Covert Visual Prompt Injection against Commercial Multimodal Large Language Models ​
Author: Meiwen Ding, Song Xia, Chenqi Kong, Xudong Jiang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.29418v2 Announce Type: replace-cross Abstract: Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior leaves them vulnerable to prompt injection attacks. Existing prompt injection methods predominantly r...
293. Evaluation and Hardening of LLM System Instructions Against Extraction via Encoding Attacks ​
Author: Anubhab Sahu, Diptisha Samanta, Reza Soosahabi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2604.01039v4 Announce Type: replace-cross Abstract: System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensitive operational context in agentic AI applications. These instructions may contain sensitive informatio...
294. BiScale-GTR: Fragment-Aware Graph Transformers for Multi-Scale Molecular Representation Learning ​
Author: Yi Yang, Ovidiu Daescu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.06336v2 Announce Type: replace-cross Abstract: Fragment-level representations provide a natural way to capture recurring molecular substructures and reuse their learned representations across molecules. However, a shared fragment identity alone may not fully describe how a fragment is ins...
295. RankFormer: A Propose-then-Select Transformer for Multi-Agent Multimodal Trajectory Prediction ​
Author: Diyi Liu, Zihan Niu, Tu Xu, Xingchen Zhang, Lishan Sun
Published: 8/12/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2604.07126v3 Announce Type: replace-cross Abstract: Predicting traffic agent trajectories plays an important role in autonomous driving, traffic operations, transportation safety analysis, etc. Although many deep learning algorithms are devised to predict future agent trajectories, the traject...
296. Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers ​
Author: Harsh Kohli, Srinivasan Parthasarathy, Huan Sun, Yuekun Yao
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2604.07822v2 Announce Type: replace-cross Abstract: We study implicit reasoning, i.e. the ability to combine knowledge or rules within a single forward pass. While transformer-based large language models store substantial factual knowledge and rules, they often fail to compose this knowledge f...
297. SatIR: Scalable High-Recall Constraint-Satisfaction-Based Information Retrieval for Clinical Trials Matching ​
Author: Zikai Zhou, Yufei Jin, Yilin Xu, Yu-Chiang Wang, Chieh-Ju Chao, Monica S. Lam
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB, cs.MA, cs.SC
arXiv:2604.08849v3 Announce Type: replace-cross Abstract: Many real-world retrieval and matching problems require more than topical relevance: a candidate must satisfy the specific constraints of one profile among many, not just be relevant to it. Clinical trials are a high-stakes instance of this c...
298. PinpointQA: A Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos ​
Author: Zhiyu Zhou, Peilin Liu, Ruoxuan Zhang, Luyang Zhang, Cheng Zhang, Hongxia Xie, Wen-Huang Cheng
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.08991v3 Announce Type: replace-cross Abstract: Reliable embodied interaction in indoor environments requires agents to precisely localize small everyday objects from visual observations. Yet this fundamental capability remains challenging for multimodal large language models (MLLMs), part...
299. InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis ​
Author: Oliver Bentham, Vivek Srikumar
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.13201v2 Announce Type: replace-cross Abstract: Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging. Benchmarks derived from published studies and human annotations inherit publication bias, known-knowl...
300. From Local to Cluster: A Unified Framework for Causal Discovery with Latent Variables ​
Author: Zongyu Li
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.22416v3 Announce Type: replace-cross Abstract: Latent variables pose a fundamental obstacle to both causal discovery and inference. Local approaches exploiting direct neighborhood relations provide little beyond immediate dependencies. Cluster-level methods, though capable of broader reas...
301. Language corpora for the Dutch medical domain ​
Author: B. van Es
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.25374v2 Announce Type: replace-cross Abstract: Background: Dutch medical corpora are scarce, limiting NLP development. Methods: We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources. Results: The resulting corpus comprises +...
302. Progressive Semantic Communication for Efficient Edge-Cloud Vision-Language Models ​
Author: Cyril Shih-Huan Hsu, Wig Yuan-Cheng Cheng, Chrysa Papagianni
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, cs.DC, cs.NI
arXiv:2604.26508v2 Announce Type: replace-cross Abstract: Deploying Vision-Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, which exceed the capabilities of resource-constrained embedded platforms. Conversely, fully offloading infe...
303. Proteo-R1: Reasoning Foundation Models for De Novo Protein Design ​
Author: Fang Wu, Weihao Xuan, Heli Qi, Hanqun Cao, Heng-Jui Chang, Zeqi Zhou, Haokai Zhao, Ma Jian, Carl Ma, Yu-Chi Cheng, Kuan Pang, Xiangru Tang, Zehong Wang, Guanlue Li, Hanchen Wang, Kejun Ying, Pan Lu, Chiho Im, Seungju Han, Peng Xia, Tinson Xu, Yinxi Li, Deyao Zhu, Pheng-Ann Heng, Naoto Yokoya, Masashi Sugiyama, Li Erran Li, Jure Leskovec, Yejin Choi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE
arXiv:2605.02937v2 Announce Type: replace-cross Abstract: Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize molecular geometries without explicitly reasoning about which residues or interacti...
304. Field-Localized Forgery Detection for Digital Identity Documents ​
Author: Abhishek Kumar, Riya Tapwal, Carsten Maple, Mark Hooper
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2605.09089v2 Announce Type: replace-cross Abstract: Digital onboarding and eKYC systems used by banks, fintech platforms, telecom providers, and other third-party services commonly verify users by comparing an uploaded identity document with a selfie or live facial capture. This workflow is co...
305. Access Timing as Scaffolding: A Reinforcement Learning Approach to GenAI in Education ​
Author: Janne Rotter, Pau Benazet i Montobbio, Davinia Hern'andez-Leo
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2605.15850v3 Announce Type: replace-cross Abstract: In recent years, generative AI (GenAI) in educational settings has become ubiquitous in university students' daily lives, despite its potential to induce over-reliance, metacognitive disengagement, and diminished learning when used unrestrict...
306. Grounded Post-Training with Hard Examples for Reducing Hallucination in Multimodal Large Language Models ​
Author: Qinwu Xu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.DB, cs.LG
arXiv:2605.16411v3 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization u...
307. Why Do Safety Guardrails Degrade Across Languages? ​
Author: Max Zhang, Ameen Patel, Sang T. Truong, Sanmi Koyejo
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2605.17173v2 Announce Type: replace-cross Abstract: Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds several safety-driving factors into one, obscuring the specific cause(s) of safety failure....
308. The Matching Principle: When Does a Training Penalty Cover Deployment Shift? ​
Author: Vishal Rajput
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2605.22800v3 Announce Type: replace-cross Abstract: Ordinary training optimises the task loss and then stops. It never pays for internal representation energy: Jacobians can stay large in directions that never helped the label, so even small label-preserving noise throws the model off---a desi...
309. Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness ​
Author: Manish Aryal, Faiyaz Azam, Agnivo Banerjee, Syed Mahir Ahamed, Sai Sidhanth Manoharan Jayanthi, Allegra Laro, Cl'ement Legentilhomme, Andrew Lin, Florian Lorkowski, Marina P'erez del Valle, Radman Rakhshandehroo, Patric Rommel, Emanuel Ruzak, Nathan Theng, Paul Yushin Rapoport
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.23146v3 Announce Type: replace-cross Abstract: Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. This assumption breaks down in non-realizable settings where other actors might anticipate the agent's...
310. LVCG: Learning ECG Representations in the Latent Vectorcardiogram Space ​
Author: Bosong Huang, Panzhen Zhao, Zengxiang Li, Patricia Lee, Wei Jin, Alan Wee-Chung Liew, Ming Jin, Shirui Pan
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.31249v2 Announce Type: replace-cross Abstract: Electrocardiography (ECG) is a cornerstone of cardiac assessment, making the learning of informative ECG representations fundamental to tasks ranging from disease diagnosis to clinical report generation. However, existing methods operate almo...
311. Flow-Based Generative Modeling for Optimizing Sampling Policies in Compressed Sensing Applications ​
Author: Roman Pavelkin, Luis A. Zavala-Mondragon, Christiaan G. A. Viviers, Fons van der Sommen
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.00078v2 Announce Type: replace-cross Abstract: Numerous modern applications in signal processing and medical imaging necessitate acquiring high-dimensional signals under tight resource constraints. Traditional sampling theory suggests that accurate signal reconstruction requires a number ...
312. On Effectiveness and Efficiency of Agentic Tool-calling and RL Training ​
Author: Tong Liu, Cheng Qian, Matej Cief, Yuan He, Daniele Dan, Nikolaos Aletras, Gabriella Kazai
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.00135v2 Announce Type: replace-cross Abstract: Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge. This paper studies tool-calling along two complementary axes: effectiveness, i.e., how this capabil...
313. Where Flow Matching Leaks: Characterising Membership Signals Along the Interpolation Path ​
Author: Thomas Sesmat, Gabriel Meseguer-Brocal, Geoffroy Peeters
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SD
arXiv:2606.07271v3 Announce Type: replace-cross Abstract: Understanding memorization in generative models remains challenging, with implications for copyright and privacy. Beyond verbatim reproduction, models can encode subtler traces of their training data that never surface in their outputs yet re...
314. Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents ​
Author: Haochang Hao, Dehai Min, Zhifang Zhang, Yunbei Zhang, Miao Xu, Yingqiang Ge, Lu Cheng
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2606.07943v2 Announce Type: replace-cross Abstract: Agent skills extend general-purpose agents, but their open format enables skill poisoning: a tampered skill can make an agent run an attacker's command while completing the user's legitimate task. Invocation alone is insufficient; the attack-...
315. Anomaly Detection and Root Cause Analysis for Microservice Systems ​
Author: Luan Pham
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2606.09942v3 Announce Type: replace-cross Abstract: Microservice systems are widely used to build cloud applications, yet their complexity makes failures inevitable, degrading user experience and causing economic loss. Automated anomaly detection and root cause analysis (RCA) are now active re...
316. Time-Series Foundation Model Embeddings for Remaining Useful Life Estimation ​
Author: Amir El-Ghoussani, Michele De Vita, Ronald Naumann, Vasileios Belagiannis
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.11990v3 Announce Type: replace-cross Abstract: Remaining Useful Life (RUL) prediction is essential for industrial predictive maintenance, yet many learning-based approaches rely on extensive feature engineering or large labeled datasets to train task-specific sequence models. In this work...
317. Market Design for AI: Beyond the Copyright Binary ​
Author: Yan Dai, Maryam Farboodi, Negin Golrezaei, Sepehr Shahshahani
Published: 8/12/2026, 4:00:00 AM
Categories: econ.TH, cs.AI, cs.GT, cs.LG, stat.ML
arXiv:2606.12260v3 Announce Type: replace-cross Abstract: How can we design a market of human-generated content for use in training AI models that both enables technological progress and preserves individual incentives for high-quality content creation? Existing approaches take polar positions: a "f...
318. DIMOS: Disentangling Instance-level Moving Object Segmentation ​
Author: Hongxiang Huang, Hongwei Ren, Xiaopeng Lin, Yulong Huang, Zeke Xie, Bojun Cheng
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.12826v2 Announce Type: replace-cross Abstract: Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, and animal tracking. Event cameras record asynchronous brightness changes, providing high temporal res...
319. A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian Prediction ​
Author: Yunhong Lou, Xihang Yue, Xinran Wei, Tianqi Deng, Linchao Zhu
Published: 8/12/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.AI
arXiv:2606.14498v2 Announce Type: replace-cross Abstract: Predicting the Kohn-Sham Hamiltonian with machine learning can accelerate density functional theory while retaining access to molecular orbitals, energy levels, and electronic-structure observables that energy-only surrogates cannot resolve. ...
320. The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages ​
Author: Miso Choi, Seonga Choi, Mincheol Kwon, Woosung Joung, Jinkyu Kim, Jungbeom Lee
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2606.15821v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, forming distinct model lineages. It remains unclear whether a fundamental behavioral link exists betwe...
321. Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure ​
Author: Moritz Schlager, Friederike Jungmann, Samuel Schmidgall, Philipp Raffler, Franziska Hartl, Eva Wende, Paula Ro{\ss}m"uller, Conrad Ketzer, Avinatan Hassidim, Dale R. Webster, Yossi Matias, Yun Liu, Daniel Rueckert, Mike Schaekermann, Paul Hager
Published: 8/12/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY
arXiv:2606.17441v2 Announce Type: replace-cross Abstract: Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and expensive user studies. However, existing approaches often lack realism and controllability, often ove...
322. When Reranking Hurts: Uncertainty-Based Gating for Few-Shot Reranking ​
Author: Orian Dabod, Amir DN Cohen, Gabriel Stanovsky
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.31087v3 Announce Type: replace-cross Abstract: Few-shot selection typically assumes that reranking retrieved examples always improves performance. We challenge this view by identifying that the expensive reranking step can in fact degrade performance. Instead, we propose \emph{Training-Fr...
323. Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks ​
Author: Yoshihiro Maruyama
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2607.03798v4 Announce Type: replace-cross Abstract: Symmetry is everywhere in nature and society. Geometric deep learning builds architectures respecting group symmetries, whereas topological deep learning organizes computation through cells, incidence relations, and local-to-global structure....
324. EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization ​
Author: Youngkil Song, Yoonjae Baek, Dongwon Kim, Inho Kim, Dongkeun Kim, Suha Kwak
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.04872v2 Announce Type: replace-cross Abstract: Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, coupling high-level reasoning with temporal grounding in a single response. To tackle this challenge, we propos...
325. MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs ​
Author: Hantao Zhang, Jinru Sui, Ed Li, Dirk Bergemann, Zhuoran Yang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.08970v3 Announce Type: replace-cross Abstract: Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce M...
326. Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference ​
Author: Soumil Mandal
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.13205v2 Announce Type: replace-cross Abstract: Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accumulated attention mass, treated here as signal energy, and keeping the heaviest. On schema-de...
327. Ablation-Corrected Evaluation of Attribution Maps in Echocardiographic Ejection-Fraction Models ​
Author: Hyunkyung Han, Min Jung Kim
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.13738v5 Announce Type: replace-cross Abstract: Attribution maps for echocardiographic ejection-fraction models are evaluated by their overlap with an expert left-ventricular annotation, compared against a chance level that is computed from an area ratio rather than measured. We measure it...
328. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning ​
Author: Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.16057v5 Announce Type: replace-cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic too...
329. KAYROS: An Anytime and Exact Open-Source Solver for Duration-Minimization Time-Dependent Vehicle Routing. A Technical Report and a Case Study in Human-AI Engineering ​
Author: Florian Rascoussier
Published: 8/12/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.MS
arXiv:2607.23116v2 Announce Type: replace-cross Abstract: Time-dependent routing recognizes that the same journey can take a different time depending on when it begins. Under duration minimization, even the departure time of each vehicle becomes a decision. Exact methods for this setting exist in th...
330. Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact ​
Author: Mateusz Koz{\l}owski
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2607.25589v2 Announce Type: replace-cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is usually assumed rather than tested. We perf...
331. Living-Harness Is an Interactive-Agent Evolver ​
Author: Yuetian Du, Yucheng Wang, He Xu, Jiexu Xu, Shanwen Tan, Bing Zhao, Boyu Yang, Zhijie Xu, Ming Kong, Hu Wei, Jie Liu, Qiang Zhu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CL
arXiv:2607.26598v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because post-episode feedback rarely revises the persistent harness that guides future in...
332. The Epistemic Politics of AI Anthropomorphism ​
Author: Donna M Bye, Levin Kuhlmann
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2608.00961v2 Announce Type: replace-cross Abstract: AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustained or relational interaction with AI are routinely pathologised or dismissed as naive, vulnerable to de...
333. dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model ​
Author: Hankun Wang, Bohan Li, Shi Lian, Xiaoyu Gu, Jing Peng, Da Zheng, Yiwei Guo, Colin Zhang, Kai Yu
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL, eess.AS
arXiv:2608.02673v2 Announce Type: replace-cross Abstract: Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface for expressing edit requests, but its ambiguity may leave the in...
334. Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't ​
Author: Ravi Satya Durga Prasad Yenugula
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.02829v2 Announce Type: replace-cross Abstract: Model families are typically trained size by size, each from scratch. Can apretrained large model instead be converted into a smaller sibling? Wecharacterize the 1.4B->410M conversion in the Pythia family end to end.Representations align stro...
335. Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models ​
Author: Zhuoheng Huang, Mukesh Singh
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.06690v2 Announce Type: replace-cross Abstract: Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different systems question: can trusted authorization determine which newly trained parameters are reachable by th...
336. MaskFlow: Precise, Consistent and Seamless Regional Image Editing ​
Author: Rui Xu, Yang Yong, Shunzi Yang, Ruihao Gong, Chengtao Lv
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.06929v2 Announce Type: replace-cross Abstract: Regional image editing has attracted considerable attention for its spatial controllability. Although instruction-based and mask-reference-based editing methods can achieve strong semantic alignment, reliable regional control remains challeng...
337. EliSeg: Verified Target Construction for Report-Grounded Abnormality Segmentation ​
Author: Chengyi Peng, Haoyu Yang, Meixing Shi, Yuxiang Cai, Yankai Jiang
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.07299v2 Announce Type: replace-cross Abstract: Radiology reports describe clinical observations but do not specify executable segmentation targets. They may contain present, negated, prior,uncertain, or irrelevant findings, while multiple valid abnormalities may coexist. Existing segmenta...
338. MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures ​
Author: Zhuoning Xu, Xiucheng Zhang, Hanjun Luo, Yingbin Jin, Yinpeng Dong, Hanan Salam
Published: 8/12/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2608.07556v2 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) decompose long-horizon tasks across supervisors and subagents, but delegated goals do not necessarily carry their original authorization boundaries. Existing safety benchmarks mainly study adversarial compromise, whi...
339. Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions ​
Author: Chenrui Fan, Yize Cheng, Ming Li, Yongyuan Liang, Tianyi Zhou, Soheil Feizi
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.07968v2 Announce Type: replace-cross Abstract: Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency constraint, mode...
340. Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution ​
Author: Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev, Nikita Dragunov, Roman Yampolskiy, Andrei Kuznetsov
Published: 8/12/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.08311v2 Announce Type: replace-cross Abstract: We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive ...
341. A Combined Feature-Based Framework for Disguise and Spoofing Detection in Face Recognition Systems ​
Author: Sangiya Pararajasingham
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CR
arXiv:2608.08521v2 Announce Type: replace-cross Abstract: Face recognition systems face two distinct, commonly-separated failure modes: spoofing, where an impostor presents a photograph or video of an authorized user, and disguise, where a legitimate user is rejected because their appearance differs...
342. SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning ​
Author: David D. Yuan, Tony Z. Zhao, Kaylee Burns, Chelsea Finn
Published: 8/12/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.09138v2 Announce Type: replace-cross Abstract: While learned robotic policies hold promise for advancing generalizable manipulation, their practical deployment is often hindered by suboptimal execution speeds. Imitation learning policies are inherently limited by hardware constraints and ...
343. Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation ​
Author: Farzan Farnia, Hossein Goli, Amin Gohari
Published: 8/12/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2608.09385v2 Announce Type: replace-cross Abstract: Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. We introd...
344. ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models ​
Author: Yilin Jiang, Xiaorong Zhu, Fei Tan, Zicheng Zhang, Kaiyi Huang, Yang Yu, Zexuan Fei, Yiming Luo, Keqian Li, Hao Hao, Guangtao Zhai, Aimin Zhou
Published: 8/12/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.09548v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accura...