arXiv cs.AI - 2026-08-06 ​
287 items collected.
1. A Long-Run Persistence Theory for AI Systems under the Redundancy-Adjusted Artificial Age Score (AAS) ​
Author: Seyma Yaman Kayadibi
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04012v1 Announce Type: new Abstract: Artificial intelligence systems are increasingly expected to operate over repeated cycles of interaction, adaptation, and update rather than through isolated one-shot outputs. This raises a fundamental theoretical question: can an AI system persist ind...
2. The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents ​
Author: Mohsen Arjmandi
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04066v1 Announce Type: new Abstract: How do you verify a long-horizon agent when its own state and self-reports are exactly what you cannot trust? We present an agent instrument built so that verification is structural rather than post-hoc. A deterministic Executive owns all belief; a lan...
3. Monte Carlo Tree Search for Table-to-Multimodal Report Generation ​
Author: Teng Lin, Zhiyang Zhang, Yuyu Luo, Nan Tang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04071v1 Announce Type: new Abstract: Automatically generating professional multimodal reports comprising both textual analysis and visual charts from structured tabular data is a critical challenge in data intelligence. Existing methods suffer from fixed linear pipelines and isolated subt...
4. FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables ​
Author: Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.04077v1 Announce Type: new Abstract: Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner deliverables. We introd...
5. FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents ​
Author: Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.04095v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing ...
6. BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding ​
Author: Yangxuan Zhou, Sha Zhao, Yuning Chen, Chen Wu, Jiquan Wang, Shijian Li, Gang Pan
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.04156v1 Announce Type: new Abstract: Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific interpretation. We term this capabi...
7. Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models ​
Author: Mario Leiva, Yue Ma, Qinru Qiu, Gerardo Simari, Paulo Shakarian
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG, cs.LO
arXiv:2608.04190v1 Announce Type: new Abstract: Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone does not recover it: combiners such as majority voting trade recall for precision and are brittle to coordinated...
8. MatrAIx: Simulating the World with 8.3 Billion Persona Agents ​
Author: Xiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu, Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang, Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon, Yilun Du, Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr, Emily Fox, Asu Ozdaglar, Dawn Song
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04205v1 Announce Type: new Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simu...
9. Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent ​
Author: St John Grimbly, Nicolas Kuske, Evert A. Boonstra, Bruce A. Bassett, Charel van Hoof, Rowan Hodson, Benjamin Rosman, Ryan Smith, Mark Solms, Jonathan P. Shock
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04232v1 Announce Type: new Abstract: Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sharpen the others. Any fixed-budget system therefore has to decide where to allocate its perceptual precision. We ...
10. The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning ​
Author: Agnese Chiatti, Michael Cochez, Cristina Cornelio, Sebastijan Dumancic, Artur d'Avila Garcez, Luis C. Lamb, Lia Morra, Mathias Niepert, Robert Peharz, Alberto Speranzon, Maarten Stol, Annette Ten Teije, Thiviyan Thanapalasingam, Frank Van Harmelen, Emile Van Krieken, Antonio Vergari, Benjie Wang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.04285v1 Announce Type: new Abstract: Neurosymbolic AI systems that integrate machine learning and symbolic reasoning are rapidly gaining attention. They complement the data-intensive statistical approaches of neural networks and language models with symbolic reasoning algorithms to functi...
11. SafeCommit: Certifying When Memory-Grounded Agents May Safely Act ​
Author: Mayur Akewar, Ravi Ranjan
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.04289v1 Announce Type: new Abstract: Long-horizon agents increasingly use persistent memory and tools to take actions with external side effects. A central failure mode is premature commitment: an agent acts before resolving whether its memory grounding is stale, conflicting, incomplete, ...
12. NeuMoSync: End-to-End Neuromodulatory Control for Plasticity and Adaptability in Continual Learning ​
Author: Seyed Roozbeh Razavi Rohani, Khashayar Khajavi, Wesley Chung, Mandana Samiei, Mo Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.04358v1 Announce Type: new Abstract: Continual learning (CL) requires models to learn tasks sequentially, yet deep neural networks often suffer from plasticity loss and poor knowledge transfer, which can impede their long-term adaptability. Drawing high-level inspiration from global neuro...
13. Improving Auto-Design of Neural PDE Solvers with a Domain-Specific Language ​
Author: Shengxin Kong, Liwen Xu, Jingwen Fu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04384v1 Announce Type: new Abstract: Neural PDE solver auto-design is fundamentally a search-space representation problem. In the space of unrestricted Python programs, valid solvers form an extremely sparse subset: most candidate programs are syntactically incorrect, semantically incompa...
14. Architectural Implications of Agentic AI Workflows ​
Author: Jirong Yang, Peizhe Liu, Chaojie Zhang, Jovan Stojkovic
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.AR, cs.OS
arXiv:2608.04458v1 Announce Type: new Abstract: Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy and present its first architectural characterization with a production study at Microsoft Azure and a controlled st...
15. CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models ​
Author: De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04509v1 Announce Type: new Abstract: Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable models must identify the trustworthy source and abstain when neither is adequate. Existing post-training objective...
16. Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness ​
Author: Haoting Qian, Qingjie Zhang, Zhicong Huang, Cheng Hong, Han Qiu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.04519v1 Announce Type: new Abstract: Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Current unlearning benchmarks include mainly single-hop questions and a narrow set of multi-hop questions...
17. What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent Skills ​
Author: Tao Li, Junfeng Liu, Qinghua Zhao, Yifan Li, Lei Wang, Bo Shao, Xuejun Liu, Linjun Shou
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04562v1 Announce Type: new Abstract: Agent skills are increasingly optimized by automated feedback loops, producing long structured artifacts whose internal value remains unclear. We study skill valuation: assigning credit to the internal units of a fixed skill, such as rules, examples, s...
18. Joint UAV Flight and Opportunistic Routing under Reinforcement Learning for Delay-Tolerant Networks ​
Author: Xiao Wang, Shun-Ren Yang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04590v1 Announce Type: new Abstract: The growing deployment of delay-tolerant networks (DTNs) has made store-carry-forward (SCF) communication indispensable under sparse connectivity. However, intermittent contacts, finite buffers, and limited message time-to-live (TTL) often give rise to...
19. Agreement Before Diversity: Verification-First Complementarity for Heterogeneous Language-Model Coordination ​
Author: Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Jie Li, Ru Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04618v1 Announce Type: new Abstract: Heterogeneous language-model ensembles expand the space of candidate responses, yet they lack a principled criterion for when a newly generated answer should supersede an already supported one. We decouple candidate headroom from replacement authority,...
20. A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing ​
Author: Zhuohang Jiang, Yuxin Chen, Yongsen Pan, Zheng Hu, Wenqi Fan, Qing Li, Hongyang Wang, Jun Wang, Wenwu Ou
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04625v1 Announce Type: new Abstract: Industrial recommendation strategy iteration heavily relies on large-scale A/B experimentation. Traditional tuning requires experts to repeatedly design strategies, configure experiments, analyze results, and adjust parameters, making the process labor...
21. AI Literacy for Legal Translation: Developing Digital Resilience ​
Author: {\L}ucja Biel
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY, cs.HC
arXiv:2608.04641v1 Announce Type: new Abstract: Generative AI is transforming legal translation by introducing opportunities alongside linguistic, technical, legal, ethical and cognitive risks. This chapter examines the implications of AI for professional legal translation and proposes an AI literac...
22. Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning ​
Author: Aaditya Mehta, Arya Shah
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04663v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning often adds social terms to individual rewards, yet the scale of those terms is usually chosen by hand. We ask whether a guilt signal can instead be calibrated from human neural and behavioural data and tra...
23. Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports ​
Author: Cristian Mascia, Roberto Pietrantuono, Daniel Rodriguez, Stefano Russo
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04697v1 Announce Type: new Abstract: Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-syste...
24. Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools ​
Author: Atul Anand, Sourav Chattaraj
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness. ...
25. When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning ​
Author: Yongxin Wang, Ruizhe Zhou, Yueling Tang, Yingying Zhu, Xuemin Zhao, Xiaojun Chang, Xiaodan Liang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.04726v1 Announce Type: new Abstract: Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet benchmarks usually place questions in text, leaving it unclear whether models use the same instruction equally well ...
26. Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings ​
Author: Agatha Duzan, Asa Cooper Stickland
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04735v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit-influence settings: setups where the prompt directly incentivizes the model to hide som...
27. EviGraph: Evidence-Guided Autonomous Research Agents ​
Author: Zhenjiang Ren, Ruiji Li, Xujing Zhang, Ziliang Pang, Shuo Ren, Jiajun Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04738v1 Announce Type: new Abstract: Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported claims and inconsistencies between research questions, experiments, results, and conclusions. We argue that this...
28. Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning ​
Author: Qiyuan Zhu, Dezhi Li, Pengyu Cheng, Tianle Chen, Jiacheng Wang, Ruijie Shen, Hao Gu, Sida Lin, Zirui Liu, Jiacheng Liu, Sirui Han
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04771v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) excel on complex tasks through long chain-of-thought (CoT) reasoning, but their lengthy intermediate steps cause severe overthinking that inflates inference cost. KV-cache compression is a common solution, yet existing rea...
29. NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment ​
Author: Yu Zhao, Jiangyu Pan, Tao Hu, Ming Yin, Fan Yang, Jiangfan Liu, Xiubo Liang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04776v1 Announce Type: new Abstract: The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existing research has made progress in collision prediction, accurately quantifying risk levels from monocular vision in...
30. Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation ​
Author: Sarthak Harne, Chinmay Karkar, Yash Pandya, Ahmed Awadallah, Akshay Nambi
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.04794v1 Announce Type: new Abstract: Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, conditioned on privileged information (PI) about the answer such as a reference solution, supplies dense per-token ...
31. ContextWeave: A Real-World Workflow Benchmark ​
Author: Bo Wang, Yuqian Yao, Enxi Wang, Luozhijie Jin, Yang Liu, Yiran Suo, Yuxuan Cai, Enyu Zhou, Yufei Gao, Honglin Guo, Tianyu Huai, Li Ji, Zhikai Lei, Bufan Li, Lizhi Lin, Jinxiu Liu, Jie Yang, Jiazheng Zhou, Maosen Zhou, Pengfang Qian, Shichun Liu, Guanshan Liu, Hao Zheng, Yunhao Yu, Hang Yan, Jihua Kang, Xinchi Chen, Xipeng Qiu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.04830v1 Announce Type: new Abstract: Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether...
32. When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit ​
Author: Ziang Wei, Minjun Yu, Zheyuan Lai, Mingjie Pang, Wei Li
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.04896v1 Announce Type: new Abstract: Defensive driving scores are useful only when they preserve distinctions between policies that observe surrounding actors and those that do not. Re-simulation benchmarks may use reference-conditioned forgiveness, under which an agent receives credit wh...
33. WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models ​
Author: Bohai Gu, Yueyang Yuan, Taiyi Wu, Dazhao Du, Jian Liu, Xiaoyi Pang, Jie Zhang, Xiaocheng Lu, Haobin Zhong, Xiaotong Zhao, Alan Zhao, Song Guo
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.04964v1 Announce Type: new Abstract: Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learning (RL) can improve these models, but they hit a verification bottleneck:...
34. Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load ​
Author: Thomas Bartz-Beielstein
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.05018v1 Announce Type: new Abstract: Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law designate as critical. Determinism, reproducibility, and auditability are engineering requirements rather than op...
35. From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking ​
Author: Shaopeng Liang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.05030v1 Announce Type: new Abstract: Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical match...
36. Item Response Theory for AI Safety ​
Author: Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandb...
37. Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite ​
Author: Xiawei Yue, Boran Wang, Xiaoqing Zhang, Shuxin Zheng, Ziwei Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.05095v1 Announce Type: new Abstract: Agents for long term reasoning require a memory that can be efficiently and effectively updated over time, as new facts and external feedback continue to arrive. Recently, graph memory has been adopted to offer structural organization for multi-hop ret...
38. ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment ​
Author: Yijun Lu, Rui Ye, Jiajun Wang, Yuwen Du, Tian Jin, Songhua Liu, Siheng Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.05102v1 Announce Type: new Abstract: Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a trajectory unif...
39. CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs ​
Author: Hung Truong Thanh Nguyen, H'el`ene Fournier, Piper Jackson, Makoto Itoh, Shannon Freeman, Rene Richard, Hung Cao
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SE
arXiv:2608.05107v1 Announce Type: new Abstract: AI-supported care planning can help clinicians, patients, caregivers, and care teams coordinate complex decisions across clinical, functional, psychosocial, and environmental needs. However, many AI systems present recommendations as fixed outputs, lim...
40. OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling ​
Author: Indraneil Paul, Falko Helm, Goran Glava\v{s}, Iryna Gurevych
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SE
arXiv:2608.05141v1 Announce Type: new Abstract: Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora, however, are dominated by books, academic articles...
41. Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning ​
Author: Boxiu Li, Zimo Wen, Yijia Fan, Junxiang Lei, Sufeng Guo, Jiaao Wu, Ruize Tang, Mukai Li, Yifei Shen, Xiaoyu Chen, Wanbo Zhang, Runjing Gu, Yifei Gao, Yuheng Wu, Xuyao Huang, Zelong Zhao, Jiachen Zhang, Shibo Hu, Hangxi Guo, Yilin Chen, Yuzhe Zhang, Fan Yang, Chuan Wen, Xian Zhang, Xuanhe Zhou, Zhijie Deng
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.05144v1 Announce Type: new Abstract: Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving r...
42. AutoProteinEngine: A Large Language Model Driven Agent Framework for Multimodal AutoML in Protein Engineering ​
Author: Yungeng Liu, Zan Chen, Yu Guang Wang, Yiqing Shen
Published: 8/6/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI
arXiv:2411.04440v1 Announce Type: cross Abstract: Protein engineering is important for biomedical applications, but conventional approaches are often inefficient and resource-intensive. While deep learning (DL) models have shown promise, their training or implementation into protein engineering rema...
43. TourSynbio-Search: A Large Language Model Driven Agent Framework for Unified Search Method for Protein Engineering ​
Author: Yungeng Liu, Zan Chen, Yu Guang Wang, Yiqing Shen
Published: 8/6/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI
arXiv:2411.06024v1 Announce Type: cross Abstract: The exponential growth in protein-related databases and scientific literature, combined with increasing demands for efficient biological information retrieval, has created an urgent need for unified and accessible search methods in protein engineerin...
44. Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models ​
Author: Prashant Kulkarni, Assaf Namer
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2503.15560v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly vulnerable to sophisticated multi-turn manipulation attacks, where adversaries strategically build context through seemingly benign conversational turns to circumvent safety measures and elicit harmful or...
45. Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't ​
Author: Ravi Satya Durga Prasad Yenugula
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.02829v1 Announce Type: cross Abstract: Model families train every size from scratch. Can a pretrained large model be converted into a smaller sibling? We characterize the 1.4B->410M conversion in the Pythia family end-to-end: (i) representations align strongly across sizes (ridge R^2=0.84...
46. RAG-Stack: Co-Optimizing RAG Serving Performance and Quality ​
Author: Haiqiang Zhang, Yuanqing Lei, Wanting Li, Tao Zhang, Wenqi Jiang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.IR
arXiv:2608.03487v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG), which augments large language model (LLM) generation with information retrieved from databases, has become a widely used approach for knowledge-intensive applications. Modern RAG systems, however, expose many con...
47. Towards a New Grammar of Reasoning for Artificial Legal Intelligence and the Mecelle as Its Semantic Protocol ​
Author: Ali Goksu, F. Gozde Kardes, Mustafa Yaylali
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.04011v1 Announce Type: cross Abstract: This article examines the enduring epistemic and methodological crisis of traditional legal practice in light of the opportunities and constraints introduced by artificial intelligence. It proposes an ontologically grounded framework termed the Mecel...
48. C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning ​
Author: Yuntao Shou, Tao Meng, Wei Ai, Keqin Li
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04013v1 Announce Type: cross Abstract: Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs. However, real-world data often suffer from missing modalities due to transmission errors or user behavior, severely degrad...
49. On Hamming-Lipschitz Type Stability of the Subdominant (Minmax) Ultrametric: Theory and Simple Proofs ​
Author: Alokendu Mazumder, Arnab Roy, Punit Rathore
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04014v1 Announce Type: cross Abstract: The subdominant (minmax) ultrametric is a canonical tree-structured summary of a dissimilarity matrix, arising equivalently as the ultrametric induced by single-linkage clustering. While its classical stability theory is usually formulated in $\ell_...
50. AI-driven Multimodal Representation Learning for Latent Mediation Structure Discovery of Socioeconomic Disadvantage, Psychosocial Factors, and Cardiometabolic Multimorbidity: Insights from the All of Us Research Program ​
Author: Cong Cao, Shuangge Ma
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG, stat.AP
arXiv:2608.04016v1 Announce Type: cross Abstract: Social disadvantage is associated with multimorbidity, but the pathways linking social conditions to disease burden remain poorly understood. We developed an AI-driven multimodal mediation framework that integrates socioeconomic, psychosocial, clinic...
51. Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming ​
Author: Zhihao Zhu, Yi Yang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.04018v1 Announce Type: cross Abstract: AI agents are increasingly embedded in organizational workflows, where they interact with external information sources and invoke digital tools to perform operational tasks. As organizations adopt such systems, a critical challenge is identifying and...
52. A Trust-region Framework for Moment Estimation ​
Author: Oluwasegun A. Somefun
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SY, eess.SP, eess.SY
arXiv:2608.04026v1 Announce Type: cross Abstract: In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment estimation mechanisms, such as \textsc{Adam}, in stochastic gradient optimization. Specifically, in this framework, the magnitude of the update step ...
53. Lindblad-Inspired Multi-Timescale Reservoir Computing with Separable Rotation and Dissipation ​
Author: Jyotiranjan Beuria, Amit Shukla
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.ET, cs.NE
arXiv:2608.04028v1 Announce Type: cross Abstract: Echo-state networks enable efficient temporal learning by fixing the recurrent dynamics and training only a linear readout. However, conventional reservoirs typically accommodate signal mixing, memory retention, and stability within a single random r...
54. NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts ​
Author: Mohammed I. Radaideh, Jeremy Moon, Andre Gala-Garza, Emma Son, Yug Shah, Majdi I. Radaideh
Published: 8/6/2026, 4:00:00 AM
Categories: cs.GR, cs.AI, cs.CV, cs.CY, cs.LG
arXiv:2608.04030v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) has transformed text-to-image synthesis, yet its ability to represent specialized engineering domains remains largely unexplored. As an exmaple in nuclear engineering, general-purpose foundation models frequent...
55. EDATracer: An Agentic Framework for Large-Scale EDA Artifact Analysis ​
Author: Phat Tieu, Sayanti Jana, Matthew DeLorenzo, Jiawen Wu, Narendran Srinivasan, Srinivas Shakkottai, Jiang Hu, Jeyavijayan Rajendran
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.MA, cs.SE
arXiv:2608.04032v1 Announce Type: cross Abstract: Modern chip design relies on electronic design automation (EDA) tools that generate large, heterogeneous artifacts, including source files, scripts, logs, netlists, and reports. Analyzing these artifacts is critical for debugging, optimization, and d...
56. CheckOne: Lightweight Fault Detection and Mitigation for Vision Transformers ​
Author: Mohammad Hasan Ahmadilivani, Sven-Markus Loorits, Jaan Raik
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AR, cs.AI
arXiv:2608.04035v1 Announce Type: cross Abstract: The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns related to hardware faults. Algorithm-Based Fault Tolerance (ABFT) methods have emerged as lightweight and symmetric protection mechanisms for...
57. Reconstructing Persistent Worlds from Narratives for Narrative-Grounded Interactive Experiences ​
Author: Yi-Chun Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.GR, cs.HC
arXiv:2608.04037v1 Announce Type: cross Abstract: Designing narrative-grounded interactive experiences remains labor-intensive because interactive content must align with the underlying world implied by the narrative. Existing approaches formulate problems such as narrative planning, scene generatio...
58. Robust and Personalized Federated Learning for Aircraft-Engine Prognostics under Benign and Adversarial Client Heterogeneity ​
Author: Chinmoy Mitra, Md. Mehedi Hasan Nipu, Mohammad Sakib Mahmood, Md. Rakibul Islam, M. F. Mridha
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2608.04045v1 Announce Type: cross Abstract: Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry without sharing raw data. This study examines two complementary challenges: benign heterogeneity, where honest o...
59. Beyond the QBER Threshold: A Temporal QBER Based Machine Learning Framework for Multi Attack Detection in BB84 QKD ​
Author: Isha, Deepak Singh, Devesh Kumar, S. K Pal, Praful Hambarde, Amit Shukla
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.04047v1 Announce Type: cross Abstract: Conventional BB84 Quantum Key Distribution (QKD) systems rely on a fixed 11% Quantum Bit Error Rate (QBER) threshold to detect eavesdropping. However, stealthy attacks can remain below this threshold while still compromising channel security. This pa...
60. Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs ​
Author: Yu Luo, Bo Dong, Wenhua Cheng, Haihao Shen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04048v1 Announce Type: cross Abstract: Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput. However, conventional quantization methods typically require a separate checkpoint for each tar...
61. AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection ​
Author: Shihao Weng, Yang Feng, Xiaofei Xie, Jiongchi Yu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.04053v1 Announce Type: cross Abstract: Prompt injection remains a critical threat to LLM agents, yet existing defenses treat each task as a self-contained problem, independent of previous encounters. In practice, user requests are often underspecified: they describe the desired outcome wi...
62. Modality Agreement- and Conflict-Aware Prototype Hypergraph Learning for Multimodal Intent Understanding ​
Author: Mohnish Raj, Suraj Kumar, Soumi Chattopadhayay, Chandranath Adak, Ayan Dutta
Published: 8/6/2026, 4:00:00 AM
Categories: cs.MM, cs.AI, cs.CL, cs.LG
arXiv:2608.04054v1 Announce Type: cross Abstract: Multimodal intent recognition requires understanding not only what textual, acoustic, and visual signals share, but also how they disagree. Such disagreement is frequently class-informative; for example, lexical positivity accompanied by incongruent ...
63. LaPrune: Controllable Differentiable Sparsity at Million Scale ​
Author: Jakub Antczak, Joanna Wojciechowicz, {\L}ukasz Struski, Jacek Tabor
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04057v1 Announce Type: cross Abstract: Top-$k$ selection determines which components of a sparse model remain active. Hard selection blocks gradients, while continuous relaxations often couple mask hardness to the selected mass. We introduce LaPrune, a mathematically exact-budget differen...
64. SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors ​
Author: Yongchao Huang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04060v1 Announce Type: cross Abstract: Joint-embedding predictive architectures learn abstract states by predicting target embeddings from context embeddings, but their transition models are typically opaque neural maps. We introduce SJEPA, a reconstruction-free JEPA framework that learns...
65. An Inline Control Architecture for Language Models in Intelligent Transportation Systems ​
Author: Narendra Kumar Dewangan, Mounira Msahli
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.04065v1 Announce Type: cross Abstract: Vehicle-to-everything (V2X) systems increasingly incorporate large language models (LLMs) for semantic tasks such as message summarization, operator assistance, and decision support at roadside units and edge nodes. Although these components are not ...
66. FBID: Adaptive Personalized Federated Learning for Robust Out-of-Distribution Attack Detection in IoT Networks ​
Author: An Khanh Bui, Cong Thanh Nguyen, Hoang-Anh Pham, Hoang Thai Dinh, Diep N. Nguyen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.04073v1 Announce Type: cross Abstract: Personalized Federated Learning (PFL) has emerged as a promising solution for intrusion detection in heterogeneous IoT environments, as it can improve local adaptation under highly Non-Independent and Identically Distributed (non-IID) data distributi...
67. Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms ​
Author: Samuel Fern'andez-Mendui~na, Amir Ziashahabi, Eduardo Pavez, Antonio Ortega, Salman Avestimehr
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, eess.SP, math.IT
arXiv:2608.04074v1 Announce Type: cross Abstract: Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throughput is bandwidth-bound. Hence, reducing the cache size can raise both decoding speed and serving capacity. The...
68. Spatiotemporal Graph Transformer for Traffic Intelligence in Edge Computing ​
Author: Laha Ale, Letian Lin, Na Cao, Zheng Ma, Peng Yu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04075v1 Announce Type: cross Abstract: Accurate traffic forecasting is essential for proactive resource management in edge computing, where service demand evolves dynamically across both space and time. In practical cellular edge systems, traffic exhibits strong spatial correlations among...
69. Out-Of-The-Loop Multi-Fidelity Bayesian Optimization ​
Author: Gustavo Sutter, Hao Wang, Luis Ricardez-Sandoval, Pascal Poupart, Agustinus Kristiadi
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04113v1 Announce Type: cross Abstract: Black-box optimization is a ubiquitous problem in science and engineering, often dealing with expensive objective functions with cheaper lower-fidelity proxies available. Multi-fidelity Bayesian optimization (MF-BO) is a principled approach to this p...
70. Interpretable Fuzzy Inference for UAV Target Tracking Using Bounding-Box Geometry ​
Author: Reza Ahmari, Ahmad Mohammadi, Vahid Hemmati, Nicholas Edmond, Hossein Z. Saghazadeh, Olusola Odeyomi, Parham Kebria, Abdollah Homaifar
Published: 8/6/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.04121v1 Announce Type: cross Abstract: Vision-based guidance of unmanned aerial vehicles (UAVs) toward unmanned ground vehicles (UGVs) supports cooperative aerial--ground robotics, but reliable continuous yaw estimation from onboard vision remains challenging because of sensing uncertaint...
71. Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering ​
Author: Haotian Xia, Zilin Xiao, Junbo Zou, Vicente Ordonez, Hanjie Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04124v1 Announce Type: cross Abstract: Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence across time. Existing methods typically rely on long textual chain-of-thought rationales, even though many questions...
72. InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion ​
Author: Alon Ziv, Harel Pogoda, Yossi Adi
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.LG
arXiv:2608.04142v1 Announce Type: cross Abstract: Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to compute aggregated statistics of clean audio samples. In this work, we p...
73. LiNC: Lightweight Noise Correction via Per-Sample Trust and Gaussian Mixture Modeling ​
Author: Abhishek Moturu, Babak Taati, Anna Goldenberg
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2608.04147v1 Announce Type: cross Abstract: Label noise is common in medical imaging datasets due to factors such as inter-rater variability, annotation errors, and ambiguous cases. This can severely undermine the reliability and clinical effectiveness of machine learning models trained using ...
74. AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering ​
Author: Zihan Fang, Yueke Zhang, Yu Huang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.HC
arXiv:2608.04148v1 Announce Type: cross Abstract: Agentic AI is increasingly used to coordinate planning, implementation, review, and testing in software development, yet it often offers limited transparency into its decisions and interactions. Many such systems also assume that users can effectivel...
75. TRNet: Topography-Guided Frequency Rectification and Structure-Aware Decoding for Multimodal Paddy Rice Segmentation ​
Author: Kaiwen Xiao, Chunlong Fu, Liping Zheng, Yanfeng Su
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04154v1 Announce Type: cross Abstract: Mapping paddy rice from very-high-resolution imagery in mountainous and hilly regions is difficult because terrain alters optical appearance and increases confusion with visually similar vegetation. We present TRNet for 0.5-m GaoJing-1 red--green--bl...
76. Visualizing Graph-to-Answer Mechanism Recovery in Materials-Science Hypothesis Generation ​
Author: Shashwat Sourav, Subhadeep Pal, Markus J. Buehler, Sanjay Das, Fiona Y. Wang, Dominik Soos, Tirthankar Ghosal
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04170v1 Announce Type: cross Abstract: AI co-scientists can generate fluent materials-science hypotheses, but fluency does not show that an answer preserves a scientifically meaningful mechanism. We present a graph-to-answer mechanism-tracing case study for Graph-PRefLexOR-8B, a Qwen3-8B ...
77. Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills ​
Author: Peichun Hua, Haoxuan Xu, Mengyuan Li
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.04192v1 Announce Type: cross Abstract: Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while keeping the underlying packages hidden. Prior work focuses on prompt injection attacks that directl...
78. Patients-like-me: A Variational LM--GNN Framework for Explainable Clinical Prediction ​
Author: Xinyu Wang, Yixuan Li, Hanwei Wu, Qincheng Lu, Chi-Kuang Yeh, Xiao-Wen Chang, Ziyang Song
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04193v1 Announce Type: cross Abstract: Language models (LMs) offer strong textual representations for electronic health records (EHRs), but they encode patient sequences in isolation and provide limited explainability. Graph neural networks (GNNs) complement LMs by incorporating inter-pat...
79. A Unified Model for Cross-Domain Clone Detection via Model Merging ​
Author: Palash R. Roy, Banani Roy, Kevin A. Schneider, Chanchal K. Roy
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.04215v1 Announce Type: cross Abstract: The growing diversity of code clone types, from syntactic copies to cross-language semantic clones to AI-generated duplicates, has created a fragmentation crisis in clone detection. Current deep learning detectors are domain specialists that degrade ...
80. Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary ​
Author: S. Ashwin Hebbar, Peiyao Sheng, Sewoong Oh, Pramod Viswanath
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04240v1 Announce Type: cross Abstract: Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to experts and non-experts a...
81. Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems ​
Author: J. de Curt`o, I. de Zarz`a
Published: 8/6/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.SY, eess.SY
arXiv:2608.04265v1 Announce Type: cross Abstract: Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems, a stronger question is whether the planning architecture remains appropriate after autonomous participants res...
82. Compass: Continuously Aligning Social Media Feeds via In-Situ Reflections ​
Author: Aadit Barua, Leijie Wang, Amy X. Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.04274v1 Announce Type: cross Abstract: Social media recommendation feeds often optimize for users' immediate impulses rather than preferences they would hold after deeper reflection. Some systems address this misalignment by incorporating users' explicit preferences via a configuration pa...
83. EA-Graph: Artifact-Anchored Verification Memory for Coding Agents under Upstream Drift ​
Author: Hwai-Jung Hsu, Cheng-Jan Chi, Hanna Everett
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.04278v1 Announce Type: cross Abstract: Coding agents increasingly work across sessions, but prose notes can preserve a conclusion without the program state that supported it. After an upstream change, a repository may still build even though earlier verification claims are no longer valid...
84. MIDAS: Multi-LLM Iterative Data-Adaptive Summarization ​
Author: Karen Lee, Dhanashree Balaram, Seojun Shon, Umair Rasheed
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA
arXiv:2608.04307v1 Announce Type: cross Abstract: Text summarization is deceptively difficult. While condensing information seems straightforward, real-world enterprise summarization of support tickets, legal documents, incident reports, and more, demands strict adherence to domain-specific guidelin...
85. Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic) ​
Author: Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi, Sanggeon Yun, Hyunwoo Oh, SungHeon Jeong, Nathaniel D. Bastian, Mahdi Imani, Mohsen Imani
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, cs.MA
arXiv:2608.04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats...
86. Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits ​
Author: Bo Xue, Ji Cheng, Haodong Jing, Hongzong Li, Shuang Qiu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04324v1 Announce Type: cross Abstract: This paper studies generalized low-rank matrix bandits with multiple prioritized objectives. At each round, the learner selects a matrix-valued arm and observes a vector-valued reward, whose components correspond to multiple objectives with different...
87. ATLAS: Adaptive Topological Learning with Abstract Successors for Continual Learning ​
Author: R. Blake Lawlor, Daniel S. Brown
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04334v1 Announce Type: cross Abstract: Contemporary model-free reinforcement learning algorithms can achieve very high performance, but have low sample efficiency and are not robust to changes in the environment. Model-based algorithms have much higher sample efficiency, but still fail wh...
88. COMPAS: Difficulty-Aware Joint Search for Optimizing Code Generation ​
Author: Jingzhi Gong, Jie M. Zhang, Gunel Jahangirova, Dong Huang, Mohammad Reza Mousavi, Mark Harman
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.04336v1 Announce Type: cross Abstract: Code generation systems make each LLM call with a model, a prompt, and decoding settings. However, existing optimization methods usually tune only part of these choices or use one fixed configuration for all tasks: global optimizers search one config...
89. Equitable System-Prompt Selection via Constrained Mixed-Strategy GroupDRO ​
Author: Mengyu Xu, Qiaoxin Yang, Zhihan Liu, Ruiyao Xu, Zachary Liu, Kezhen Chen, Chongyang Gao
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CE, stat.ML
arXiv:2608.04339v1 Announce Type: cross Abstract: Large language models are increasingly used for information seeking, yet semantically equivalent questions phrased in different ways can receive answers of considerably different quality. System prompts are widely employed to steer response behavior,...
90. iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data ​
Author: Al Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Gyawali, Gianfranco Doretto, Donald A. Adjeroh
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, stat.ML
arXiv:2608.04348v1 Announce Type: cross Abstract: Multimodal learning of images and tabular data is often impaired by ineffective representations, resulting in redundancy, dispersion, and generalization problems. To tackle this challenge, we introduce Graph-Enhanced Descriptor Sequencing (GEDS), a s...
91. HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models ​
Author: Tian Jin, Ruikang Zhang, Zefeng Zhao, Ding Luo, Jin Zeng
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.04351v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) excel at general speech understanding; however, adapting them to fine-grained tasks like Speech Emotion Recognition (SER) remains a significant bottleneck. Current Parameter-Efficient Fine-Tuning (PEFT) methods typ...
92. Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework ​
Author: Zhaoqi Wang, Daqing He, Zijian Zhang, Ye Liu, Jiamou Liu, Zhirui Zeng, Zhan Qin, Zhen Li, Xin Li, Hongwei Yao, Jincheng An, Yong Liu, Yi Li, Qi Sun, Xiulei Liu, Liehuang Zhu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.04366v1 Announce Type: cross Abstract: While retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vulnerabilities to knowledge corruption attacks. Adversaries exploit these vulnerabilities by poisoning documents...
93. FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation ​
Author: Yinghao Tang, Tan Zhenwei, Yiyao Wang, Wanli Gu, Xiaolu Zhang, Jun Zhou, Wei Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04374v1 Announce Type: cross Abstract: Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institutional delivery. We introduce FinReportBench, an expert-grounded benchmark for measuring and improving instituti...
94. Towards Trustworthy Hypergraph Neural Networks under Label Noise ​
Author: Mengyao Zhou, Zhiheng Zhou, Xiao Han, Guiying Yan
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04377v1 Announce Type: cross Abstract: Hypergraph neural networks (HGNNs) have demonstrated remarkable capabilities in processing complex higher-order relationships. However, their performance is highly dependent on labeled data, making them vulnerable to label noise. Despite advances in ...
95. Image Classification Using CNN-QNN Hybrid Model with Optimized Correlated Features ​
Author: Minseo Seong, Youngwook Kim
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04379v1 Announce Type: cross Abstract: We propose a method to optimize the correlation among convolutional neural network (CNN) features that are used as inputs to quantum neural network (QNN) to enhance image classification accuracy. Unlike prior approaches that employ orthogonal decompo...
96. NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning ​
Author: Tinghe Zhang, Jian Xu, Jiaheng Chen, Jiaxing Li, Yucheng Xiao, Qiang Wang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04381v1 Announce Type: cross Abstract: Self-supervised learning on graphs is largely shaped by contrastive methods that depend on carefully designed augmentations, and by generative methods that reconstruct node attributes in the input space. Both paradigms can entangle representations wi...
97. Approximate Multi-Objective Search Under Rulebooks ​
Author: Omar Muhammetkulyyev, Oren Salzman, Tichakorn Wongpiromsarn
Published: 8/6/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.04398v1 Announce Type: cross Abstract: Robotic planning often involves multiple objectives with complex priority relationships, such as safety, efficiency, and regulatory compliance. Rulebooks formalize these relationships, allowing partial ordering of objectives that generalizes both Par...
98. Training-Free Hashing-Based Attention via Binary Principal Components ​
Author: Daohai Yu, Zhanpeng Zeng, Keyu Chen, Wenhao Li, Zhifeng Shen, Luxi Lin, Ruizhi Qiao, Xing Sun, Rongrong Ji
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.04405v1 Announce Type: cross Abstract: Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly processing ever-growing key-v...
99. MESH: Memory-Efficient Sinkhorn Optimization for Mixture-of-Experts Training ​
Author: Masato Fujitake
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.04407v1 Announce Type: cross Abstract: Memory-efficient matrix optimizers such as Sinkhorn gradient descent remove most AdamW optimizer state for dense Transformer matrices, but direct application to Mixture-of-Experts (MoE) training is unreliable. We study this failure in a controlled 11...
100. Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation ​
Author: De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04408v1 Announce Type: cross Abstract: On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this decision as counterfactual recoverability and replay each error state...
101. SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation ​
Author: Zikun Qu, Min Zhang, Mingze Kong, Zhiwei Shang, Yikun Ban, Shuang Qiu, Zhongxiang Dai
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04419v1 Announce Type: cross Abstract: On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether un...
102. Generative Optimization for Incentivized Advertising with Global Level Constraints ​
Author: Gege Chen, Ning Luo, Hao Jiang, Da Li, Wenzheng Shu, Teng Sha, Yanxiang Zeng, Wenxin Tai, Fan Zhou, Xialong Liu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04421v1 Announce Type: cross Abstract: Incentivized advertising allocates monetary or virtual rewards to drive user engagement, where a key challenge is optimizing continuous incentive magnitudes under strict global constraints. This problem is complicated by high-frequency interactions, ...
103. MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages ​
Author: Qiongqiong Wang, Ai Ti Aw, Nancy F. Chen, Ying Lay Chiu, Yang Ding, Yingxu He, Ridong Jiang, Zhuohan Liu, Yanfeng Lu, Yi Ma, Muhammad Huzaifah, Nabilah Binte Md Johan, Nattadaporn Lertcheva, Pham Minh Duc, Sailor Hardik Bhupendra, Siti Umairah Binte Mohammad Salleh, Shuo Sun, Tarun Kumar Vangani, Jeremy H. M. Wong, Jinyang Wu, Longyin Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04433v1 Announce Type: cross Abstract: We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages. The model finetunes MERaLiON-SpeechEncoder-2, a large conformer based transformer pre-train...
104. ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation ​
Author: Yiru Dong, Richong Zhang, Fanshuang Kong, Si Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.04439v1 Announce Type: cross Abstract: Large language models (LLMs) have made notable progress in code generation, but they still struggle on challenging tasks that require sophisticated algorithms or complex implementations. Recent methods increasingly use code execution as feedback, esp...
105. D$^2$F-ReAG: Dynamic Decomposition and Filtering for Multi-Hop Reasoning-Augmented Generation ​
Author: Jiaoyang Li, Junhao Ruan, Shengwei Tang, Kaiyan Chang, Zhengtao Yu, Tong Xiao, Jingbo Zhu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04444v1 Announce Type: cross Abstract: Large language models (LLMs) often generate inaccurate answers due to their reliance on static internal knowledge. Retrieval-augmented generation (RAG) addresses this limitation by integrating external knowledge and excelling at single-hop queries. H...
106. When does training on downscaled images yield the same gradients? ​
Author: Seunghyun Ji
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04448v1 Announce Type: cross Abstract: Diffusion transformers deliver strong image generation, but their training cost grows superlinearly with resolution. Recent work justifies training or sampling at reduced resolution on a spectral premise: at high noise, a downscaled latent preserves ...
107. Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning ​
Author: Pengcheng Pan, Xinfang Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.04452v1 Announce Type: cross Abstract: High-resolution pixels and crop or zoom tools give multimodal large language models the ability to inspect an image, but they do not provide a reliable task-conditioned policy for deciding where to inspect. Q-CueGraph makes this decision explicit. It...
108. TwinIR: Coordinated Invisible Dual-Point Attacks on Online HD Map Construction ​
Author: Haibo Hu, Jianghuai Deng, Chen Tang, Yang Lou, Qian Xu, Jianping Wang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04453v1 Announce Type: cross Abstract: Online HD map construction is critical to prediction and planning in autonomous driving. We find that existing physical attacks against online map construction are limited by a cross-boundary compensation effect: after the target boundary is perturbe...
109. Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning ​
Author: Hans-Martin Will, Allen L. Brown Jr., Matthew Fuchs
Published: 8/6/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.LO
arXiv:2608.04457v1 Announce Type: cross Abstract: As "AI Scientists" emerge to drive research via the Model Context Protocol (MCP), systems relying on ephemeral scripts will fail. The sheer scale of stateful, interconnected evidence requires a machine-walkable warranty grounded in a purpose-built da...
110. Tropical Algebraic Geometry for Neuronal Representations: An Arakelov-Green Measure Based Descriptor for Graph Learning ​
Author: Yuyang Zhang, Weihan Xu, Xuehai Zhou, Shucheng Cao, Qihuang Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CG
arXiv:2608.04460v1 Announce Type: cross Abstract: The quantitative analysis of 3D neuronal morphologies requires capturing both graph topology and spatial geometry. Current message-passing Graph Neural Networks (GNNs) are bounded by the 1-Weisfeiler-Lehman (1-WL) test, limiting their ability to capt...
111. Beyond Linear Dynamics: Neural Bilinear Dynamical Models for Time Series Forecasting ​
Author: Mengzhou Gao, Huangqian Yu, Pengfei Jiao
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04471v1 Announce Type: cross Abstract: Time series in real-world applications are often generated by nonlinear dynamical systems, making accurate forecasting challenging. Existing approaches that explicitly model system dynamics typically rely on linear assumptions or Koopman-based linear...
112. EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment ​
Author: Zhenyu Yi, Jianwei Xu, Yue Hu, Zhongwei Qiu, Sijing Li, Liang Huang, Bin Lv, Ling Zhang, Yingda Xia
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.04472v1 Announce Type: cross Abstract: The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-supervised learning from uni-modal images or videos, overlooking the rich semantic knowledge contained ...
113. AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation ​
Author: Jinting Wang, Yuguang Yang, Shengyu Li, Yan Rong, Shan Yang, Xiaoda Yang, Li Liu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.04479v1 Announce Type: cross Abstract: Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions. However, determining whether generated audio faithfully satisfies complex textual instructions remains challe...
114. AFD-Ledger: Deployment Provisioning for Attention--FFN Disaggregation ​
Author: Chengyu Qiu, Xiao Fu, Fengcun Li, Yulei Qian, Yuchen Xie, Xunliang Cai, Yingdi Shan, Yongwei Wu, Mingxing Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2608.04502v1 Announce Type: cross Abstract: Attention--Feed-Forward Network (FFN) Disaggregation (AFD) is emerging as a promising architecture for serving Mixture-of-Experts (MoE) language models. While existing AFD systems improve the efficiency of disaggregated execution, they leave a deploy...
115. GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction ​
Author: Shuo Liu, Huixiang Cai, Weiru Zhang, Xiaoyi Zeng
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04504v1 Announce Type: cross Abstract: Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visual-textual cues while underestimating sparse but decision-critical contextual variables. This issue,...
116. GUARD: Grounding Uncertainty and Ablation-Based Risk Detection for Diffusion-Based VLAs ​
Author: Suhas Hegde, Jitendra Yasaswi Bharadwaj Katta
Published: 8/6/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.04510v1 Announce Type: cross Abstract: Diffusion-based vision-language-action (VLA) policies can generate plausible actions even when their predictions are weakly grounded in the visual and language evidence defining the task. We introduce GUARD, a test-time failure detection method that ...
117. CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding ​
Author: Zhenyu Yi, Qiang Hu, Zhenhao Li, Jiaxuan Zhao, Yusong Sun, Lichi Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.04515v1 Announce Type: cross Abstract: Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices. However, this slice-wise formulation produces thousands of visual tokens that burden the LLM backbone, many of which capture overlapping visual eviden...
118. A Model Merging Approach for Continual MLLM Unlearning ​
Author: Yuhang Wang, Linlin Zhang, Haoxuan Ji, Xianmin Ye, Zhenxing Niu, Haichang Gao
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04548v1 Announce Type: cross Abstract: Multimodal large language model (MLLM) unlearning methods have been proposed to remove private, sensitive, or proprietary information from well-trained models. However, most existing MLLM unlearning methods are designed for one-shot requests and fail...
119. EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks ​
Author: Pau Arnal, Khaled Denfir, Danylo Smahliuk, Amrut Avhad, Marcus A. Castro
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.04549v1 Announce Type: cross Abstract: Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated on. We dedicate more than 4,000 human expert hours to evaluate a selection of six frontier LLMs on a member of t...
120. Breadcrumbing Search Agents ​
Author: Xuebin Li, Hanqing Zhao, Siyuan Liang, Kejiang Chen, Weiming Zhang, Dacheng Tao, Nenghai Yu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.04565v1 Announce Type: cross Abstract: LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical security risk: web content retrieved during execution is untrusted, exposing agents to prompt injection and goal ...
121. PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning ​
Author: Chen Yang, Shenxiang Zeng, Haoyang Zhao, Zhouyuan Xu, Youquan He, Haoyu Li, Mingyi Deng, Jiansheng Fan, Chen Wang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04575v1 Announce Type: cross Abstract: Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language models (VLMs) often struggle to interpret these dynamics and reason reliably about future and counterfact...
122. Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders ​
Author: Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Chengpeng Fu, Yu Wang, Ming Liu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04586v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from the curse of multil...
123. EASy: Towards Efficient LLM-Based Agentic System ​
Author: Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04588v1 Announce Type: cross Abstract: Agentic systems have emerged as a promising paradigm for solving complex tasks by coordinating specialized LLM-based agents. However, most existing systems primarily optimize task success while giving limited consideration to execution efficiency und...
124. The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering ​
Author: Yuqian Fu, Tianwen Qian, Yanjun Li, Yu Li, Kunyu Peng, Xu Zheng, Yongqin Xian, Alessio Tonioni, Yanwei Fu, Xiaoling Wang, Danda Paudel, Federico Tombari, Luc Van Gool, Leyi Wu, Yifan Zhao, Jinjie Zhang, Yinchuan Li, Yingcong Chen, Zixu Li, Zhiwei Chen, Zhiheng Fu, Wenbo Wang, Yupeng Hu, Weili Guan, Liqiang Nie, Takuya Murakawa, Toru Tamaki, Yi Wen, Zhenglin Du, Zhengyang Li, Lingling Li, Licheng Jiao, Wenping Ma
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04589v1 Announce Type: cross Abstract: EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scenarios. The first EgoCross Challenge was hosted at the Third EgoVis Work...
125. When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models ​
Author: Byoungjae Min, Kennedy Edemacu, Sae-Hong Cho, Yoonhyuk Choi, Beakcheol Jang, Jong Wook Kim
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04591v1 Announce Type: cross Abstract: Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context. Yet non-observation licenses a negative answer only when evidence completely covers the query scope; otherwise, the answer should rema...
126. Rethinking Reservoir Pruning: A Dynamical Perspective for Echo State Networks ​
Author: Sudip Laudari, Puspa Raj Adhikari
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.DS
arXiv:2608.04593v1 Announce Type: cross Abstract: Echo State Networks (ESNs) offer an efficient framework for temporal prediction, but their randomly initialized reservoirs are often over-parameterized and dynamically redundant. Existing pruning methods largely rely on static connectivity or activat...
127. The Order Is the Guarantee: Verifier-Budgeted Code Deletion with Static-First Learned Proposals ​
Author: Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Han Wang, Jie Li, Ru Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.04611v1 Announce Type: cross Abstract: Frontier coding models now match or exceed strong human reference points on programming benchmarks, yet benchmark success does not imply maintainable software. Prompt-driven "vibe coding" is additive: new branches, guards, and fallbacks accumulate fa...
128. Masked diffusion enables coherent beat tracking ​
Author: Francesco Foscarin, Filip Korzeniowski, Richard Vogl
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.04624v1 Announce Type: cross Abstract: Current neural networks for beat tracking generate invalid outputs, such as consecutive downbeats and erratic tempo changes, even when these are not present in the training data. Heavy post-processing techniques can alleviate these problems, but the ...
129. DisMix: Order-Aware Mixup for Medical Imaging via Disentangling Ordinal and Non-Ordinal Features ​
Author: Dileepa Pitawela, Gustavo Carneiro, Hsiang-Ting Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04652v1 Announce Type: cross Abstract: Image mixup is a widely adopted data augmentation strategy, yet it is ill-suited for ordinal classification tasks such as medical disease grading, where labels encode a progression of severity. By indiscriminately blending disease-severity cues (ordi...
130. CSGen: A Multi-Domain Curvilinear Structure Generation Model via Hierarchical Multimodal Diffusion ​
Author: Zhe Shan, Ziming Yang, Lei Zhou, Wenwen Zhang, Cong Lin, Xia Xie
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04655v1 Announce Type: cross Abstract: Curvilinear structure analysis is an important and fundamental task in multimedia. However, the controllable generation of images with precise curvilinear structure objects remains an open challenge. To address this, we propose CSGen, a hierarchical ...
131. Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark ​
Author: Enrico Mensa, Lorenzo Zane, Calogero Jerik Scozzaro, Matteo Delsanto, Tommaso Milani, Daniele Paolo Radicioni
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04670v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed computational linguistics and achieved remarkable performance across numerous natural language processing tasks, yet significant gaps persist in understanding how these systems process culturally embedded...
132. Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports ​
Author: Haobin Li, Ping Deng, Weizhong Qian, Liang Jiang, Zhenyu Huang, Mouxing Yang, Xi Peng
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.04682v1 Announce Type: cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase. However, existing SWE benchmarks typically assume that high-quality issu...
133. Personalized Federated Sparse Adaptation of Time-Series Foundation Models ​
Author: Priyanka Nihalchandani, Naman Srivastava, Varun Ojha, Pandarasamy Arjunan
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2608.04695v1 Announce Type: cross Abstract: Federated adaptation of time-series foundation models (TSFMs) is attractive for building energy forecasting because meter data are private, distributed, and highly non-IID. However, a single parameter-sharing strategy is unlikely to serve all pretrai...
134. Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO ​
Author: Xuzheng Yang, Jun Ling, Tao Huang, Caiyan Qin, Peng Wang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04698v1 Announce Type: cross Abstract: We tackle the challenging yet underexplored task of Generalized Referring Expression Comprehension (GREC), which requires a model to localize the object described by a textual expression when it exists (positive sample) and to refuse output when it d...
135. Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification ​
Author: Maryam Gholami Shiri, Eva Tuba, Sa\v{s}o D\v{z}eroski, Tome Eftimov, Ana Nikolikj
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2608.04702v1 Announce Type: cross Abstract: Benchmarking deep learning (DL) models for multi-label classification (MLC) of remote sensing images (RSI) typically yields rankings that do not generalize beyond the evaluated datasets. In this work, we move beyond rankings by employing functional a...
136. A 6G Integrated Sensing and Communication Framework for Railway Intrusion Detection and Collision Prediction ​
Author: Ajeet Kumar Yadav, Sankaran Balasubramaniam, Aritra Chatterjee, Vinod Aduru, Yogesh Simmhan, Pandarasamy Arjunan
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.SP
arXiv:2608.04710v1 Announce Type: cross Abstract: Integrated Sensing and Communication (ISAC) combines sensing and communication to efficiently utilize wireless resources and is emerging as a key paradigm for next-generation wireless networks. By leveraging the wide bandwidth, high frequencies, and ...
137. What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend ​
Author: Shahed Masoudian, Passant Shafaei, Monorama Swain, Markus Schedl
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2608.04714v1 Announce Type: cross Abstract: Benchmark scores are reported as properties of a model, yet the inference framework used to produce them, such as HuggingFace, vLLM, or Ollama, are considered non-influential and their names and versions are almost never disclosed. In this work we in...
138. Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control ​
Author: Mahshad Rastegarmoghaddam, Davoud Nikkhouy, Shima Samadzadeh
Published: 8/6/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.RO, cs.SY
arXiv:2608.04732v1 Announce Type: cross Abstract: Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control. We develop an integrated architecture in which the uncertain...
139. PURPOSE: Poisoning Conflict Resolution in RAG via Proxy-Fact-Grounded Updates ​
Author: Zijian Wang, Yubo Zhu, Muzhi Dong, Yanjun Lou, Yisheng Li, ZiLiang Zhang, Wei Tong, Yuan Zhang, Jingyu Hua, Sheng Zhong
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.04756v1 Announce Type: cross Abstract: In Retrieval-Augmented Generation (RAG), post-retrieval conflict resolution arbitrates among noisy or contradictory retrieved passages. However, the robustness of this safeguard against knowledge poisoning has not been adequately studied. Existing bl...
140. InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval ​
Author: Tsz Ting Chung, Jiangnan Li, Jie Zhou, Mo Yu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04761v1 Announce Type: cross Abstract: Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated experience into actionable guidance. At each decision step, retrieving the right insight can help the agent pr...
141. Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models ​
Author: Houze Xu, Jizhong Li, Ziyi Ye
Published: 8/6/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.04765v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However, existing VLA models still face major challenges in long-horizon tasks: sparse expert demonstrations...
142. FUSEP: A Multi-Center Benchmark for Diverse Tasks in Early Pregnancy Fetal Ultrasound Screening ​
Author: Bin Pu, Jiewen Yang, Liwen Wang, Ying Tan, Guannan He, Xingbo Dong, Qika Lin, Jiarong Guo, Lixian Yang, Zuozhu Liu, Shengli Li, Kenli Li
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04766v1 Announce Type: cross Abstract: A large number of infants with congenital anomalies are born each year globally, especially in areas with underdeveloped medical resources. Currently, fetal ultrasound screening is the most common modality for early pregnancy anatomy detection. This ...
143. Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent ​
Author: Chenyu Wang, Yi Liu, Baoqing Li, Min Tu, Diping Song
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.04772v1 Announce Type: cross Abstract: Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are privacy-restricted. We introduce Guideline-as-Oracle (GAO), which compiles American Academy of Ophthalmology gu...
144. IMFACT: Counterfactual Explanations for Time Series via Intrinsic Mode Function Substitution ​
Author: Udo Schlegel, Julian Rakuschek, Thomas Seidl, Andreas Holzinger, Tobias Schreck, Javier Del Ser
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04777v1 Announce Type: cross Abstract: Oscillatory signals, such as vibration, carry class-discriminative information in specific frequency bands; perturbing them in raw feature space for counterfactual analysis easily destroys their temporal structure and produces physically implausible ...
145. RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists ​
Author: Yuexi Yang, Alyssa Wu, Ji Luo, Richeng Xuan, Zhichao Hu, Yuhong Liu, Zhen Qin
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.04783v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repository-scale assistance. However, existing benchmarks largely rely on bug reports from GitHub Issues, which often al...
146. Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation ​
Author: Yi Yang, Cong Qin, Xiaodan Liu, Chishui Chen, Qing Dong, Yan Zhang, Cao Liu, Zhao Yang, Lu Pan, Jiaye Lin, Yi Feng
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated. On-Policy Self-Distillation (OPSD) addresses this ...
147. Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First ​
Author: Ishaan Bhola, Adithyan Krishnan, Mukunda NS
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.04804v1 Announce Type: cross Abstract: Frontier language models can resolve repository-level software issues, but each attempt is expensive, and existing routers select a model from the issue text alone. We present SuperScout, which routes after scouting the repository: a 7B searcher, Sup...
148. Towards a satellite image manipulation and deepfake localization benchmark dataset ​
Author: Jacob Arndt, Debvrat Varshney, Philipe Dias, Nivedita Nukavarapu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.04840v1 Announce Type: cross Abstract: Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative artificial intelligence. Highly realistic synthetic imagery produced for malicious purposes (deepfakes) can have major consequences in the r...
149. A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination ​
Author: Wenxiao Zhao, Dong Liu, Kaiyi Xu, Feng Liu, Zhen Zhao, Fei Ben, Shu Wang, Wenhao Li, Yingnian Wu, Fenghua Ling, Haobo Li, Lei Bai
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.04872v1 Announce Type: cross Abstract: Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We propose A-SR, a self...
150. When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs ​
Author: Jiaming Cheng, Subhransu Das, Rajiv Ramnath
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.04893v1 Announce Type: cross Abstract: Multi-agent LLM systems relay key--value caches instead of text and credit their gains to exchanged ``latent thoughts''. That credit is a claim about \emph{which} example's cache is relayed, not merely that one is. We audit it causally in released sy...
151. A Chain Is Only as Strong as Its Weakest Link: A Scoping Review of System Integration Audits in AI ​
Author: Leah Davis, Dominic Martin, AJung Moon
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.04921v1 Announce Type: cross Abstract: As AI systems become increasingly integrated into diverse interfaces and applications, model-centric audits are insufficient to address risks arising from interactions among system components and deployment environments. System integration has long b...
152. Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning ​
Author: Xuehang Guo, Pengyuan Li, Tom Hope, Tirthankar Ghosal, Manling Li, Qingyun Wang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.04926v1 Announce Type: cross Abstract: As chart images, tabular data, and visualization code play increasingly important roles across diverse domains, cross-representation understanding across these modalities poses fundamental challenges for AI systems: the relationships across represent...
153. SVI-DAG: A Structured Variational Inference Approach to Bayesian Causal Discovery ​
Author: Shrenik Zinage
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.04930v1 Announce Type: cross Abstract: Bayesian causal discovery seeks to determine the posterior distribution of causal theories, which are interpreted as directed acyclic graphs (DAGs) that explain the observed data. The resulting posterior allows systematic reasoning regarding epistemi...
154. CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications ​
Author: Brendan Smith, Susana Lopez-Moreno, Eric Dolores-Cuenca, Sangil Kim, Jose L. Mendoza-Cortes, Nijamudheen Abdulrahiman
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.mtrl-sci, cond-mat.other, cs.AI, physics.chem-ph
arXiv:2608.04942v1 Announce Type: cross Abstract: CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications. CheMLFlow targets a common bottleneck in scientific machine learning development, where ...
155. A General Sufficient Condition for Rewriting Horn-ALCHI Atomic Queries into GQL ​
Author: David Carral, Calixte Gruson, Quentin Mani`ere
Published: 8/6/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2608.04945v1 Announce Type: cross Abstract: The emergence of the ISO standard GQL introduces a powerful query language extending first-order logic with controlled recursion, raising the question of its applicability to evaluation of ontology-mediated queries (OMQs). We focus on OMQs consisting...
156. SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models ​
Author: Sihan Hu, Lyuhan Huang, Youjin Deng, Kun Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.04975v1 Announce Type: cross Abstract: SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theory and its implementation as working numerical code. It is a component of the Artificial Analysis In...
157. Protoreasoning in Tiny Transformers ​
Author: Eduardo Valle, Fergal Reid
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.04980v1 Announce Type: cross Abstract: We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study step-by-step reasoning on ~1M-parameter models and opening up opportunities for much more detailed experimentat...
158. ORACLE: A Multi-Objective Reinforcement Learning-Based Analog Circuit Design Optimizer with Large Language Models-Guided Exploration ​
Author: Osei Brempong, Mohammed Ayman Habib, Vivan Poddar, Morteza Fayazi
Published: 8/6/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY
arXiv:2608.04999v1 Announce Type: cross Abstract: Analog circuit design automation using reinforcement learning (RL) has emerged as a promising approach for reducing manual effort. However, many existing RL-based methods focus on single-objective optimization. Even methods designed for multi-objecti...
159. OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents ​
Author: Jingsheng Zheng, Xinyuan Fang, Jintian Zhang, Zhengke Gui, Huajun Chen, Ningyu Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC, cs.LG, cs.MA
arXiv:2608.05013v1 Announce Type: cross Abstract: LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints across many steps while navigati...
160. Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems ​
Author: Isaiah Andrews
Published: 8/6/2026, 4:00:00 AM
Categories: econ.TH, cs.AI, cs.LG
arXiv:2608.05015v1 Announce Type: cross Abstract: Representation theorems in decision theory establish that behavior satisfies certain axioms if and only if it can be rationalized by a well-defined objective. I argue that this ``if and only if'' structure provides a potentially useful foundation for...
161. ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation ​
Author: Xiaoyan Gu, Yifang Wang, Wenqing Zheng, Haozhong Liu, Yixia Zheng, Peiyi Jiang, Wenjie Ning, Wei Zhang, Wei Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.05026v1 Announce Type: cross Abstract: High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reliance on culturally grounded meanings and deep contextual knowledge behind the images. Current AI-ass...
162. Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning ​
Author: Yuxuan Huang, Xingyu Zeng, Tianhang Zheng, Chaochao Lu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.05045v1 Announce Type: cross Abstract: Released aligned large language models remain vulnerable to malicious downstream finetuning. Existing defenses are largely designed for the fine-tuning-as-a-service (FTaaS) paradigm or rely on downstream users to follow additional safety procedures, ...
163. The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations ​
Author: Sandra C. Sandoval, Navita Goyal, Rashawn Ray, Long Doan, Rachel Rudinger, Hal Daum'e III
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL
arXiv:2608.05050v1 Announce Type: cross Abstract: Against the backdrop of violence in police interactions with the U.S. public, we explore how deferentially police officers speak to virtual characters depicted as Black adult males in vir- tual reality (VR) simulations. We evaluate the effect of seei...
164. MarsCast: Transfer Learning of AI Weather Foundation Models to Planetary Atmospheres ​
Author: M. L. Carroll, J. Li, S. D. Guzewich, G. Villanueva, J. A. Caraballo-Vega, M. J. Frost
Published: 8/6/2026, 4:00:00 AM
Categories: astro-ph.EP, cs.AI, cs.CV, cs.LG
arXiv:2608.05054v1 Announce Type: cross Abstract: We investigate the transferability of Earth weather foundation models to planetary atmospheres by adapting the GraphCast graph neural weather forecasting model to Mars. While GraphCast achieves state-of-the-art performance for terrestrial forecasting...
165. RepairFormer: Automated Repair of Structured Inputs Using Transformers ​
Author: Ovi Paul, Tom J King, Ali Shokri
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.05060v1 Announce Type: cross Abstract: Structured input files such as JSON, DOT, OBJ, INI, S-expression, and TinyC are widely used in software systems, but small corruptions can cause parsers to reject otherwise useful data. Repairing such inputs is important because malformed configurati...
166. Hardware Design and Security in the Era of Chiplets and LLMs ​
Author: Johann Knechtel, Ozgur Sinanoglu, Paul V. Gratz, Ramesh Karri
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.AR
arXiv:2608.05063v1 Announce Type: cross Abstract: The semiconductor industry is undergoing a dual revolution: the shift toward heterogeneous 2.5D chiplet systems and the integration of Large Language Models (LLMs) into Electronic Design Automation (EDA) flows. While these paradigms offer unprecedent...
167. Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models ​
Author: Jianru Shen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.05064v1 Announce Type: cross Abstract: Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human. We study whether verbalized confidence can su...
168. VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection ​
Author: Narges Rashvand, Ghazal Alinezhad Noghre, Shanle Yao, Gabriel Maldonado, Hamed Tabkhi
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.05069v1 Announce Type: cross Abstract: Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance footage, including changes in lighting, viewpoint, and human appearance. To mitigate visual noise and address pr...
169. MultiPathFormer: Towards a Foundation Model for Multipath Wireless Propagation ​
Author: Blessed Guda, Kayley Sze, Carlee Joe-Wong
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.SP
arXiv:2608.05076v1 Announce Type: cross Abstract: Recent advances in machine learning have enabled training of wireless foundation models, which aim to support tasks such as channel estimation, beam prediction, and localization based on wireless signals. Existing wireless foundation models typically...
170. Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection ​
Author: Ahmed Hassoon, Mark Dredze
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.05085v1 Announce Type: cross Abstract: Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, which tool to build, and when to stop. Many systems make these decisions by maximizing a myopic score such as expected information ga...
171. Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth ​
Author: Arunava Majumder, Marius Krumm, Hendrik Poulsen Nautrup, Hans J. Briegel
Published: 8/6/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG, stat.ML
arXiv:2608.05110v1 Announce Type: cross Abstract: Near-term quantum hardware limits circuit depth and often imposes geometrically local connectivity for quantum generative models, restricting the output distributions accessible to shallow unitary Born models. Introducing stochasticity into a unitary...
172. Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition ​
Author: Paritosh Parmar, Landy Lan, Hong Yang, Chen Yi, Chiat Pin Tay
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.ET, cs.HC, cs.LG
arXiv:2608.05115v1 Announce Type: cross Abstract: Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom incident recognition from CCTV-style observations. This setting remains underexplored, with limited benchmarks a...
173. Chained Recursive Language Models for Multi-Iteration Reasoning ​
Author: Purbesh Mitra, Sennur Ulukus
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IT, cs.LG, eess.SP, math.IT
arXiv:2608.05124v1 Announce Type: cross Abstract: Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer. This beco...
174. SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant ​
Author: Adel Javanmard, David P. Woodruff, Vahab Mirrokni
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2608.05127v1 Announce Type: cross Abstract: Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD, use high-dimensional geometric constructions but incur unfavorable dim...
175. OPD-V: Visual On-Policy Self-Distillation with Modality Balance ​
Author: Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.05131v1 Announce Type: cross Abstract: On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distilla...
176. Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains ​
Author: Ayoub Kirouane, Christos Petrocheilos
Published: 8/6/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL
arXiv:2608.05138v1 Announce Type: cross Abstract: Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end...
177. GRALS: GCN-Guided Redundancy-Aware Local Search for Minimum Vertex Cover ​
Author: Chanjuan Liu, Qiqi Bao, Yu Zhang, Enqiang Zhu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2503.06396v2 Announce Type: replace Abstract: The minimum vertex cover (MVC) problem seeks to identify the smallest set of vertices that cover all edges in an undirected graph. As a fundamental NP-hard combinatorial optimization problem, MVC has been widely studied due to its applications in n...
178. The Yokai Learning Environment: Tracking Beliefs Over Space and Time ​
Author: Constantin Ruhdorfer, Matteo Bortoletto, Johannes Forkel, Jakob Foerster, Andreas Bulling
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2508.12480v3 Announce Type: replace Abstract: The ability to cooperate with unknown partners is a central challenge in cooperative AI and widely studied in the form of zero-shot coordination (ZSC), which evaluates an algorithm by measuring the performance of independently trained agents when p...
179. Zero-shot reasoning for simulating scholarly peer-review ​
Author: Khalid M. Saqr
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.ET
arXiv:2510.02027v2 Announce Type: replace Abstract: Scholarly publishing requires scalable scrutiny supported by auditable evidence. This paper presents a two-component benchmark of xPeer, the peer-review simulation engine delivered through the xPeerd.com web front-end. The operational component ana...
180. Corrigibility Transformation: Constructing Goals That Accept Updates ​
Author: Rubi Hudson
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2510.15395v2 Announce Type: replace Abstract: An AI agent will learn a desired goal more effectively if it does not resist the training process, but many partially learned goals incentivize an AI to avoid further goal updates. We would like goals to be corrigible, meaning they allow changes re...
181. Calibrating Transformer Attention via Task-Space Sensitivity Feedback ​
Author: Yawei Liu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2512.20661v2 Announce Type: replace Abstract: Transformer-based pre-trained language models (PLMs) excel in text classification but suffer from attention dilution and attention sink effects, forcing models to over-focus on task-irrelevant tokens. Existing attention supervision methods rely on ...
182. XGrammar-2: Dynamic and Efficient Structured Generation Engine for Agentic LLMs ​
Author: Linzhang Li, Yixin Dong, Guanjie Wang, Ziyi Xu, Alexander Jiang, Tianqi Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.04426v4 Announce Type: replace Abstract: Modern LLM agents increasingly rely on dynamic structured generation, such as tool calling and response protocols. Unlike traditional structured generation with static structures, these workloads vary both across requests and within a request, posi...
183. MemFly: On-the-Fly Memory Optimization via Information Bottleneck ​
Author: Zhenyuan Zhang, Xianzhang Jia, Zhiqin Yang, Zhenbo Song, Wei Xue, Sirui Han, Yike Guo
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2602.07885v2 Announce Type: replace Abstract: Long-term memory enables large language model agents to tackle complex tasks through historical interactions. However, existing frameworks encounter a fundamental dilemma between compressing redundant information efficiently and maintaining precise...
184. Text2GraphQuery-Bench: A Text to Graph Query Benchmark ​
Author: Songlin Lyu, Lujie Ban, Zihang Wu, Tianqi Luo, Jirong Liu, Ayoub Moussaid, Oskar van Rest, Heng Lin, Chenhao Ma, Nan Tang, Shipeng Qi, Yongchao Liu, Zhan Qiu, Juelu Zhang, Jiajun Zheng
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2602.11745v2 Announce Type: replace Abstract: Graph models are fundamental to data analysis in domains rich with complex relationships. Unlike SQL, which benefits from a rel- atively unified standard and widespread familiarity, graph query languages are diverse (e.g., Cypher, GQL, SQL/PGQ) and...
185. SimMOF: AI agent for Automated MOF Simulations ​
Author: Jaewoong Lee, Taeun Bae, Jihan Kim
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.29152v2 Announce Type: replace Abstract: Metal-organic frameworks (MOFs) offer a vast design space, and as such, computational simulations play a critical role in predicting their structural and physicochemical properties. However, MOF simulations remain difficult to access because reliab...
186. AI Assistance Reduces Persistence and Hurts Independent Performance ​
Author: Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel A. Bakker, Rachit Dubey
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.04721v4 Announce Type: replace Abstract: People often optimize for long-term goals in collaboration: A mentor or companion doesn't just answer questions, but also scaffolds learning, tracks progress, and prioritizes the other person's growth over immediate results. In contrast, current AI...
187. Assessing and Explaining the Persuadability of Large Language Models as Legal Decision Tools ​
Author: Oisin Suttle, David Lillis
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2604.26233v3 Announce Type: replace Abstract: As Large Language Models (LLMs) are proposed as legal decision assistants, and even first-instance decision-makers, across a range of judicial and administrative contexts, it becomes essential to explore how they answer legal questions, and in part...
188. Contextual Agentic Memory is a Memo, Not True Memory ​
Author: Binyan Xu, Xilin Dai, Kehuan Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2604.27707v2 Announce Type: replace Abstract: Current agentic memory systems (vector stores, retrieval-augmented generation, scratchpads, and context-window management) do not implement memory: they implement lookup. We argue that treating lookup as memory is a category error with provable con...
189. Online Goal Recognition using Path Signature and Dynamic Time Warping ​
Author: Douglas Tesch, Nathan Gavenski, Leonardo Amado, Odinaldo Rodrigues, Felipe Meneguzzi
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.07736v2 Announce Type: replace Abstract: Online goal recognition in continuous domains poses two central challenges: efficiently encoding large trajectories and effectively comparing them. Recent work addresses these challenges by using custom state-space representations and metrics to co...
190. CogniFold: Always-On Proactive Memory via Cognitive Folding ​
Author: Suli Wang, Yiqun Duan, Yu Deng, Rundong Zhao, Dai Shi, Minghua Deng, Chen Chen, Yiqi Wang, Xinliang Zhou
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2605.13438v4 Announce Type: replace Abstract: Existing agent memory remains predominantly reactive and retrieval-based, lacking the capacity to autonomously organize experience into persistent cognitive structure. Toward genuinely autonomous agents, we introduce CogniFold, a brain-inspired "al...
191. Tree of Thoughts as a Classical Heuristic Search Problem: Formal Foundations and Design Patterns ​
Author: Guni Sharon
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2605.28566v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities, yet their standard generation process -- auto-regressive token prediction -- is inherently myopic and prone to cascading errors. To address this, the Tree-of-Thoughts...
192. Trivium: Temporal Regret as a First-Class Objective for Causal-Memory Controllers ​
Author: Edward Y. Chang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2606.04421v3 Announce Type: replace Abstract: Many agentic systems and LLM pipelines correct mistakes by optimizing outcome reward. This addresses only the what of failure; the why and when may go unlogged, allowing the same error to recur across episodes. We propose long-horizon temporal regr...
193. Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims ​
Author: Philip Quirke
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.08292v2 Announce Type: replace Abstract: Mechanistic studies often assign a component a role when removing it damages a behavior, its activation linearly encodes task information, and restoring that activation repairs the damage. We test the stronger implication: can the same activation c...
194. Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases ​
Author: Marcus J. Min, Mike He, Zhaoyu Li, Zixuan Yi, Sharad Malik, Aarti Gupta, Xujie Si, Osbert Bastani
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.PL
arXiv:2607.13292v2 Announce Type: replace Abstract: Autoformalization translates informal natural language into formal, machine-verifiable languages. While most work focuses on individual statements, real formalization efforts are inherently theory-level: they require an entire web of axioms, defini...
195. Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification ​
Author: Yanni Dong, Minghua Liu, Meilin Zhu, Xiaowei Huang, Lijun Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.16868v2 Announce Type: replace Abstract: Large Language Models often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applications. Existing uncertainty metrics such as semantic entropy capture agreement at the level of seman...
196. Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals ​
Author: Nicolas Caron, Christophe Guyeux, Hassan Noura, Maxime Coulmeau, Benjamin Aynes
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21597v2 Announce Type: replace Abstract: Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event prediction accuracy, not the operational coherence of a continuous risk signal. This work proposes ...
197. Fragility of Value under Imperfect Alignment ​
Author: Winter Cross
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.28881v2 Announce Type: replace Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imper...
198. ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction ​
Author: Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.29677v2 Announce Type: replace Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We p...
199. Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation ​
Author: William Caban
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00794v3 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims. No formal framework has yet characterized how validity degrades across the stages of these pipeline...
200. MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents ​
Author: YuFei Luo, Xiucheng Xu, Zhen Yang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.01742v2 Announce Type: replace Abstract: Long-term memory is critical for LLM agents operating over long-horizon interactions. However, several persistent limitations of existing memory systems can be traced to two recurring misalignment patterns in long-term interaction settings: Tempora...
201. HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents ​
Author: Daeyoung Roh, Donghee Han
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02009v2 Announce Type: replace Abstract: Retrieval-augmented search agents answer multi-hop questions by repeatedly issuing search queries and accumulating evidence. This creates a stopping problem: after the necessary evidence has appeared, further retrieval often adds cost, latency, and...
202. Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy ​
Author: Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.02087v2 Announce Type: replace Abstract: Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introduces challenges distinct from classical RL, with implications for in...
203. Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling ​
Author: Qinwen Wang, Jieping Luo, Aoxiang Qin, Ruoyu Zhao, Jianxiong Tang, Wei Zhang, Zhichao Lu, Luziwei Leng
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02347v2 Announce Type: replace Abstract: Recurrent linear attention models (RLAs) such as Mamba offer efficient linear-time sequence modeling as an alternative to Transformers, yet their fixed-capacity recurrent states limit long-sequence modeling. Drawing inspiration from hierarchical hu...
204. Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions ​
Author: Nicole Mitchell, Dhruv Agarwal, Maty Bohacek, Remi Denton, Roma Patel
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02491v2 Announce Type: replace Abstract: Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features can introduce longitudinal risks---cognitive, developmental and soc...
205. When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO ​
Author: Zhe Cao, Miaowen Wen, Fangjiong Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.03467v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. In GRPO, this completion-level uniformity creates structure-level skew: recurring correct solution forms accumulat...
206. PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud ​
Author: Chenghua Wang, Daliang Xu, Dongqi Cai, Duojin Sun, Hao Zhang, Haoze Qian, Huaiyuan Zhang, Jinshuo Cui, Junbo Cui, Kezhao Zhao, Longxi Gao, Mengwei Xu, Rongjie Yi, Tam Sikyuen, Tianyue Zhang, Weikai Xie, Xuanzhe Liu, Yingying Qin, Yiwen Lu, Yuan Yao, Yuezhi Zu, Yunhan Guo, Yuxin Zheng, Ziqi Guo
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.03682v2 Announce Type: replace Abstract: Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, t...
207. When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Diagnostic for Machine Collectives ​
Author: Molood Arman
Published: 8/6/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.03722v2 Announce Type: replace Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking argu...
208. On The Suitability of Differential Dataflow For Datalog Interpretation In Highly Dynamic Settings ​
Author: Bruno Rucy Carneiro Alves de Lima, Merlin Kramer, Kalmer Apinis
Published: 8/6/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.PL
arXiv:2308.04214v2 Announce Type: replace-cross Abstract: In the domain of knowledge representation and reasoning within AI, datalog engines play an ever-increasingly crucial role. The crux of their operation lies in materialization: the evaluation of a data- log program and its incorporation into a...
209. Large-Small Model Collaboration for Enhancing Edge-Deployed Small Models ​
Author: Peigen Liu, Yijiang Fan, Zixuan Xu, Yuren Mao, Longbin Lai, Ying Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2503.10367v2 Announce Type: replace-cross Abstract: Edge devices host domain-specific small language models (SLMs) with limited resources, while private clouds offer larger LLMs. We propose G-Boost, an adaptive edge-cloud framework that improves a deployed SLM's task performance without parame...
210. Curiosity-Diffuser: Curiosity Guide Diffusion Models for Reliability ​
Author: Zihao Liu, Xing Liu, Yuhang Dong, Haitao Chang, Zhengxiong Liu, Panfeng Huang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2503.14833v2 Announce Type: replace-cross Abstract: One of the bottlenecks in robotic intelligence is the instability of neural network models. This leads to risks when applying intelligence in the physical world. Specifically, imitation policy based on neural network may generate hallucinatio...
211. ZoomV: Temporal Zoom-in for Efficient Long Video Understanding ​
Author: Yuan Zhang, Junwen Pan, Rui Zhang, Xin Wan, Qizhe Zhang, Ming Lu, Qi She, Shanghang Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2504.01407v3 Announce Type: replace-cross Abstract: Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to the overwhelming number of frames and the risk of losing essential context through naive downsampling. Inspired by the way humans watch vide...
212. Review Text as a Leading Indicator of Displayed Reputation in Platform Rating Systems: Evidence from 34 U.S. Short-Term Rental Markets ​
Author: Ali Safari
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL, cs.LG
arXiv:2504.14053v2 Announce Type: replace-cross Abstract: Rating systems on accommodation platforms suffer from a familiar problem: nearly every listing displays a nearly perfect score, so the number that is supposed to separate good listings from bad ones barely varies. Whether the review text accu...
213. Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings for CLEF JOKER 2025 Task 2 ​
Author: Russell Taylor, Benjamin Herbert, Michael Sana
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.MA
arXiv:2507.06506v2 Announce Type: replace-cross Abstract: Translating wordplay across languages presents unique challenges that have long confounded both professional human translators and machine translation systems. This research proposes a novel approach for translating puns from English to Frenc...
214. Emergence of Hierarchical Emotion Organization in Large Language Models ​
Author: Maya Okawa, Bo Zhao, Eric J. Bigelow, Rose Yu, Tomer Ullman, Ekdeep Singh Lubana, Hidenori Tanaka
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2507.10599v3 Announce Type: replace-cross Abstract: As large language models (LLMs) increasingly power conversational agents, understanding how they model users' emotional states is critical for ethical deployment. Inspired by emotion wheels, i.e., a psychological framework that argues emotion...
215. Uncertainty-aware Predict-Then-Optimize Framework for Equitable Post-Disaster Power Restoration ​
Author: Lin Jiang, Dahai Yu, Rongchao Xu, Tian Tang, Guang Wang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SI
arXiv:2508.04780v2 Announce Type: replace-cross Abstract: The increasing frequency of extreme weather events, such as hurricanes, highlights the urgent need for efficient and equitable power system restoration. Many electricity providers make restoration decisions primarily based on the volume of po...
216. Arnold: A multi-task, multi-embodiment muscle transformer policy ​
Author: Boshi An, Alberto Silvio Chiappa, Merkourios Simos, Chengkun Li, Alexander Mathis
Published: 8/6/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG, q-bio.QM
arXiv:2508.18066v2 Announce Type: replace-cross Abstract: Controlling high-dimensional and nonlinear musculoskeletal models of the human body is a foundational scientific challenge. Recent machine learning breakthroughs have heralded in-silico policies that master individual skills like reaching, ob...
217. Memorization in Large Language Models in Medicine: Prevalence, Characteristics, and Implications ​
Author: Anran Li, Lingfei Qian, Mengmeng Du, Yu Yin, Yan Hu, Zihao Sun, Yihang Fu, Hyunjae Kim, Erica Stutz, Xuguang Ai, Qianqian Xie, Rui Zhu, Jimin Huang, Yifan Yang, Siru Liu, Yih-Chung Tham, Lucila Ohno-Machado, Hyunghoon Cho, Zhiyong Lu, Hua Xu, Qingyu Chen
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2509.08604v5 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated significant potential in medicine, with many studies adapting them through continued pre-training or fine-tuning on medical data to enhance domain-specific accuracy and safety. However, a key ope...
218. Dynamic Jailbreaking Attack ​
Author: Kedong Xiu, Yunhan Yang, Churui Zeng, Tianhang Zheng, Xinzhe Huang, Di Wang, Puning Zhao, Zhan Qin, Kui Ren
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2510.02422v4 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks typically optimize a fixed-length adversarial suffix toward a predefined target response with a static optimization strategy. However, this fully static formulation undermines the effectiveness, effic...
219. RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation ​
Author: Yuquan Xue, Guanxing Lu, Zhenyu Wu, Chuanrui Zhang, Bofang Jia, Zhengyi Gu, Ziwei Wang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2510.17640v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown strong manipulation capability when trained with large-scale imitation learning datasets. However, these datasets that predominantly consist of successful trajectories rarely provide the correcti...
220. Bi-Level Reinforcement Learning Pathway for Sim-to-Real Optimality ​
Author: Akhil S Anand, Shambhuraj Sawant, Paavo Parmas, Jasper Hoffmann, Dirk Reinhardt, Sebastien Gros
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2510.17709v2 Announce Type: replace-cross Abstract: Training Reinforcement Learning (RL) policies using simulation models before deployment in real-world environments is a common strategy when real-world interaction is expensive. This approach is used in sim-to-real RL and in dyna-style model-...
221. When Large Language Models Know the Table: A Framework for Assessing Data Contamination in Tabular Datasets ​
Author: Matteo Silvestri, Fabiano Veglianti, Flavio Giorgi, Fabrizio Silvestri, Gabriele Tolomei
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.20351v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly exposed to data contamination, i.e., performance gains driven by prior exposure of test datasets rather than generalization. However, in the context of tabular data, this problem is largely unexpl...
222. Neural Diversity Regularizes Hallucinations in Language Models ​
Author: Kushal Chakrabarti, Nirmal Balachundhar
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2510.20690v3 Announce Type: replace-cross Abstract: Language models continue to hallucinate despite increases in parameters, compute, and data. We propose neural diversity -- decorrelated parallel representations -- as a principled mechanism that reduces hallucination rates at fixed parameter ...
223. Reinforcement Learning and Consumption-Savings Behavior ​
Author: Brandon Kaplowitz
Published: 8/6/2026, 4:00:00 AM
Categories: econ.GN, cs.AI, cs.LG, q-fin.EC
arXiv:2510.20748v2 Announce Type: replace-cross Abstract: This paper demonstrates how reinforcement learning can explain two puzzling empirical patterns in household consumption behavior during economic downturns. I develop a model where agents use Q-learning with neural network approximation to mak...
224. MediRec: Enhancing Chinese Medication Recommendation with Explainable Clinical Reasoning ​
Author: Juntao Li, Haobin Yuan, Ling Luo, Yuanyuan Sun, Jian Wang, Hongfei Lin
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.21084v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong potential for clinical decision support through their advanced language understanding and reasoning capabilities. However, their application to Chinese clinical medication recommendation remains ...
225. FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation ​
Author: Song Jin, Shuqi Li, Shukun Zhang, Rui Yan
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2511.07322v3 Announce Type: replace-cross Abstract: While LLMs have shown great success in financial tasks like stock prediction and question answering, their application in fully automating Equity Research Report generation remains uncharted territory. In this paper, we formulate the Equity R...
226. Stabilizing Multi-Attack Adversarial Training via Bandit Optimization ​
Author: Rui Wang, Zeming Wei, Xiyue Zhang, Meng Sun
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR, cs.CV, math.OC
arXiv:2511.12265v2 Announce Type: replace-cross Abstract: Deep Neural Networks (DNNs) remain vulnerable to diverse adversarial perturbations, motivating multi-attack adversarial training (AT) for improved robustness. However, existing methods either incur prohibitive overhead by computing all attack...
227. DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models ​
Author: Cheng Yin, Yankai Lin, Wang Xu, Sikyuen Tam, Xiangrui Zeng, Zhiyuan Liu, Zhouping Yin
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO
arXiv:2511.15669v3 Announce Type: replace-cross Abstract: Does Chain-of-Thought (CoT) reasoning genuinely improve Vision Language Action (VLA) models, or does it merely add overhead? Existing CoT-VLA systems report limited and inconsistent gains, yet no prior work has rigorously diagnosed when and w...
228. Interpreting GFlowNets for Drug Discovery: What probes can and cannot show ​
Author: Amirtha Varshini A S, Duminda S. Ranasinghe, Hok Hei Tam
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.BM
arXiv:2511.19264v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) construct molecules through sequential decisions, but their internal policies remain opaque, limiting adoption in drug discovery, where chemists need interpretable rationales for proposed structures. We pr...
229. Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens ​
Author: Yiming Qin, Bomin Wei, Jiaxin Ge, Konstantinos Kallidromitis, Stephanie Fu, Trevor Darrell, XuDong Wang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2511.19418v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) excel at reasoning in linguistic space but struggle with perceptual understanding that requires dense visual perception, e.g., spatial reasoning and geometric awareness. This limitation stems from the fact that c...
230. MODEST: Multi-Optics Depth-of-Field Stereo Dataset ​
Author: Nisarg K. Trivedi, Vinayaka A. Belludi, Li-Yun Wang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV
arXiv:2511.20853v4 Announce Type: replace-cross Abstract: Training and evaluation of state-of-the-art computer vision algorithms for reliable shallow depth of field (DoF) rendering and defocus deblurring remain constrained by a persistent lack of large-scale, full-frame, high fidelity, real-image da...
231. Revisiting Generalization Across Difficulty Levels: It's Not So Easy ​
Author: Yeganeh Kordi, Nihal V. Nayak, Max Zuo, Ilana Nguyen, Stephen H. Bach
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2511.21692v3 Announce Type: replace-cross Abstract: We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curation and evaluation. Existing research is mixed regarding whether training on easier or harder data lead...
232. Feedback Loops and Code Perturbations in LLM-based Software Engineering: A Case Study on a C-to-Rust Translation System ​
Author: Martin Weiss, Jesko Hecking-Harbusch, Jochen Quante, Matthias Woehrle
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2512.02567v2 Announce Type: replace-cross Abstract: The advent of strong generative AI has a considerable impact on various software engineering tasks such as code repair, test generation, or language translation. While tools like GitHub Copilot are already in widespread use in interactive set...
233. HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize ​
Author: Kun Zhao, Guodong Liu, Hui Ji, Siyuan Dai, Pan Wang, Jifeng Song, Chenghua Lin, Liang Zhan, Haoteng Tang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.03321v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcement learning (RL) remains challenging due to heterogeneous medical supervision. Vanilla Group Relative...
234. Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation ​
Author: Jinmei Liu, Haoru Li, Zhenhong Sun, Chaofeng Chen, Yatao Bian, Bo Wang, Daoyi Dong, Chunlin Chen, Zhi Wang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.12401v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks. A fundamental limitation remain...
235. Multi-Task GRPO: Reliable LLM Reasoning Across Tasks ​
Author: Shyam Sundhar Ramesh, Xiaotong Ji, Matthieu Zimmer, Sangwoong Yoon, Zhiyong Wang, Haitham Bou Ammar, Aurelien Lucchi, Ilija Bogunovic
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2602.05547v3 Announce Type: replace-cross Abstract: RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deployment requires reliable performance across diverse tasks. A straightforward multi-task adaptation of GRPO...
236. A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents ​
Author: Wei-Chieh Huang, Weizhi Zhang, Yueqing Liang, Yuanchen Bei, Yankai Chen, Tao Feng, Xinyu Pan, Zhen Tan, Yu Wang, Tianxin Wei, Shanglin Wu, Ruiyao Xu, Liangwei Yang, Rui Yang, Wooseong Yang, Chin-Yuan Yeh, Hanrong Zhang, Haozhen Zhang, Siqi Zhu, Henry Peng Zou, Wanjia Zhao, Song Wang, Wujiang Xu, Zixuan Ke, Zheng Hui, Dawei Li, Yaozu Wu, Langzhou He, Chen Wang, Xiongxiao Xu, Baixiang Huang, Juntao Tan, Shelby Heinecke, Huan Wang, Caiming Xiong, Ahmed A. Metwally, Jun Yan, Chen-Yu Lee, Hanqing Zeng, Yinglong Xia, Xiaokai Wei, Ali Payani, Yu Wang, Haitong Ma, Wenya Wang, Chenguang Wang, Yu Zhang, Xin Eric Wang, Yongfeng Zhang, Jiaxuan You, Hanghang Tong, Xiao Luo, Xue Liu, Yizhou Sun, Wei Wang, Julian McAuley, James Zou, Jiawei Han, Philip S. Yu, Kai Shu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2602.06052v4 Announce Type: replace-cross Abstract: Research in artificial intelligence is shifting from model innovations and benchmark scores towards problem definition and rigorous real-world evaluation. As the field enters the "second half," the central challenge becomes real utility in lo...
237. Can Post-Training Transform LLMs into Causal Reasoners? ​
Author: Junqi Chen, Sirui Chen, Chaochao Lu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2602.06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts. While large language models (LLMs) show promise in this domain, their precise causal estimation capabilities are still limited, and the impact of post-...
238. RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis ​
Author: Zhen Bi, Xueshu Chen, Luoyang Sun, Yuhang Yao, Qing Shen, Jungang Lou, Cheng Deng
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR, cs.PF
arXiv:2602.11506v4 Announce Type: replace-cross Abstract: The transition toward localized intelligence through Small Language Models (SLMs) has intensified the need for rigorous performance characterization on resource-constrained edge hardware. However, objectively measuring the theoretical perform...
239. AdaCorrection: Adaptive Offset Cache Correction for Accurate Diffusion Transformers ​
Author: Dong Liu, Yanxuan Yu, Ben Lengerich, Ying Nian Wu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2602.13357v3 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art performance in high-fidelity image and video generation but suffer from expensive inference due to their iterative denoising structure. While prior methods accelerate sampling by caching ...
240. Formal Analysis and Supply Chain Security for Agentic AI Skills ​
Author: Varun Pratap Bhardwaj
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE
arXiv:2603.00195v2 Announce Type: replace-cross Abstract: 32 pages, 5 theorems with full proofs, 68 references, open-source tool: https://github.com/qualixar/skillfortify. v2: corrects the bibliography (22 entries had author lists that did not match the papers at the cited arXiv identifiers; all ver...
241. Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion ​
Author: Bohai Gu, Taiyi Wu, Dazhao Du, Jian Liu, Shuai Yang, Xiaotong Zhao, Alan Zhao, Song Guo
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.06140v2 Announce Type: replace-cross Abstract: Video object insertion is fundamental to video editing, yet existing diffusion methods often produce visually plausible but physically inconsistent results. We present Place-it-R1, an end-to-end framework for physically plausible video object...
242. Seeking Physics in Diffusion Noise ​
Author: Chujun Tang, Lei Zhong, Fangqiang Ding
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.RO
arXiv:2603.14294v3 Announce Type: replace-cross Abstract: Do video diffusion models encode signals predictive of physical plausibility? We probe intermediate denoising representations of pretrained Diffusion Transformers (DiTs) and find that physically plausible and implausible videos are partially ...
243. Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions ​
Author: Qiuchi Xiang, Haoxuan Qu, Hossein Rahmani, Jun Liu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2603.21194v2 Announce Type: replace-cross Abstract: Multi-agent discussions have been widely adopted, motivating growing efforts to develop attacks that expose their vulnerabilities. In this work, we study a practical yet largely unexplored attack scenario, the discussion-monitored scenario, w...
244. The Luna Bound Propagator for Formal Analysis of Neural Networks ​
Author: Henry LeCates, Haoze Wu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO
arXiv:2603.23878v3 Announce Type: replace-cross Abstract: The parameterized CROWN analysis, a.k.a., alpha-CROWN has emerged as a practically successful abstract interpretation method for neural network verification. However, existing implementations of alpha-CROWN are limited to Python, which compli...
245. SleepVLM: A Rule-Grounded Vision-Language Model for Auditable Sleep Staging ​
Author: Guifeng Deng, Pan Wang, Mengfan Niu, Jiquan Wang, Shuying Rao, Junyi Xie, Xi'ang Chen, Sha Zhao, Gang Pan, Wanjun Guo, Tao Li, Haiteng Jiang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2603.26738v4 Announce Type: replace-cross Abstract: Sleep staging is essential for sleep assessment and disorder diagnosis. In recent years, automatic sleep staging systems have achieved accuracy approaching that of human experts, but the black-box nature of their predictions hinders clinical ...
246. Terminal Agents Suffice for Enterprise Automation ​
Author: Patrice Bechard, Orlando Marquez Ayala, Emily Chen, Jordan Skelton, Sagar Davasam, Srinivas Sunkara, Vikas Yadav, Sai Rajeswar
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL
arXiv:2604.00073v3 Announce Type: replace-cross Abstract: There has been growing interest in building agents that can interact with digital platforms to execute meaningful enterprise tasks autonomously. Among the approaches explored are tool-augmented agents built on abstractions such as Model Conte...
247. MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding ​
Author: Junxian Wu, Chenghan Fu, Zhanheng Nie, Daoze Zhang, Bowen Wan, Wanxian Guan, Chuan Yu, Jian Xu, Bo Zheng
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, cs.IR
arXiv:2604.00513v3 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention. Although recent multimodal large language models (MLLMs) have driven significant progress in product und...
248. Generative Experiences for Digital Mental Health Interventions: Evidence from a Randomized Study ​
Author: Ananya Bhattacharjee, Michael Liut, Matthew J"orke, Diyi Yang, Emma Brunskill
Published: 8/6/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2604.07558v3 Announce Type: replace-cross Abstract: Digital mental health (DMH) tools have extensively explored personalization of interventions to users' needs and contexts. However, this personalization often targets what support is provided, not how it is experienced. Even well-matched cont...
249. Multi-Modal Learning meets Genetic Programming: Analyzing Alignment in Latent Space Optimization ​
Author: Benjamin L'eger, Kazem Meidani, Christian Gagn'e
Published: 8/6/2026, 4:00:00 AM
Categories: cs.NE, cs.AI
arXiv:2604.08324v4 Announce Type: replace-cross Abstract: Symbolic regression (SR) aims to discover mathematical expressions from data, a task traditionally tackled using Genetic Programming (GP) through combinatorial search over symbolic structures. Latent Space Optimization (LSO) methods use neura...
250. Topology-Aware Reasoning over Incomplete Knowledge Graph with Graph-Based Soft Prompting ​
Author: Shuai Wang, Xixi Wang, Yinan Yu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.12503v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have shown remarkable capabilities across various tasks but remain prone to hallucinations in knowledge-intensive scenarios. Knowledge Base Question Answering (KBQA) mitigates this by grounding generation in Knowl...
251. From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs ​
Author: Itay Itzhak, Eliya Habba, Gabriel Stanovsky, Yonatan Belinkov
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2604.14137v3 Announce Type: replace-cross Abstract: Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness. Instead, users often rely on ``vibe-testing'': informal experience-based evaluation, such as comparing models on coding tasks related to ...
252. Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models ​
Author: Danae S'anchez Villegas, Samuel Lewis-Lim, Nikolaos Aletras, Desmond Elliott
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.LG
arXiv:2604.14888v3 Announce Type: replace-cross Abstract: Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains unclear. We analyze reasoning dynamics in 18 VLMs covering instruction-tuned and reasonin...
253. Just Repair: A Minimal Denoising Network for Time Series Anomaly Detection ​
Author: Kadir-Kaan "Ozer, Ren'e Ebeling, Markus Enzweiler
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.17388v3 Announce Type: replace-cross Abstract: Time series anomaly detectors have grown steadily more complex, incorporating attention mechanisms, adversarial training, and stochastic latent variables. Yet, it is unclear how much of this machinery detection actually requires. We test this...
254. A Systematic Review and Taxonomy of Reinforcement Learning-Model Predictive Control Integration for Linear Systems ​
Author: Mohsen Jalaeian Farimani, Roya Khalili Amirabadi, Davoud Nikkhouy, Malihe Abdolbaghi, Mahshad Rastegarmoghaddam, Shima Samadzadeh, Mahdi Ghane
Published: 8/6/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.RO, cs.SY, math.OC
arXiv:2604.21030v2 Announce Type: replace-cross Abstract: The integration of Model Predictive Control (MPC) and Reinforcement Learning (RL) has emerged as a promising paradigm for constrained decision-making and adaptive control. MPC offers structured optimization, explicit constraint handling, and ...
255. DeepImagine: Clinical Trial Outcome Prediction via Stepwise Local Counterfactual Imaginations ​
Author: Youze Zheng, Jianyou Wang, Yuhan Chen, Matthew Feng, Longtian Bao, Hanyuan Zhang, Maxim Khan, Aditya K. Sehgal, Christopher D. Rosin, Umber Dube, Ramamohan Paturi
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2604.23054v2 Announce Type: replace-cross Abstract: Predicting the outcomes of prospective clinical trials remains a major challenge. Clinical trial outcomes result from complex interactions among experimental factors such as drug interventions, participant demographics, and protocols. Here, w...
256. IConFace: Fine-Grained Identity Conditioning for Reference-Aware Face Restoration ​
Author: Axi Niu, Jinyang Zhang, Senyan Qing
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2605.02814v2 Announce Type: replace-cross Abstract: Severe face degradation can remove person-specific evidence, making restoration underdetermined. A generative prior may recover a sharp, plausible face yet miss localized traits that persist across images of the same person. Same-identity ref...
257. CIDR: A Large-Scale Industrial Source Code Dataset for Software Engineering Research ​
Author: Vladislav Savenkov
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2605.12153v2 Announce Type: replace-cross Abstract: We present the Curated Industrial Developer Repository (CIDR), a large-scale dataset of real-world software repositories collected from industrial partners. The dataset comprises 4,225 repositories spanning 75 programming languages, totaling ...
258. Stable Attention Response for Reliable Precipitation Nowcasting ​
Author: Penghui Wen, Zexin Hu, Sen Zhang, Patrick Filippi, Xiaogang Zhu, Allen Benter, Thomas Bishop, Zhiyong Wang, Kun Hu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.13181v3 Announce Type: replace-cross Abstract: Precipitation nowcasting remains challenging due to the highly localized, rapidly evolving, and heterogeneous nature of atmospheric dynamics. Although recent methods increasingly adopt attention-based architectures in both unimodal and multim...
259. When Bits Break Recourse: Counterfactual-Faithful Quantization ​
Author: Chaymae Yahyati, Ismail Lamaakal, Khalid El Makkaoui, Ibrahim Ouahbi
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2605.17160v4 Announce Type: replace-cross Abstract: Model quantization is widely used to reduce memory, latency, and deployment cost, and is typically judged by whether predictive accuracy is preserved. In decision systems that provide algorithmic recourse, however, accuracy preservation is no...
260. Spectral Integrated Gradients for Coarse-to-Fine Feature Attribution ​
Author: Soyeon Kim, Seongwoo Lim, Kyowoon Lee, Jaesik Choi
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2605.19607v2 Announce Type: replace-cross Abstract: Integrated Gradients (IG) is a widely adopted feature attribution method that satisfies desirable axiomatic properties. However, the choice of integration path significantly affects the quality of attributions, and the standard straight-line ...
261. Distributionally Robust Transfer Learning with Structurally Missing Covariates, with Application to Cross-National Cardiac Arrest Prediction ​
Author: Siqi Li, Chuan Hong, Ziye Tian, Benjamin Sieu-Hon Leong, Koshi Nakagawa, Hideharu Tanaka, Sang Do Shin, Khuong Quoc Dai, Do Ngoc Son, Marcus Eng Hock Ong, Nan Liu, Molei Liu
Published: 8/6/2026, 4:00:00 AM
Categories: stat.AP, cs.AI, cs.LG, stat.ML
arXiv:2605.24212v2 Announce Type: replace-cross Abstract: Deploying clinical prediction models across healthcare systems often fails when key training covariates are unavailable at deployment and labeled outcomes are limited in the target domain. For example, high-performing models for out-of-hospit...
262. VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild ​
Author: Xiaohongshu Inc
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.27882v2 Announce Type: replace-cross Abstract: LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience gap. We attribute this gap to existing benchmarks' reliance on over-specified queries, singl...
263. The Hamilton-Jacobi Theory of Deep Learning ​
Author: Jose Marie Antonio Mi~noza, Erika Fille T. Legara, Christopher P. Monterola
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.DS, math.RT, physics.comp-ph
arXiv:2605.28983v2 Announce Type: replace-cross Abstract: In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient step selects the initial data of a viscous Hamilton--Jacobi equation whose Hopf--Cole propagator best ...
264. An Enhanced Geometric-Spectral Feature Learning Framework for Airborne Multispectral Point Cloud Classification ​
Author: Xian Li, Yanfeng, Gu Linghua Xu, Aleksandra Pi\v{z}urica
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.09123v2 Announce Type: replace-cross Abstract: Multispectral point cloud (MPC) is composed of 3D spatial-spectral information, which holds tremendous potential for accurate land-cover classification. However, the representation power of classification models is limited by inherent high-di...
265. AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models ​
Author: Hui Geng, Yi Su, Zijian Gao, Tianjiao Wan, Qisheng Xu, Jiaxin Chen, Hengzhu Liu, Kele Xu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2606.14591v2 Announce Type: replace-cross Abstract: Recent advances in pretrained large audio-language models (LALMs) have demonstrated strong capabilities across speech, sound, and music. To adapt these models to downstream tasks without the cost of pretraining from scratch, post-training has...
266. Human-in-the-Loop Atlas-Based 3D Asset Segmentation for Interactive Content Workflows ​
Author: Paul Julius K"uhn, Saptarshi Neil Sinha, Jakob Hansen, Robin Horst
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.17824v2 Announce Type: replace-cross Abstract: Segmenting 3D assets into meaningful regions remains challenging, especially when segmentation criteria are application-dependent and require user control. We present a human-in-the-loop pipeline for generating a segmented 2D parameterized at...
267. Delta-Diffusion: Modeling Longitudinal Brain Amyloid-PET Trajectories via Conditional Poisson Diffusion Bridge ​
Author: Yongheng Sun, Minhui Yu, Mengqi Wu, Maureen Kohi, Mingxia Liu
Published: 8/6/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV
arXiv:2606.22216v2 Announce Type: replace-cross Abstract: While longitudinal brain PET imaging is the gold standard for quantifying the spatiotemporal accumulation of Beta-amyloid, its widespread clinical utility is constrained by high operational costs and cumulative radiation risks. Recent deep ge...
268. War in the Abstract: The Rise and Consequences of Militarized Language in Scientific Communication ​
Author: Sovesh Mohapatra, David Lydon-Staley, Dani S. Bassett
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.DL
arXiv:2606.23462v2 Announce Type: replace-cross Abstract: Scientists do not, by profession, wage war. Yet warfare's vocabulary consistently appears in their abstracts. To quantify the extent to which warfare's vocabulary pervades scientific abstracts, we analyze 21.4 million papers (2010-2025; OpenA...
269. Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks ​
Author: Yoshihiro Maruyama
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2607.03798v3 Announce Type: replace-cross Abstract: Symmetry is everywhere in nature and society. Geometric deep learning builds architectures respecting group symmetries, whereas topological deep learning organizes computation through cells, incidence relations, and local-to-global structure....
270. Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment ​
Author: Taehyung Yu, Seongjae Kang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.SD
arXiv:2607.08256v2 Announce Type: replace-cross Abstract: Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting among multiple candidates with an automatic speech recognition (ASR) verifier. We identify an evaluation confound: the apparent quality of a ver...
271. Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results ​
Author: Sheng Xu, Zhen Chen, Junhua Wang, Boyuan Huang, Ke Jia, Jiadun Zhu, Yiming Xu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.11183v3 Announce Type: replace-cross Abstract: Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plausible responses. We study inference-time feed-forward network (FFN) intervention as a way to i...
272. TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale ​
Author: Zhouchonghao Wu, Akshay Rangesh, Weixin Li, Wei-Jer Chang, Zachary Lee, Saeed Bonab, Tim Wang, Wei Zhan
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO
arXiv:2607.13028v2 Announce Type: replace-cross Abstract: Training robust autonomous driving agents requires a simulator fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that lo...
273. Decision Making Needs Uncertainty Quantification [Lecture Notes] ​
Author: Osvaldo Simeone
Published: 8/6/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, cs.LG, math.IT
arXiv:2607.14407v3 Announce Type: replace-cross Abstract: Many signal processing systems ultimately exist to {act}. Whenever the state variable that determines the action to be taken by a decision maker, or agent, is uncertain, the way that uncertainty is represented decides how well the agent perfo...
274. Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models ​
Author: Jie Zhang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21636v3 Announce Type: replace-cross Abstract: Synthetic tabular data are valued for preserving not just column-wise marginals but inter-column dependency, which carries much of the minority-class signal in domains such as fraud detection and clinical risk. Yet standard certification is l...
275. SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving ​
Author: Yihui Zhang (Beihang University), Tianyu Wo (Beihang University), Jinghao Wang (Beihang University), Xiaoyang Sun (University of Leeds), Menghao Zhang (Beihang University), Cangzhou Yuan (Beihang University), Li Li (Beihang University), Chunming Hu (Beihang University), Albert Y. Zomaya (The University of Sydney), Renyu Yang (Beihang University)
Published: 8/6/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG, cs.PF
arXiv:2607.23933v2 Announce Type: replace-cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency. Persistent...
276. CORF-GS: Real-Time Wireless Radiance Field Reconstruction via Coupled Optical-RF Gaussian Splatting ​
Author: Jinya Zhang, Jiajia Guo, Chao-Kai Wen, Shi Jin
Published: 8/6/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.CV, cs.IT, math.IT
arXiv:2607.25569v2 Announce Type: replace-cross Abstract: Recent advances in 3D Gaussian Splatting (3DGS)-based wireless radiance field (WRF) reconstruction provide an efficient solution for wireless channel modeling. However, existing WRF reconstruction methods rely on pre-collected observations an...
277. SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups ​
Author: Yikun Bai, Binghang Lu, Yikai Liu, Elaheh Akbari, Soheil Kolouri, Linxuan Wang, Ping He, Shuchan Wang, Ruqi Zhang, Guang Lin
Published: 8/6/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.27431v3 Announce Type: replace-cross Abstract: Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Existing diffusion and flow-matching models produce high-quality backbones on SE(3)^N, but inference requir...
278. ORCA-bench: How Ready Are Language Model Agents for Oncall? ​
Author: Albert Gong, Kyuseong Choi, Abhineet Agarwal, Jason Schechner, Ryan Huang, Raj Agrawal, Anish Agarwal, Raaz Dwivedi
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SE
arXiv:2607.28545v2 Announce Type: replace-cross Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours a...
279. DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing ​
Author: Sowjanya Puligadda, Mengdie Zhang, Ali Zamani, Dhruva Dixith Kurra, Eric Chen, Juan Marcano
Published: 8/6/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.28750v2 Announce Type: replace-cross Abstract: As mobile applications grow in complexity, traditional End-to-End (E2E) testing frameworks struggle with UI volatility, maintenance overhead, and cross-platform scalability. This paper presents DragonCrawl, an AI-driven mobile testing system ...
280. It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling ​
Author: Puzhuo Zheng, Hasan Kurban
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01207v2 Announce Type: replace-cross Abstract: Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe transfers poorly to vision-language models (VLMs): recent work shows that simple majority voting beats s...
281. PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation ​
Author: Xiaohan Jiang, Zeyu Li, Wei Zhang, Jiang Xu
Published: 8/6/2026, 4:00:00 AM
Categories: cs.ET, cs.AI
arXiv:2608.01791v2 Announce Type: replace-cross Abstract: The rapid development of photonic integrated circuits (PICs) is shifting the design flow from traditional graphical user interface (GUI)-based methods to script-based methods for higher flexibility, portability, and maintainability. However, ...
282. CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship ​
Author: Yao Liu, Guangjia Chai, Yuming Huang, Jihao Huang, Lei Wang, Junchen Wan
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.02046v2 Announce Type: replace-cross Abstract: LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same...
283. PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs ​
Author: Zhongjie Ba, Shengwang Xu, Peng Cheng, Jinyang Zou, Ting Yu, Zhibo Wang, Zhan Qin
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02150v3 Announce Type: replace-cross Abstract: Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understa...
284. GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience ​
Author: Sitong Gong, Caixin Kang, Tianyu Yan, Guo Chen, Bo Zheng, Kaipeng Zhang, Yunzhi Zhuge, Xiang Ruan, Huchuan Lu, Yifei Huang
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02392v2 Announce Type: replace-cross Abstract: A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situation. Existing video-memory systems primarily support question-conditioned recall, whereas proactive assi...
285. A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models ​
Author: Shu Quan, Tianfang Hao, Sitong Fang, He Geng, Jiayi Zhou, Boyuan Chen, Kaile Wang, Donghai Hong, Juntao Dai, Yaodong Yang, Jiaming Ji
Published: 8/6/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI
arXiv:2608.02684v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in protein engineering can equally be prompted to generate predicted toxin-like sequences, pot...
286. Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents ​
Author: Shuai Wang, Haodong Chen, Yu Yin, Shengyao Zhuang, Bevan Koopman, Guido Zuccon
Published: 8/6/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2608.02751v2 Announce Type: replace-cross Abstract: Existing deep-research agents use a Search--Visit workflow that retrieves whole webpages without considering the structure they expose through titles, headings, sections, and metadata. This prevents agents from directly constraining retrieval...
287. GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models ​
Author: Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski
Published: 8/6/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB
arXiv:2608.03729v2 Announce Type: replace-cross Abstract: Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess...