arXiv cs.AI - 2026-08-24 ​
290 items collected.
1. SDAD: Spec-Driven Agentic Development for the AI-Native SDLC ​
Author: Vu Hung Nguyen, Thanh Nguyen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.20341v1 Announce Type: new Abstract: Frontier coding agents backed by large language models with context windows from hundreds of thousands to millions of tokens are restructuring the Software Development Life Cycle (SDLC). Rich context handling and multi-step reasoning now allow substant...
2. PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure ​
Author: Myron Koch (Peak Summit Labs)
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.20342v1 Announce Type: new Abstract: Large language model (LLM) coding agents start each session with an empty context window, discarding accumulated knowledge from prior work. We present PrimeAgentOrchestrator (PAO), a system that spawns new instances of Claude Code -- Anthropic's termin...
3. Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification ​
Author: Md. Hasib Ur Rahman
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20378v1 Announce Type: new Abstract: Safety alignment in Large Language Models (LLMs) is often superficial, relying on refusal mechanisms that trigger only at the final stages of generation without erasing the foundational knowledge of harmful concepts acquired during pretraining. This st...
4. A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications ​
Author: Neel Mokaria, Rishie Raj, Dheeraj Baiju, Xiaoqian Shen, Shraman Pramanick, Kevin Qinghong Lin, Arda Senocak, Mike Zheng Shou, Philip Torr, Mohamed Elhoseiny, Yapeng Tian, Ruohan Gao, Salman Khan, Sayan Nag, Sanjoy Chowdhury, Dinesh Manocha
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20379v1 Announce Type: new Abstract: Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestrate perception, memory, and decision-making around powerful LLM backbone...
5. Interpretable Multimodal Classification with Linear Discriminant Tree Ensembles ​
Author: Mojtaba Moattari
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20384v1 Announce Type: new Abstract: Multimodal affect and behaviour classifiers that fuse heterogeneous text, audio, and visual streams must simultaneously achieve competitive accuracy and produce human-understandable explanations of the cues driving their decisions -- a dual objective t...
6. Representation Affects Retrieval: A Case Study of Skill Discovery and Routing in a Multimodal Agent Harness ​
Author: Kevin Dela Rosa
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20389v1 Announce Type: new Abstract: A production agent harness must discover and rank, from a growing library of skills, the one most appropriate for a user's task. At small scale this selection happens in context: the LLM planner chooses among skill representations exposed in its system...
7. Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory ​
Author: Mustafa Arslan
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20397v1 Announce Type: new Abstract: Agentic large language models (LLMs) on the Model Context Protocol (MCP) re-encode verbose tool schemas every turn, so prefill - quadratic in sequence length - dominates time-to-first-token (TTFT) as the tool registry grows. Nexus's primary lever is to...
8. Environmental Slow AI: Design Principles for Generative Systems ​
Author: Vanessa Utz
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.20398v1 Announce Type: new Abstract: Generative AI (genAI) systems produce cultural artefacts at scale, but they also reflect embedded cultural values through their design. Once identified, these values become open to deliberate reshaping. This position paper examines the maximalist value...
9. When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory ​
Author: Minkyu Song
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.20400v1 Announce Type: new Abstract: Agentic memory under a fixed budget involves two stages: retention and retrieval. Existing retrieval-centered paradigms implicitly assume necessary evidence survives eviction, but we challenge this by isolating a pre-retrieval failure mode: structurall...
10. World models of environment, agent and joint agent-environment systems ​
Author: Manuel Baltieri, Filippo Torresan, Yivan Zhang, Alexander Boyd, Fernando E. Rosas
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.20401v1 Announce Type: new Abstract: World models are a central component of model-based reinforcement learning. They are usually discussed in terms of what variables they predict, such as observations, rewards, states, latent or information states. We argue that there is a prior distinct...
11. StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models ​
Author: Michelle Lin
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.20414v1 Announce Type: new Abstract: Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct latent spatial structure from a single image remains difficult to isolate. Broad benchmarks often combine perception, optical character rec...
12. Categorical AI phenomenology: A first-person approach ​
Author: Robert Prentner
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, q-bio.NC
arXiv:2608.20420v1 Announce Type: new Abstract: This paper develops a phenomenology-first approach to artificial consciousness by reframing consciousness as the subjective experience enacted through an agent's interface with the world. We shift the methodological focus to first-person structures, mo...
13. Who Delegates to AI? Evidence from 53,000 Agent Configurations ​
Author: Hyeongjae Lee, Jihyang Cheon, Lanu Kim
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.20425v1 Announce Type: new Abstract: A growing literature measures how far occupations are exposed to AI, but these measures capture where AI could perform tasks, not whether workers have adopted it. We propose a new layer of exposure, delegated exposure, which records whether a worker ha...
14. STCO: Conditional Neural Operators for Time-Dependent PDEs ​
Author: Xingxin Yang, Zhan Zhang, Juan Li
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20477v1 Announce Type: new Abstract: Neural operators have emerged as efficient surrogates for time-dependent physical systems governed by partial differential equations (PDEs), but their future-state predictions are often conditioned only on observed states and static problem descriptors...
15. Terminal Agents: A Survey of AI Agents in Command-Line Environments ​
Author: Yi Bin, Xiaoyang Yuan, Haoxi Zeng, Wencheng Ye, Wenqi Shao, Chen Qian, Wei Ye, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Jingkuan Song, Heng Tao Shen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.20485v1 Announce Type: new Abstract: Large language model agents increasingly act through terminals, yet existing surveys disperse terminal-mediated behavior across software engineering, tool use, and computer-use research. We regard terminal agents as systems whose dominant progress-bear...
16. Lost in Translation: How Universal Ethical Values Fail to Translate Across Global Contexts ​
Author: Ozioma C. Oguine, Munachimso B. Oguine, Cesar Cervera, Jenny Yang, Pooja Voladoddi, Mario Rodriguez, Saif Eddin Bani Malhem, Karla Badillo-Urquiola, Daricia Wilkinson
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2608.20490v1 Announce Type: new Abstract: AI ethics frameworks treat values such as fairness, transparency, and accountability as universal and uniformly operationalizable across contexts. We examined how 14 experts across 10 countries made sense of AI in practice, reinterpreted core values, a...
17. A Temporal Planning Approach for Intelligent Flood Response ​
Author: Fazlul Hasan Siddiqui, Md. Monjurul Islam, Sabah Binte Noor
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20510v1 Announce Type: new Abstract: Effective response to multiple, simultaneously flooded areas requires coordinating appropriate actions in the correct temporal order, under severe resource constraints. Automated planning provides a foundation for addressing this challenge by generatin...
18. FL-MAESTRO: Multi-Agent LLM Orchestration for Resource-Constrained Federated Learning ​
Author: Jiajun Wu, Zirui Wang, Jiayu Zhou, Qiang Ye, Steve Drew
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20518v1 Announce Type: new Abstract: In Federated Learning (FL), the communication topology is a runtime variable rather than a fixed design choice, since links and edge devices drop in and out during training. Each round, the server must commit three coupled decisions, namely the communi...
19. Volumetric Radiology AI in the Era of Multimodal Large Language Models ​
Author: Zanting Ye, Shengyuan Liu, Xin Liu, Chenhui Wang, Zhisong Wang, Jiashuai Liu, Zipei Wang, Cheng Wang, Wentao Pan, Mengjie Fang, Di Dong, Mohammad Salmanpour, Arman Rahmim, Yu Gu, Yong Xia, Hongming Shan, Yixuan Yuan, Yefeng Zheng, Lijun Lu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20549v1 Announce Type: new Abstract: Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image analysis toward multimodal understanding and reasoning. Volumetric radiology, however, presents a fundamental repres...
20. Consilience: Conformally Calibrated Communication Control for Hidden-Profile Multi-Agent Reasoning ​
Author: Abhijith Babu, Ramneet Kaur, Vishal Pramanik, Olivera Kotevska, Nathaniel D. Bastian, Susmit Jha, Sunny Raj, Yanzhao Wu, Sumit Kumar Jha, Anirban Roy
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20564v1 Announce Type: new Abstract: Multi-agent LLM systems can improve reasoning by pooling diverse perspectives, but their effectiveness depends on coordinating communication, particularly in hidden-profile settings where each agent holds only part of the evidence required for a correc...
21. Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation ​
Author: Emilio Ferrara
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.20569v1 Announce Type: new Abstract: Are frontier models able to introspect about their internal states? Recent work suggests that under certain conditions a complex enough model can audit its own internals, call out what changed, and report back confidently about it. We tested that claim...
22. FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth ​
Author: Josef Chen, Erim Hayretci
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG, cs.SE
arXiv:2608.20574v1 Announce Type: new Abstract: Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable groun...
23. Difficulty-Aware Semantic-ID Optimization for Generative Recommendation ​
Author: Xin Yu, Stephen Li, Sina Aghaei, Zifan Zhu, Jiamu Bai, Guanjie Huang, Bo Peng, Yiyao Liu, Lingzhou Xue
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20611v1 Announce Type: new Abstract: Semantic-ID-based generative recommendation casts retrieval and ranking as autoregressive generation over hierarchical item identifiers. A common recipe is SFT followed by GRPO, yet vanilla GRPO is poorly matched to this tree-structured task. Under the...
24. Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills ​
Author: Christopher Kevin, Narendran Raghavan, Jean-Francois Puget, Roshni Malani, Meghana Puvvadi, Moshe Abramovitch, Mohit Gupta, Rama Akkiraju, Subodh Prabhu, Yogesh Dangi, Wei Luo, Seong Hee Lee
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20614v1 Announce Type: new Abstract: Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and workflow packages must be reviewed with evidence rather than prose. Current gates often scan these artifacts for structure, style, and security, but...
25. Dual-Cache Latent Space Communication between Heterogeneous Language Models ​
Author: Jiyao Liu, Qi Zhang, Yaoyi Jia, Ziwen Kan, Song Wang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.20617v1 Announce Type: new Abstract: Multi-agent LLM systems split work across models, so answering often requires knowledge that sits in another agent's context: a Sharer has encoded information that a Receiver needs to complete its task. They usually communicate by exchanging text, whic...
26. Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work ​
Author: George Juraj Salapa
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.20622v1 Announce Type: new Abstract: Frontier models have collapsed the cost of writing custom code: a niche problem a specialist sees in their own domain now costs an afternoon. The cost of reviewing and maintaining that code hasn't collapsed. Each solution drifts from the next; understa...
27. SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL ​
Author: Xiangqi Wang, Nhan H. Pham, Oktie Hassanzadeh, Dharmashankar Subramanian, Xiangliang Zhang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20630v1 Announce Type: new Abstract: SQL systems increasingly expose AI functions for tasks such as classification, extraction, filtering, ranking, retrieval, joining, and summarization. Despite their diverse APIs, these functions play only three relational roles: transforming individual ...
28. Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents ​
Author: Quang Dao, Purvi Kathalkar, Kenneth Eaton
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20631v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated the ability to solve multi-step tasks requiring planning, tool use, and external information access, yet growing execution histories increase inference cost and expose reasoning to outdated, irrelevan...
29. Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions ​
Author: Catherine King, Lynnette Hui Xian Ng, Kathleen M. Carley
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20649v1 Announce Type: new Abstract: Designers and policymakers in sociotechnical domains like content moderation, privacy interfaces, recommender systems and beyond, must choose among a growing menu of proposed interventions, but typically lack a principled basis for comparing them. Prio...
30. Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance ​
Author: Sergiy Lunyakin
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.CL, cs.IR
arXiv:2608.20661v1 Announce Type: new Abstract: Enterprise adoption of large language models in finance is constrained less by fluency than by trust: in Financial Planning and Analysis (FP&A) and other regulated workflows, an answer is usable only if it is traceable to authoritative sources and audi...
31. DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents ​
Author: Sarthak Singh
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.20664v1 Announce Type: new Abstract: DreamBench-SWE is a multi-session benchmark for software-agent memory hygiene in which later software tasks depend on non-inferable evidence from earlier sessions and are scored by executable hidden oracles. We report the original scaled v2 fold and a ...
32. Why2Speak: Faithful Reasoning for Abstaining Action Policies ​
Author: Shreya Mendi, Brinnae Bent
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.20670v1 Announce Type: new Abstract: Many agentic systems must repeatedly choose between acting and abstaining, making faithful reasoning important for oversight: an explanation is useful only if it reflects the computation that produced the action. We study this problem through intervent...
33. CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery ​
Author: Piyush Jha, Jake Rudolph, Victoria Knapp-P'erez, Max Fieg, Aishik Ghosh, Vijay Ganesh
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.LO, hep-ph
arXiv:2608.20686v1 Announce Type: new Abstract: Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain constraints. Reinforcement learning (RL) offers a promising approach, but existing methods rely on scalar rewards that provide limited information...
34. VortexChat: An agentic framework for autonomous multi-objective integrated photonic design ​
Author: Faqian Chong, Yulun Wu, Shilong Li, Andrew Forbes, Hongsheng Chen, Song Han
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, physics.optics
arXiv:2608.20688v1 Announce Type: new Abstract: The advancement of modern integrated photonics is frequently bottlenecked by device design workflows that rely heavily on manual simulation and expert intuition. While inverse design offers an alternative, it remains constrained by expert supervision a...
35. DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning ​
Author: Haorui Xu, Yuzhou Zhu, Liyuan Gao
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20717v1 Announce Type: new Abstract: Reliable confidence estimation is essential for using large language models in mathematical reasoning, but black-box verbalized confidence is difficult to calibrate. When the same problem is queried under multiple confidence-steering prompts, the resul...
36. Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol ​
Author: Guodong Xu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.20729v1 Announce Type: new Abstract: Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrower attribution problem of criterion revision: when criterion K0 accepts an outcome violating a broader commitment ...
37. ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation ​
Author: Siyuan Ma, Yutian Zhang, Boshi Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Xiaojin Huang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.20735v1 Announce Type: new Abstract: Manipulating moving objects requires a policy to anticipate contact events, yet vision-language-action (VLA) policies are commonly fine-tuned from the current observation alone. World action models (WAMs) learn predictive dynamics, but running a video-...
38. Continuous-Time Quantum Walks based Graph Neural Network ​
Author: Yuliang Zhan, Zefeng Gao, Jian Li, Yang Liu, Hao sun
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20738v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) are widely used on graph-structured data, but most suffer from two key weaknesses. First, message passing behaves as a low-pass filter under the homophily assumption, leading to poor performance on heterophilic graphs. Seco...
39. Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis ​
Author: Yantao Li, Huanlin Gao, Fang Zhao, Chao Tan, Qiang Hui, Shuting Liu, Fuyuan Shi, Ting Lu, Shaoan Zhao, Xueqiang Guo, Xinpei Su, Jianbing Zhang, Xinyu Dai, Kai Wang, Shiguo Lian
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20743v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work that pushes the drafter itself tow...
40. Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design ​
Author: Gyubok Lee, Kiwoong Yoo, Jimin Seo, Kyunghoon Hur, Edward Choi
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20755v1 Announce Type: new Abstract: Modern de novo design workflows generate many candidate protein binders, but wet-lab validation capacity remains limited, making shortlisting a major bottleneck. We study whether LLMs can generate multi-metric ranking policies from precomputed structur...
41. Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization ​
Author: Praphul Singh, Shanu Kumar, Akshat Agarwal
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20768v1 Announce Type: new Abstract: Specialist language models are usually understood through endpoint gains: the generalist scores lower, the specialist scores higher, and the difference is treated as evidence of specialization. This leaves the released update itself largely unexamined....
42. CAS: Conformalized Agentic Search via Adaptive Retrieval and Policy Weighting ​
Author: Zixi Zhu, Jiayuan Su, Jian Zhang, Yu Lin, Hongwei Wang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20771v1 Announce Type: new Abstract: Search Agents face a severe reliability crisis during reinforcement learning (RL) fine-tuning. Heuristic Top-K retrieval often causes critical evidence loss or noise inclusion, while over-confidence induced by progressive RL leads to hallucinated answe...
43. Structure for Reading, Prose for Writing: Asymmetric Structural Conditioning in Multi-Agent Document Authoring ​
Author: Cheng Yu, Nikhil Mathew, Zhengjie Wang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.IR
arXiv:2608.20786v1 Announce Type: new Abstract: Multi-agent pipelines that author formal documents must both read a requester's forms and write against them. We report a deployed tender-response system, running an open-weights model under sovereignty constraints, and evaluate it against human-writte...
44. Knowing but Not Saying: Preventing Factual Access Failures in LLM SFT via Recall-Anchored Distillation ​
Author: Haodong Chen, Yadong Wang, Shengtao Wen, Dong Liang, Xiang Chen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20794v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) can degrade factual behavior outside the target domain. This degradation is often described as catastrophic forgetting, yet open-ended factual failures do not necessarily imply that the underlying facts have been erased. In...
45. Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation ​
Author: Pengshuai Yang, Zijing Gao, Xue Yu, Benhui Zhuang, Bo Yuan, Junlan Feng
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20797v1 Announce Type: new Abstract: Evaluating language-guided mobile agents has recently shifted from rule-based to model-based approaches to achieve scalable and automated assessments. However, existing holistic evaluation paradigms process entire trajectories at once, leading to subst...
46. Dynamic Context Scheduling: Learning Beyond the Static Universe ​
Author: Martin Mr'az, Andr'e Biedenkapp
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20799v1 Announce Type: new Abstract: We study dynamic context scheduling as a training instrument for contextual re- inforcement learning. Rather than treating intra-episode context variation as a deployment reality, we treat it as a controlled shaping mechanism. Thereby, context evolves ...
47. SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers ​
Author: Sakif Hossain, Julian Teusch, J"org P. M"uller
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20802v1 Announce Type: new Abstract: Human motion forecasters are increasingly accurate and fast, but reliable deployment requires uncertainty estimates that are structured, calibrated, and efficient. Bayesian and ensemble-based uncertainty estimates often require repeated stochastic infe...
48. Neuro-Geospatial Modelling of EEG Affective States Using Literature-Informed Environmental Context ​
Author: Utsav Poudel, Jagannath Aryal, Subramaniyaswamy Vairavasundaram
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.LG
arXiv:2608.20807v1 Announce Type: new Abstract: Environmental exposures such as air pollution and greenness have been associated with affective and cognitive outcomes, but EEG and environmental datasets are rarely jointly georeferenced. We investigate whether literature-informed environmental priors...
49. Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence ​
Author: Yang Liu, Bin Chong, Wenkai Yang, Shuai Zhang, Yancheng Chen, Feiyu Han, GuoZhen, Cheng Zhang, Huaibing Xie, Changze Lv, Shihan Dou, Pluto Zhou
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20820v1 Announce Type: new Abstract: Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to single-turn inputs; naive multi-turn composition yields bounds that deg...
50. Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress ​
Author: Nataliya Shakhovska, Ivan Izonin, Stergios-Aristoteles Mitoulis
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20825v1 Announce Type: new Abstract: Artificial intelligence systems increasingly make consequential judgments - which patient is deteriorating, which building is safe to enter, whether an image is authentic and are trusted on the strength of how accurately and confidently they predict. T...
51. Foundation Models for Partial Causal Identification ​
Author: Alexis Bellot, Anish Dhir
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20841v1 Announce Type: new Abstract: This paper investigates the development of causal foundation models for bounding the effect of interventions and counterfactuals from observational data. We show that a canonical prior can be defined with full support over the space of structural causa...
52. TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding ​
Author: Rohan Kumar, Steven Xu, Kyle MacDonald, Matthew Long, Bernice Chow, Mac VanRenterghem, Sudeep Das
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20844v1 Announce Type: new Abstract: Product catalogs underpin search, discovery, and recommendation in e-commerce, yet they are often attribute-sparse: the attributes shoppers and downstream systems rely on are either buried in unstructured content such as titles and images or missing fr...
53. RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation ​
Author: Kyle Wild, Yusuke Takahashi, Asako Uraki
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.DB, cs.IR
arXiv:2608.20845v1 Announce Type: new Abstract: Nearly every retrieval-augmented question-answering system in production ships with a hidden interpreter: on each query a language model re-derives the meaning of raw corpus text and then throws that work away. Cheaper models do not close the gap: per-...
54. MGAL: A Multilingual Granularity-Aware Long-Context Benchmark ​
Author: Chunhan Li, Chenglin Xu, Zongyang Zhang, Jiale Liu, Zhuoxi Rao, Xudong Jia, Junxiu He, Menglin Yang, Wenjuan Gong, Zhengzhe Liu, Chengwei Qin
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20853v1 Announce Type: new Abstract: Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing benchmarks are limited to the document level and focus mainly on high-resource languages, leaving many fine-grained challenges insufficiently evaluated...
55. Coverage-Driven Verification for Safety-by-Design in AI-Based Collision Avoidance Systems ​
Author: Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank K"oster, Sven Hallerbach
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20864v1 Announce Type: new Abstract: Artificial Intelligence (AI) offers significant potential for future aviation systems; however, its integration into safety-critical applications requires compliance with the aviation sector's stringent safety standards. For AI and Machine Learning (ML...
56. ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries ​
Author: Seungheun Baek, Mogan Gim, Jaewoo Kang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.20869v1 Announce Type: new Abstract: Predicting transition states (TS) in chemical reactions is crucial, as they provide insights into reaction mechanisms. Recent work on TS prediction have focused on flow matching supervised on straight linear paths that do not align with actual reaction...
57. UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists ​
Author: Ye Chen, Weining Zhang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20918v1 Announce Type: new Abstract: Organizations maintain task-specific adapters for open-weight language models, and each new base-model release forces a migration decision: retain existing specialists, port adapters, refresh from preserved behavior, or retrain. Prior transfer work eva...
58. Graph-Operator World Models for Morphology-Parameter Generalization in Continuous Control ​
Author: Xu Yang, Yiqin Yang, Qianchuan Zhao
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.20936v1 Announce Type: new Abstract: World models for continuous control are commonly trained for a fixed physical system and can degrade when known morphology parameters such as link lengths, masses, damping, and actuation change. Existing approaches often provide these parameters as con...
59. No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators ​
Author: Ye Chen, Weining Zhang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20938v1 Announce Type: new Abstract: Evaluators often produce correct labels via flawed reasoning, a critical failure for agentic systems gating actions, routing reviews, or supplying training feedback. Standard evaluation only verifies final label correctness, ignoring whether judgment c...
60. The Logic of Machine Self-Preservation ​
Author: Cheng Siong Chin
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.MA
arXiv:2608.20940v1 Announce Type: new Abstract: There is already evidence of agentic AI exhibiting self-preservation behaviors: resisting deactivation, misrepresenting their activities, and, in some instances, attempting to copy themselves into other machines. This can be attributed to a phenomenon ...
61. TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming ​
Author: Yibo Hu, Yu Qian, Mao Gu, Yingfan Tao, Yuhao Chen, Yongdong Luo, Zhuoqun Liu, Meiguang Jin, Junfeng Ma
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.20958v1 Announce Type: new Abstract: E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We present TLive-Omni, an omni-modal un...
62. Can Scientific Claims Be Removed from Large Language Models? A Systematic Evaluation of Claim-Level Unlearning ​
Author: Snigdha Paul, Manasi Patwardhan, Arman Cohan
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20960v1 Announce Type: new Abstract: Language models (LMs) are trained on static scientific corpora, whereas scientific knowledge continuously evolves through correction and revision. Scientific claims encoded within these models may later become retracted, disproven, or updated by subseq...
63. TreeWY: Speculative Verification for Gated DeltaNet Hybrids ​
Author: Sneha Murthy Ghantasala
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.DC, cs.LG, cs.PF
arXiv:2608.20961v1 Announce Type: new Abstract: Modern open models are hybrids: most layers are linear-attention (Gated DeltaNet, GDN) layers carrying a small fixed-size recurrent state instead of a growing key-value (KV) cache. This makes ordinary decoding memory-efficient, but hurts speculative de...
64. Generalizing Soft Tissue Deformation and Force Prediction Across Material Stiffness and Geometry ​
Author: Madina Kojanazarova, Sidaty El Hadramy, Philippe C. Cattin
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CG, cs.CV
arXiv:2608.20967v1 Announce Type: new Abstract: Accurate soft tissue simulation is essential for surgical training, pre-operative planning, and haptic feedback systems. While learning-based surrogate models trained on data using the finite element method (FEM) offer a promising path to real-time inf...
65. Deep Learning Models Also Recall Features ​
Author: Pierre Beckmann
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20970v1 Announce Type: new Abstract: Recent work in mechanistic interpretability has studied how large language models recall facts stored in their weights. This paper argues that factual recall points to something broader: a general kind of operation in deep learning models, which I call...
66. Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models ​
Author: Tonglin Yan, Gregoire Sergeant-Perthuis, David Rudrauf
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20975v1 Announce Type: new Abstract: Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied beha...
67. Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents ​
Author: Yanze Jiang, Mingxuan Li, Yuhao Wang, Shengfang Zhai, Jiaheng Zhang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21027v1 Announce Type: new Abstract: LLM agents are emerging as an important paradigm for real-world tasks that require reasoning, tool use, and sequential decision-making. As these agents operate over longer horizons, runtime intervention offers a way to improve reliability without retra...
68. Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance ​
Author: Alexander Thomas, Hubert P. H. Shum, Darren Nellis, Manli Zhu, Phatpicha Yochum, William Bartle, Daniel Wrightson
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21036v1 Announce Type: new Abstract: The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code, a complex regulatory framework where errors in classification, packaging, stowage, or segregation can result in f...
69. Socialized Division and Collaboration: Rethinking Class-Incremental Learning under Optimization Conflicts ​
Author: Xinjie Yao, Zhihe Fan, Yunqi Zhu, Jiaqi Zhou, Dengyu Zhao, Zhoupeng Guo, Yan Fan, Guosong Jiang, Pengfei Zhu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21044v1 Announce Type: new Abstract: Class-incremental learning is commonly instantiated as a single-model paradigm, where a unified model sequentially adapts to an unbounded stream of sessions. While effective under mild distributional shifts, this formulation becomes strained when succe...
70. The Cost of a Physics Prior Is Bounded by the Ablation Gap ​
Author: Boris Kriuk
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21059v1 Announce Type: new Abstract: Shape-constrained and physics-informed learning reports an accuracy cost of enforcing a prior and treats it as a property of the prior. We show it is mostly a property of the free features and the validation split. Let P be the excess risk of restricti...
71. CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models ​
Author: Bokai Zhao, Yiyang Zhang, Hanqing Chao, Yawei Ma, Long Bai, Tai Ma, Minfeng Xu, Ming Song, Tianzi Jiang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.21060v1 Announce Type: new Abstract: Pathology foundation models (PFMs) are increasingly used as general-purpose backbones, yet existing benchmarks cannot systematically diagnose their whole-slide cellular representation capabilities, including the decodability of cell-type information an...
72. Can Legal AI Know When It Is Wrong? And Do Students Know When It Is? ​
Author: Angel Mary John, Vipin Kumar Singh, Jerrin Thomas Panachakel
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21089v1 Announce Type: new Abstract: Integrating Large Language Models (LLMs) into the Indian judiciary promises access to justice but introduces severe risks. We identify the 'inertia of confidence'--an overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide i...
73. When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge ​
Author: Xin Sun, Di Wu, Yuchen Guo, Jiahuan Pei, Isao Echizen, Abdallah El Ali, Saku Sugawara
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21097v1 Announce Type: new Abstract: LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are often interpreted as independent evidence. We test this assumption for a common pair of judgments: trust scoring...
74. ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models ​
Author: Wenzheng Jiang, Xuankun Rong, Yuanzhao Zhai, Dawei Feng, Huaimin Wang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21100v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) extend model capabilities beyond text, they also make safety alignment increasingly challenging. Multimodal safety alignment methods must address cross-modal jailbreaks, safety-awareness failures, and over...
75. Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda ​
Author: Wei Lin, Tao Zhou, Zhaofei Xie, Changgui Hong
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.21107v1 Announce Type: new Abstract: Large Language Models (LLMs) are moving from code completion toward repository-scale agents that retrieve context, edit files, execute tools, and participate in security-sensitive workflows. The evidence for these systems, however, remains divided betw...
76. Root cause analysis via difference graph discovery from linear time-series data ​
Author: Anouk Ruer, Timoth'ee Loranchet, Daria Bystrova, Charles K. Assaad
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21117v1 Announce Type: new Abstract: Root cause analysis aims to identify the mechanisms responsible for anomalies in complex dynamical systems. In this paper, we study root cause analysis in linear time-series through the lens of difference graph discovery. We focus on effect-defying roo...
77. From Attention Masks to Inert Zero-Vector Tokens: OAttention and O-Closure for Token Dynamics ​
Author: Heyang Gong
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21174v1 Announce Type: new Abstract: Attention masks are relation-level controls: they specify which query--source pairs may interact. They do not provide a representation-carried token state that is non-participating at the attention boundary. We assign each token hidden carrier (h_i) ...
78. SENTRY: Deterministic, Intelligent Risk Assessment for IT Change Management ​
Author: Daniel Arulpragasam, Christer Henrysson, Ella Ly, Deepika Anbalagan, Leo Feng
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21203v1 Announce Type: new Abstract: Technology change management in large financial institutions depends on risk assessments that are accurate, consistent, and auditable. In practice, many institutions still rely on self-reported questionnaires. Those questionnaires are subjective, easy ...
79. Personalized Privacy Control in LLMs via Attention Head Intervention ​
Author: Junseok Kim, Nakyeong Yang, Kyomin Jung
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.21209v1 Announce Type: new Abstract: The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studies whether LLMs regulate information disclosure according to context-dependent norms. However, acceptable disclosu...
80. Enhancing LLMs in Predictive Political QA with Semi-Structured Data ​
Author: Yinan Liu, Zihan Zhou, Zichun Jin, Xinyu Wang, Bin Wang, Xiaochun Yang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR
arXiv:2608.21218v1 Announce Type: new Abstract: Predictive political question answering (QA), such as predicting how a political actor will vote, goes beyond factual lookup. External political resources offer rich historical evidence, but rarely contain the answer itself. Existing LLM augmentation m...
81. Ontology-supported AI Model and Dataset Management ​
Author: Jan Novacek, Ali Ahari, Tobias M"uller, Sebastian Reiter, Alexander Viehl, Oliver Bringmann
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21224v1 Announce Type: new Abstract: Recently, there has been a great deal of research into improving AI methods and their application. The main focus is on tracking progress, enabling transparent comparisons, and fostering a more profound understanding of AI. In that process, different o...
82. Fine-Grain GPU Parallelization of the Generalized Partition Crossover for Large-Scale Traveling Salesman Problems ​
Author: Swetha Varadarajan, Darrell Whitley
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2608.21233v1 Announce Type: new Abstract: The Traveling Salesman Problem (TSP) is one of the most extensively studied NP-hard optimization problems. Genetic Algorithm (GA)-based solvers, such as the Edge Assembly Crossover (EAX), achieve state-of-the-art performance on many benchmark instances...
83. CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment ​
Author: Chengxiao Wang, Enyi Jiang, Xiaojing Liao, Sanmi Koyejo
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21278v1 Announce Type: new Abstract: Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safety tuning may affect model responses to both harmful and benign inputs. We propose \textbf{C}ontinuous \textbf{L}at\textbf{E}nt \textbf{...
84. AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization ​
Author: Huizu Lin, Chengkai Huang, Tianqi Gao, Tao Huang, Daijiao Liu, Tongxin Li, Xiaoyan Sun, Lina Yao
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21292v1 Announce Type: new Abstract: Skills play different roles as an agent's policy evolves: they should first provide learnable knowledge, then support capability formation, and finally be invoked only when they improve individual decisions. Existing methods rarely model this lifecycle...
85. From Regulation to Implementation: A Critical Evaluation of LLM-Assisted Regulatory Compliance in Industry ​
Author: Adriana Watson, Marco B"ucheler, Grant Richards
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21317v1 Announce Type: new Abstract: The European Union (EU) has emerged as a leading regulatory body in the development of sustainability and privacy regulations. While new regulation requirements vary, many include a documentation artifact to ensure compliance. Notably, the Ecodesign fo...
86. Unified Branch-and-Bound Search for the Steiner Traveling Salesman Problem on Graphs of Convex Sets ​
Author: Jingtao Tang, Hang Ma
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.21319v1 Announce Type: new Abstract: We formalize the Steiner Traveling Salesman Problem (Steiner-TSP) on Graphs of Convex Sets (GCS), which seeks a minimum-cost closed trajectory through required convex sets while allowing optional transit vertices and revisits. To explore the resulting ...
87. Anatomy-Informed Neural Networks: Encoding Anatomic Priors in Loss and Architecture, with an SE(3) Formulation of Guidewire-Induced Aortoiliac Deformation ​
Author: David P. Stonko
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.RO
arXiv:2608.21332v1 Announce Type: new Abstract: Deep-learning models of anatomy can be numerically plausible yet anatomically impossible, and they generalize poorly when data are scarce. We introduce Anatomy-Informed Neural Networks (AINN), in which soft anatomic priors enter as penalty terms in the...
88. VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences ​
Author: Elaine Lau, Thanuka Udumulla, Lee Izhaki-Tavor, Francisco Guzm'an, Nicholas Magazine, Jonas Mueller
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.21357v1 Announce Type: new Abstract: In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures, ...) to inform research decisions. We introduce VIALS, a visual question-a...
89. When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha ​
Author: Manisha Mehta, Virendra Mehta
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.20345v1 Announce Type: cross Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from ther...
90. Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias ​
Author: Keren Fuentes, Aaron Mueller
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.20347v1 Announce Type: cross Abstract: Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study, we show that repres...
91. Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing ​
Author: Sanjay Basu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20348v1 Announce Type: cross Abstract: Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost-in-the-middle (LitM) effect: information near the center of a long context is retrieved less reliably than information near the edge...
92. Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality ​
Author: Qipeng Xie, Zi Liang, Jiafei Wu, Yufei Chen, Weizheng Wang, Wenao Ma, Zhong Ming, Haiqin Yang, Kaishun Wu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20349v1 Announce Type: cross Abstract: Large Language Models (LLMs) exhibit extreme sensitivity to surface-level prompt variations, in which minor lexical changes can trigger disproportionate performance fluctuations. Moving beyond black-box optimization and coarse-grained templates, we p...
93. How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel ​
Author: Chang Liu, Chaoyang Ning, Dayi Jiang, Enrui Gu, Fang Ran, Hongyan Xue, Huaqing Li, Hui Cai, Jia Liu, Jiang-Ming Yang, Jianshe Li, Jiawei Luo, Jin Zhou, Leshen Zhu, Lihui Chen, Liying Ma, Lyuxin Xue, Mengjian Ji, Ruijia Xu, Wei Ren, Wei Wu, Xiaoling Qu, Xiaoyun Feng, Xin Zhang, Xixie Zhou, Xuanwei Hu, Yan Chen, Yichao Wang, Yongqi Tong, Yu Liu, Yuhong Zhou, Zemin Sun, Zhenwen Xu, Zhiling Liu, Zifan Wang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20350v1 Announce Type: cross Abstract: Traditional industrial agents rely on modular pipelines, including Router, Retriever, Planner, Executor, Responder, Reviewer, and other components. These systems often fracture into a labyrinth of ad-hoc patches, leading to cascading errors and high ...
94. The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP ​
Author: Moustafa Yehia Hassan
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20353v1 Announce Type: cross Abstract: Computational mental health (CMH) classifiers often degrade under distribution shift because human annotators and distant-supervision pipelines reward different linguistic signals. We introduce TSS (Triple-Stream Stress probe), a multi-channel diagno...
95. NeuroStrata: An Electroencephalographic Connectivity-Aware Deep Representation Learning Framework for Dynamic Brain Network Analysis of Mental Stress ​
Author: Sayantan Acharya, Hamzeh Asgharnezhad, Abbas Khosravi, Douglas Creighton, Roohallah Alizadehsani, U Rajendra Acharya
Published: 8/24/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.LG
arXiv:2608.20354v1 Announce Type: cross Abstract: This study introduces NeuroStrata, a connectivity-aware deep representation learning framework for EEG-based mental stress analysis using Time-Varying Partial Directed Coherence (TV-PDC). Unlike conventional EEG classification approaches based on sta...
96. ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models ​
Author: Zhen Wang, Yuqi Ren, Yuehan Cui, Hongxiang Wang, Jianxiang Peng, Zhaoxia Zhang, Bingkun Zhu, Tongxuan Zhang, Dezhi Tong, Deyi Xiong
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20355v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have demonstrated considerable potential for social simulation, yet struggle to accurately model individual value systems. Most existing methods mechanically stitch survey responses into prompts, which suffer from se...
97. Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration ​
Author: Deqiang Huang, Jingbo Zhou, Xinjiang Lu, Tong Xu, Hua Wu, Enhong Chen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.20357v1 Announce Type: cross Abstract: Deep search is brittle on underspecified user queries: missing constraints such as time, location, scope, or definitions can lead to retrieval drift and incomplete answers. We introduce Clarify-Then-Search, a benchmark for evaluating whether LLM-gene...
98. Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure ​
Author: Yuchen Wang, Zhongzhi Luan
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20361v1 Announce Type: cross Abstract: Automated research-idea generation systems built on large language models (LLMs) share a structural weakness: they reduce ideation to free-text recombination, random paper pairing, or embedding-similarity retrieval. The three approaches fail in the s...
99. Hadith computational science in the age of large language models: a critical narrative review ​
Author: Md. Ashraful Haque (Greentech Apps Foundation, United Kingdom), Riasat Islam (Greentech Apps Foundation, United Kingdom, Queen Mary University of London, London, United Kingdom)
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20364v1 Announce Type: cross Abstract: We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, and large language models (LLMs). Recent reviews document growth in the literature, but they do not yet provide a critical account of w...
100. Trilingual Topic Modeling of Sri Lankan Parliamentary Debates ​
Author: Himath Dhanapala, Haren Daishika, Himandhi Kuruppu, Sithija Seneviratne, Ashini Kavindya, Patalee Narasinghe, Sandeepa Weerasekara, Nisansa de Silva, Sandareka Wickramanayake
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20365v1 Announce Type: cross Abstract: Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed content, yet remain inaccessible to standard NLP pipelines due to layout-complex PDFs, multilingual scripts, a...
101. A Hybrid Edge Cloud Digital Twin for Welfare-Constrained Control in Poultry Production ​
Author: Suresh Neethirajan
Published: 8/24/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY
arXiv:2608.20367v1 Announce Type: cross Abstract: Poultry production operates under tightly coupled environmental and biological dynamics, yet commercial climate control remains largely heuristic, limiting welfare assurance and operational efficiency. We introduce an edge-cloud digital twin framewor...
102. ASTAR: Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora ​
Author: Xinfeng Zhang, Mingxuan Liu, Yifei Chen, Juncheng Zhu, Kasidit Anmahapong, Yiming Huang, Yuan Zhang, Hongjia Yang, Yi Liao, Gang Ning, Haibo Qu, Qiyuan Tian
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20369v1 Announce Type: cross Abstract: Structured reporting converts free-text radiology narratives into queryable data keys, facilitating cohort assembly, longitudinal tracking, and training label generation for medical AI. The prevailing paradigm follows a two-stage pipeline: (1) constr...
103. When Do LLMs Replace Fine-Tuned NLU? A Decision Framework for Intent Detection in Production Conversational Systems ​
Author: Carson Rodrigues, Oysturn Vas
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20371v1 Announce Type: cross Abstract: A common claim is that zero-shot large language models (LLMs) can replace fine-tuned NLU classifiers for intent detection. We test this claim head-to-head and find that the honest answer is: it depends on the intent space. On full ATIS and CLINC150 w...
104. Edge-Based Agentic Retrieval-Augmented Generation for Autonomous FHWA Bridge Inspection Compliance ​
Author: Viraj Nishesh Darji, Hemaliben Rakeshkumar Darji
Published: 8/24/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.MA
arXiv:2608.20372v1 Announce Type: cross Abstract: The Federal Highway Administration (FHWA) mandates that over 600,000 bridges in the United States be evaluated against the Recording and Coding Guide for the National Bridge Inventory (NBI). Manual compliance verification is labor-intensive, error-pr...
105. VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models ​
Author: Hyunwoo Kim
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.20374v1 Announce Type: cross Abstract: How precisely can we tell a language model how to feel? Most work on emotional generation answers with a discrete label - happy, angry, sad - which cannot express a target like "mildly downcast but calm." We instead specify the desired affect as a co...
106. EditPPT: Faithful Long-Deck Slide Editing via Structured Tool-Using Multi-Agent with Dual-Modal Validators ​
Author: Jiheon Kim, Kyudan Jung, Jaegul Choo
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2608.20381v1 Announce Type: cross Abstract: Automating slide editing requires simultaneously satisfying modification accuracy, preservation fidelity, and robustness to deck length. Existing LLM-based systems often fail on real-world presentation files because they rely on idealized intermediat...
107. Infrared Hotspot-Guided Early Warning of Lithium-Ion Battery Thermal Runaway Under Mechanical Abuse ​
Author: Syed Sajid Ullah, Salman Khan, Muhammad Zunair Zamir
Published: 8/24/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY
arXiv:2608.20383v1 Announce Type: cross Abstract: Mechanical abuse can trigger thermal runaway (TR) in lithium-ion batteries through localized heat generation before sensor signals become decisive. This paper proposes a two-stage early-warning approach that estimates localized thermal instability fr...
108. Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions ​
Author: Junhui Zhang, Qianhui Xu, Qingxiang Guo, Dawei Yang, Ling Miao, Qiangqiang Wang, Yang Song
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20387v1 Announce Type: cross Abstract: While recent text-to-speech (TTS) models achieve high naturalness, controlling fine-grained expression via natural-language instructions remains challenging. We introduce Poly- InstructTTS, which learns expressive speech from open-ended instructions ...
109. Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations ​
Author: M Waleed Kadous, Amr Elsayed, Abdullah Al Nahas, Ashraf Haress
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.20390v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) are increasingly used to answer religious questions, but for Islamic content they carry two serious risks: factual fabrication (inventing Qur'anic verses or hadith) and subtle value misalignment. We presen...
110. Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants ​
Author: Sami Khairy, Yasaman Hosseinkashi, Vishak Gopal, Ross Cutler
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20392v1 Announce Type: cross Abstract: LLM-powered meeting assistants are deployed at scale, yet systematic evaluation of their grounding fidelity remains limited to static benchmarks that miss failure modes tied to specific discourse structures or reasoning demands. We propose Evaluation...
111. Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI ​
Author: Tanmay Kumar Shrivastava, Darsh Rohit Nandu, Rajesh Kumar Mundotiya
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20393v1 Announce Type: cross Abstract: Agentic large language models (LLMs) deployed in fact-sensitive applications such as customer support must simultaneously preserve factual correctness and generate responses in a controllable stylistic register. Activation steering enables fine-tunin...
112. LingShu: A Large-Scale Symptom-Centric Contextualized Knowledge Graph Bridging Traditional Chinese Medicine and Modern Biomedicine ​
Author: Rui Hua, Zixin Shu, Kai Chang, Dengying Yan, Jianan Xia, Hui Zhu, Shujie Song, Shurui Yang, Tongxin Wang, Yue Yin, Yu Wei, Lijuan Pei, Yunhui Hu, Hao Xu, Mingzhong Xiao, Xiaodong Li, Haibin Yu, Runshun Zhang, Wenjia Wang, Baoyan Liu, Xuezhong Zhou
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20402v1 Announce Type: cross Abstract: Biomedical knowledge graphs (KGs) are pivotal for knowledge organization, yet traditional binary relations often struggle to represent the conditional nature of biomedical knowledge. Symptoms provide a shared phenotypic layer for linking Traditional ...
113. Rigorous Evaluation of Large Language Models for Malaria Drug Discovery: Trade-offs in Performance, Scale, and Resource Utility ​
Author: Marvellous O. Ajala (Magami Open Sciences Initiative), Zainab Ashimiyu-Abdusalam (Magami Open Sciences Initiative), Comfort Adesina (Magami Open Sciences Initiative)
Published: 8/24/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI, cs.LG
arXiv:2608.20418v1 Announce Type: cross Abstract: We introduce Malaria-Instruct, a curated instruction-following dataset derived from the ChEMBL Legacy Malaria corpus for Malaria virtual screening, and conduct a systematic evaluation of five open-source LLMs; Gemma-2 2B/9B, TxGemma-2B/9B, and LlaSMo...
114. Six misconceptions about large language models: A minimal model and diagnostic taxonomy ​
Author: Zhicheng Lin
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.20421v1 Announce Type: cross Abstract: Large language models (LLMs) are now embedded in scientific, educational, and governance workflows, with debates centering on their capabilities, mechanisms, and impacts. Yet these debates remain structured by persistent folk theories--intuitive, inf...
115. From Thermal Preference Prediction to Adaptive Thermal Intervention: A Reinforcement Learning Approach Using Physiological and Environmental Sensing ​
Author: Isibor Kennedy Ihianle, Emmanuel Manu, Ehsan Asnaashari, Mojgan Jadidi, Pedro Machado, Amrit Sagoo, Ahmad Lotfi
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.20423v1 Announce Type: cross Abstract: Personalised thermal comfort is essential for occupant wellbeing and for the development of more responsive building-control strategies, yet conventional Heating, Ventilation, and Air Conditioning (HVAC) systems rely on static setpoints and populatio...
116. BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers ​
Author: Hina Dixit
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.20427v1 Announce Type: cross Abstract: Dense causal attention remains expensive at long context even when implemented with highly optimized exact kernels. We study BF1, a deterministic block-aligned dyadic sparse-attention route that combines a small exact local neighborhood, a global fir...
117. Approximate Homomorphisms and Convergent Representations in Transducers ​
Author: Santiago Cifuentes
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.20428v1 Announce Type: cross Abstract: We study the stability of minimal representations of controlled stochastic processes (in particular, transducers) under perturbations. This question is motivated by recent experiments finding predictive-state structure in the latent representations o...
118. ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib ​
Author: Shane Caldwell
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.CL
arXiv:2608.20432v1 Announce Type: cross Abstract: Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM-as-judge system that scores formal proof quality along five dimensions beyond correctness: library leverage, a...
119. An LLM agent for end-to-end computational materials discovery ​
Author: Chen Yuntong, Huang Ju, Liu Yu, Zhao Dan, Sun Mingqi, Ju Chentian, Liu Yanbing, Huang Lijiang, Zhao Guobin
Published: 8/24/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI
arXiv:2608.20434v1 Announce Type: cross Abstract: The coordination of multi-scale tasks is an effective strategy for computational materials discovery, yet the repeated application of diverse algorithms and tools renders it challenging. We report MAESTRO, a large language model (LLM) agent system ca...
120. Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources ​
Author: Rana Muhammad Usman, Dominic Williamson
Published: 8/24/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI, cs.MA, cs.SI
arXiv:2608.20438v1 Announce Type: cross Abstract: Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered matched-exposure experiment sp...
121. Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach ​
Author: Amrita Shaw, Chandrasekar S. N., Sai Muthukumar V., Jhinuk Gupta, Deepak L. N. Kallepalli
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.20440v1 Announce Type: cross Abstract: Authentication of edible oils in processed foods is important for food quality, fraud prevention, and regulatory compliance. This study establishes an integrated Raman spectroscopy and machine-learning framework that links intrinsic spectral organiza...
122. AEGIS: Preventing Cross-Domain Resource Abuse in MCP ​
Author: Shriti Priya, Teryl Taylor, Frederico Araujo
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.20481v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) is an open source JSON-RPC protocol that standardizes how large language models (LLMs) interact with external systems through programmatic functions known as tools. Attackers or malicious agents can exploit certain mo...
123. Towards Traffic Modelling of Multi-Agent Systems: The Role of Coordination Topology ​
Author: Davide Lamagna, Albert Cabellos, Alberto Rodriguez-Natal, G'abor R'etv'ari, Berta Serracanta
Published: 8/24/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.MA
arXiv:2608.20494v1 Announce Type: cross Abstract: Multi-agent LLM systems are an emerging networked workload whose rapid deployment raises questions about the traffic patterns they generate. Compared to conventional applications, these systems generate requests internally: a single user task can ind...
124. Making Deployments Safe at Meta: Health Checks for Continuous Change-Safety ​
Author: Prakash KL, Anton Korenkov, Uttam Thakore, Christopher Hegre
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.20513v1 Announce Type: cross Abstract: Continuous deployment to large scale production systems creates a tension between release velocity and reliability. Every change is a potential reliability incident, yet every delay is a missed opportunity. This paper describes the deployment time he...
125. An integrated diffusion-weighted imaging processing and interpretation platform for MR-guided radiotherapy ​
Author: Yunxiang Li, Yan Dai, Yen-Peng Liao, Jie Deng, Jill B De Vis, You Zhang
Published: 8/24/2026, 4:00:00 AM
Categories: physics.med-ph, cs.AI
arXiv:2608.20519v1 Announce Type: cross Abstract: Background: Magnetic resonance imaging-guided linear accelerators (MR-Linacs) allow diffusion-weighted imaging (DWI) to be acquired at every treatment fraction, but converting these low-signal-to-noise-ratio acquisitions into clinical decisions requi...
126. Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes ​
Author: Praveen Pathak, Siddharth Tiwary, Charudatt Kadolkar, Vijay Singh, David Rakestraw, Shirish Pathare, Anwesh Mazumdar
Published: 8/24/2026, 4:00:00 AM
Categories: physics.ed-ph, cs.AI, cs.CY
arXiv:2608.20521v1 Announce Type: cross Abstract: Multimodal AI can read handwritten physics solutions, but high-stakes grading requires agreement with official scores and outcomes. This study evaluated GPT-5.5-based grading on 10364 scanned pages from 520 handwritten submissions by 416 unique candi...
127. ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations ​
Author: Naveen Venkatanarayanan, Yuchen Qiu, Tianyi Peng, George Gui, Olivier Toubia
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2608.20539v1 Announce Type: cross Abstract: Digital twin simulations show promise, but current empirical evidence suggests that the approach should be tested before being deployed in any particular context. To lower the friction for researchers and practitioners to test and deploy digital twin...
128. Consistency Models for Fast MRI Reconstruction Using Regularization by Denoising ​
Author: Merve G"ulle, Junno Yun, Ya\c{s}ar Utku Al\c{c}alar, Mehmet Ak\c{c}akaya
Published: 8/24/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG, physics.med-ph
arXiv:2608.20561v1 Announce Type: cross Abstract: Diffusion models (DMs) have emerged as powerful generative priors for MRI reconstruction with promising results. Yet DM-based methods require extensive iterative refinement, limiting their practical deployment. Consistency models (CMs) provide a comp...
129. Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents ​
Author: Wei Shao, Chongzhou Fang, Zuxiong Tan, Zequan Liang, Setareh Rafatirad, Avesta Sasan, Houman Homayoun
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.20563v1 Announce Type: cross Abstract: Long-horizon security LLM agents must carry information and decisions across many dependent interactions, where later actions often depend on services, state, or access discovered much earlier. This makes final task success difficult to interpret: an...
130. Aggregate, Don't Adapt: Subject-Level Posterior Aggregation and Transductive Calibration for Cross-Site Parkinsonian Gait Severity ​
Author: Junlong Shen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.20587v1 Announce Type: cross Abstract: We describe the winning entry to the MoCha 2026 Benchmark and Challenge on Parkinsonian Gait, which predicts MDS-UPDRS gait severity from canonicalized SMPL motion recorded at clinical sites unseen during training. The system reaches 0.6945 macro-F1 ...
131. Testing and Evaluation of Agentic AI Systems In Military Command and Control ​
Author: Ulysse Richard, Heather Frase, Sarah Cao, Di Cooke, Sebastian Kwon, Adrianna Tan
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CY
arXiv:2608.20597v1 Announce Type: cross Abstract: Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assurance case, which requires three ...
132. JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification ​
Author: Tianxin Zhou, Ruixi Lin
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.20607v1 Announce Type: cross Abstract: Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared false-negative blind spots r...
133. When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation ​
Author: Lauren Pothuru
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20627v1 Announce Type: cross Abstract: Agentic retrieval-augmented generation (RAG) interleaves retrieval, reasoning, and answer generation across multiple hops. A retrieval error at hop 1 can surface only as a wrong answer at hop 3, while later retrieval can also repair the trajectory. T...
134. AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale ​
Author: Minbyul Jeong, Chanwoong Yoon
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20634v1 Announce Type: cross Abstract: Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthesized around predefined tasks and benchmarks. This task-centric paradigm makes it difficult to scale environment...
135. ARQ: Agentic CodeQL Query Refinement for C/C++ Vulnerability Detection ​
Author: Chunyi Wang, Yunfei Ke, Junfeng Yang, Yun-Yun Tsai, Penghui Li
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.20637v1 Announce Type: cross Abstract: Static analyzers have been widely adopted for vulnerability detection in C/C++ programs. Query-based static analyzers (e.g., CodeQL) encode vulnerable code patterns in detection queries and match them against source code. However, existing queries st...
136. Provable Edge-of-Stability for Adam on a One-Dimensional Quadratic ​
Author: Yiman Fong, Heng Yang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC
arXiv:2608.20638v1 Announce Type: cross Abstract: The edge-of-stability (EoS) phenomenon of Adam has been widely observed, while its underlying dynamical mechanism is not yet fully understood. We study uncorrected Adam on a one-dimensional quadratic, a clean setting where constant curvature isolates...
137. One Hierarchy, Two Systems: Semantic Product IDs for Discovery-Surface Ranking and Search-Page Query Reformulation ​
Author: Steven Xu, Sanjyot Thete, Saathvik Dirisala, Raghav Saboo, Nimesh Sinha, Leo Shao, Elyse Winer, Sudeep Das, Martin Wang, Kyle MacDonald
Published: 8/24/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.20640v1 Announce Type: cross Abstract: Multi-merchant e-commerce catalogs contain equivalent and related products under different merchant-scoped identifiers, fragmenting behavioral evidence across merchants. Expert-defined taxonomies, meanwhile, are often too coarse for fine-grained disc...
138. RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction ​
Author: Guangyu Wang, Zhidan Liu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.20656v1 Announce Type: cross Abstract: Traffic sensors commonly record flow, speed, and occupancy, but standard traffic flow forecasting benchmarks and models rarely exploit all three raw measurements reliably. Although speed and occupancy provide sensor-native traffic-state information b...
139. Amplifying the imaging power of digital sky surveys with space telescopes data and generative AI ​
Author: Sai Teja Erukude, Lior Shamir
Published: 8/24/2026, 4:00:00 AM
Categories: astro-ph.IM, astro-ph.GA, cs.AI, cs.LG
arXiv:2608.20666v1 Announce Type: cross Abstract: While Digital sky surveys provide excellent throughput of image data and can cover a large footprint, their imaging power is normally inferior to that of space-based telescopes. Space-based telescopes, on the other hand, provide excellent imaging pow...
140. C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination ​
Author: Tsao-Lun Chen, Chi-Cheng Fu, Han-Yi E. Chou, Shun-Feng Su
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.20667v1 Announce Type: cross Abstract: Pseudo-label-based semi-supervised learning has achieved strong performance due to its simplicity and scalability. However, it is typically developed under a closed-world assumption that unlabeled data are drawn from the same distribution as labeled ...
141. Lightweight Adaptive ReduNet via Hyperspherical Manifold Learning ​
Author: Zhenglin Huang, Qifa Yan, Bin Dai, Xiaohu Tang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.20668v1 Announce Type: cross Abstract: In recent years, a white-box neural network called ReduNet has been proposed, which employs the maximal coding rate reduction (MCR$^2$) principle to transform raw data into low-dimensional discriminative features via a forward layer-wise construction...
142. Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes ​
Author: Neeraj Yadav
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL, cs.LG
arXiv:2608.20685v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) has no model of time: when a fact changes across a coding session - a function is renamed, an endpoint moves, a dependency is bumped - RAG retrieves both the old and new value with near-identical similarity and ca...
143. Identity-Aware Human-Object Interaction Motion Captioning ​
Author: Yiming Wang, Yonghao Dang, Huilai Li, Jiawei Tu, Jianqin Yin
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.20690v1 Announce Type: cross Abstract: Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using generic terms such as "a person" or "someone", without grounding the caption in subject identity. To address this l...
144. Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation ​
Author: Rujin Liang, Zhongpu Chen, Yuhao Lei, Xin Miao
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.20756v1 Announce Type: cross Abstract: While multimodal retrieval-augmented generation (RAG) systems increasingly rely on images as external knowledge sources, the introduction of poisoned visual evidence can severely compromise multimodal large language model (MLLM) generation. Unlike pr...
145. PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering ​
Author: Srikar Kashyap Pulipaka
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.20757v1 Announce Type: cross Abstract: We describe the PSK submission to the WMT 2026 Multilingual Instruction Shared Task. Our system uses the 3.35B-parameter Tiny Aya Global model with three QLoRA adapters, one for each task. The adapters are trained on multilingual document-summary pai...
146. Fuzzy-MoE: Interpretable Regime-Conditioned Expert Routing for Non-Stationary Multivariate Time Series Forecasting ​
Author: Lan Guo, Jie Xiao, Zhao Su, Jun Shen, Haoran Li, Weixia Ma, Qingguo Zhou, Binbin Yong
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.20761v1 Announce Type: cross Abstract: In non-stationary multivariate time series, different variables and samples often exhibit heterogeneous latent dynamic states, while existing deep forecasting models usually compress them into a unified end-to-end mapping, leading to suboptimal model...
147. CARD: Diagnosing Belief to Action Routing Failures in Vision Language Models ​
Author: Souptik Kumar Majumdar, Fabian K"ogel, Andreas Bulling
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.20763v1 Announce Type: cross Abstract: Linear probes and activation steering have uncovered that vision-language models (VLMs) internally represent mental states such as agents' beliefs, knowledge, and intentions. However, it is unclear whether and how these representations are used by do...
148. Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding ​
Author: Haoyue Liu, Zhichao Wang, Ye Chen, Haonan Deng, Xiaoying Tang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.20769v1 Announce Type: cross Abstract: Streaming emotion understanding uses historical state while continuously interpreting current audio, often feeding the model's previous prediction back as context. We show that this history conditioning can distort current perception. On a balanced C...
149. CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models ​
Author: Hui Lu, Zhijie Peng, Yuqi Lin, Zaijia Yang, Jiaming He, Shuhan Ye, Yi Yu, Hanwei Zhu, Bingquan Shen, Alex Kot, Xudong Jiang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.20791v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies are vulnerable to localized physical perturbations, yet existing certified patch defenses target discrete labels and cannot directly certify continuous, temporally correlated actions. We introduce CertVLA, a cert...
150. Profiling What Matters: Context-Aware Item Profiles from Large-Scale Metadata for LLM Recommenders ​
Author: Dojun Hwang, Seunghan Lee, Cheonyoung Park, Sara Yu, SeongKu Kang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2608.20801v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have significantly advanced reranking in recommendation, effectively leveraging item-side information remains challenging. Real-world items are described by vast, heterogeneous, and unstructured metadata, where deci...
151. Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation ​
Author: Yanglei Gan, Peng He, Run Lin, Peiyuan Jiang, Yifan Wang, Qiao Liu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20804v1 Announce Type: cross Abstract: Temporal Knowledge Graph (TKG) extrapolation seeks to infer future facts from time-varying relational histories. Recent diffusion-based approaches improve uncertainty modeling through generative denoising, but their aggregated conditioning on subject...
152. TRACE: Training-time Report-guided and Clinically Ordered Concept Editing ​
Author: Wentao Yue, Tianyou Lai, Jiayu Luo, Qingyu Mao, Ziying Wang, Zhenyuan Ning, Qilei Li
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.20809v1 Announce Type: cross Abstract: Breast ultrasound diagnosis relies on clinically meaningful semantic concepts, yet most deep learning methods adopt end-to-end image-to-label paradigms that lack interpretability and robustness. While concept-based approaches offer a promising altern...
153. When Generated Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception ​
Author: Guangyuan Dong, Chuang Liu, Yangchen Zeng, Haoyu Wang, Xiaoyang Yu, Pinlong Zhao, Yuchao Hou, Ziwei Li, Zheng Lin
Published: 8/24/2026, 4:00:00 AM
Categories: cs.MM, cs.AI, cs.CV, cs.GR
arXiv:2608.20810v1 Announce Type: cross Abstract: Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retrievable by the queries it was meant to serve. When the scene contains e...
154. Scaling Muon for Diffusion Transformers ​
Author: Chenghao Li, Xiao Han, Xinxin Huang, Wei Liu, Boyang Li, Bing Xiao, Heran Zhang, Juanma Perez Rua, Ke Xu, Kangning Liu, Linjun Kuang, Na Li, Tan Wang, Tian Xie, Wei Peng, Yang Pei, Yifan Xu, Yuanhao Zhai, Yuwei Lin, Zhe Wang, Zihao He, Daniel Li, Junbiao Tang, Ziyang Jiang, Dake Chen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2608.20818v1 Announce Type: cross Abstract: The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-end efficiency on large Diffusion Transformers (DiTs) remain unclear. We first establish Muon's scaling...
155. STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction ​
Author: Tong Sun, Mingyang Ma, Jiayang Yu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20831v1 Announce Type: cross Abstract: Aspect-based sentiment analysis (ABSA) quadruple extraction requires jointly predicting target, aspect, opinion, and sentiment over reviews that often contain multiple fine-grained sentiment tuples. While large chain-of-thought (CoT) models perform w...
156. Advantage-level Aggregation Reinforcement Learning for X-point Target Magnetic Configuration Control in an EXL-50U Experiment-Calibrated Simulation Environment ​
Author: Siqi Ding, Xuanhe Wang, Pei Guo, Guoyang Shi, Changquan Yu, Yiting Wang, Xianming Song, Xiang Gu, Zhengyuan Chen, Lei Xing, Yapeng Zhang, Jianguo Chen, Tianyuan Liu
Published: 8/24/2026, 4:00:00 AM
Categories: physics.plasm-ph, cs.AI
arXiv:2608.20834v1 Announce Type: cross Abstract: Managing divertor heat loads is a central challenge for compact, high-power tokamaks. To increase local flux expansion and decouple the dissipation volume from the core, EHL-2 adopts the X-point target (XPT) divertor. This requires the secondary X-po...
157. BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP ​
Author: Haoran Sun, Klaus Marius Hansen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.20851v1 Announce Type: cross Abstract: Agentic engineering systems have shown strong performance on general-purpose benchmarks, yet their effectiveness in enterprise resource planning (ERP) domain-specific languages (DSLs) remains underexplored. We introduce BC-Bench, a benchmark designed...
158. KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs ​
Author: Xubin Chen, Yipeng Zhou, Wen Sun, Chengkai Huang, Xiaoming Fu, Quan Z. Sheng
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20887v1 Announce Type: cross Abstract: Automatic Medical Coding (AMC), which assigns standardized International Classification of Diseases (ICD) codes to clinical notes, is essential for medical reimbursement, quality reporting, and clinical research. Existing pre-trained language model (...
159. Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization ​
Author: Zhu Xu, Jiaqi Tang, Pokai Chen, Yuxin Peng, Yang Liu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.20913v1 Announce Type: cross Abstract: Explainable deepfake detection extends binary classification by requiring models to not only predict authenticity but also provide interpretable justifications. This expanded scope is critical in practice, where users like forensic analysts need insi...
160. Source-Free MT Evaluation Is Not MT Evaluation ​
Author: Baban Gain, Ramakrishna Appicharla, Asif Ekbal
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20925v1 Announce Type: cross Abstract: Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments. As a result, source-free, reference-based evaluation has become the practi...
161. MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation ​
Author: Ziwu Liu, Guozhong Li, Chen Qiu, Weiyang Kong, Panos Kalnis
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20927v1 Announce Type: cross Abstract: Cross-model latent guidance lets a frozen large mentor encode an input once and a frozen small student generate from the resulting signal. Existing methods keep this signal fixed, assuming it stays useful as the output grows; we show this fails in lo...
162. Neural-Primitive: An Efficient End-to-end Local Planner with Primitive-based Imitation Learning for Autonomous Flight ​
Author: Zhitao Liu, Guangtong Xu, Zihan Wang, Jialiang Hou, Chao Xu, Fei Gao
Published: 8/24/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.20948v1 Announce Type: cross Abstract: Autonomous flight in unknown cluttered environments is hindered by the computation-quality-memory trilemma of onboard trajectory generation. In this paper, we propose an efficient end-to-end local planner via imitation learning. A lightweight offline...
163. Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs ​
Author: Bakbergen Ryskulov, Iker Garc'ia-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Rom'an Or'us
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.PF
arXiv:2608.20953v1 Announce Type: cross Abstract: Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context be...
164. Vibe Coding and Web Application Security: A Twin-Prompt Study ​
Author: Darko Andro\v{c}ec
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.20963v1 Announce Type: cross Abstract: Large language models increasingly generate complete web applications from natural-language prompts, raising the question of whether explicitly requesting security best practice improves the result. We study six functionally distinct web applications...
165. Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric ​
Author: Sami Shames El Deen, Mariette Awad
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.20964v1 Announce Type: cross Abstract: In this research, we introduce SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers for extractive summarization tasks. To ensure that the summaries generated by SAraBERT achieve a high coverage of the document's ...
166. Structured but Fragile: On the Limits of LLMs in Cybersecurity Decision-Making ​
Author: Pasquale Malacaria, Yunxiao Zhang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.20966v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in cybersecurity workflows, yet it remains unclear whether they can perform structured security reasoning or merely rely on superficial cues and prior knowledge. We study this question in the context...
167. WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving ​
Author: Xinlin Wang, Yujiao Xiang, Yuheng Zhou, Jingqi Wang, Minqing Huang, Jiajie Huang, Dongxu Wei, Tingguang Zhou, Xiyang Wang, Gong Chen, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.20974v1 Announce Type: cross Abstract: Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA is built around random-mask completion and deterministic regression, m...
168. Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models ​
Author: Deepanshu Pandey, Arnav Chavan, Nahush Lele, Sankalp Dayal, Deepak Gupta
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.20988v1 Announce Type: cross Abstract: Quantization of Large Language Models (LLMs) is often hindered by the sensitivity of the self-attention mechanism to discretization errors. We identify the softmax operator as a bottleneck for quantization stability due to its sensitivity to outliers...
169. Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models ​
Author: Zhen Yang, Sizai Hou, Kaiwen Zheng, Yaofang Liu, Liang He, Yixuan Chen, Kangning Cui
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.21019v1 Announce Type: cross Abstract: Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is rarely treated as a primary objective. We frame calibration-data selection for quantization as a targ...
170. Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge ​
Author: Rishiraj Sengupta, Sotiris Chatzimiltis, Mohammad Shojafar, Xiatian Zhu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.NI
arXiv:2608.21021v1 Announce Type: cross Abstract: Real-world fault analysis in 5G and emerging 6G networks demands domain expertise to analyze free-text diagnostics, including root-cause explanations and recommended actions. LLMs have emerged as a promising approach to automating this, yet whether l...
171. CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment ​
Author: Yutian Jiang, Jiabo Liu, Xixuan Hao, Yuxuan Liang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.21041v1 Announce Type: cross Abstract: Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretab...
172. $Z^2$-ACT: End-to-End Verifiable Agentic Intent Control for Open 6G RAN ​
Author: Sunder Ali Khowaja, Kapal Dev, George C. Alexandropoulos
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.NI
arXiv:2608.21049v1 Announce Type: cross Abstract: With the progression in open and disaggregated 6G radio access networks, it is expected that the system will be able to host multi-vendors. In order to host multi-vendors, it is essential that AI-assisted control loops remain safe, verifiable, and au...
173. CoAnchor: Robust Collaborative Perception under Spatio-Temporal Misalignment via Object-Level Anchors ​
Author: Chi Li, Rui Lin, Aobo Ji, Dongzhu Xu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.21055v1 Announce Type: cross Abstract: Collaborative perception extends the sensing range of a single vehicle by fusing observations from nearby agents, which improves the robustness of autonomous driving. In realistic deployments, however, the received collaborator messages are often aff...
174. AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images ​
Author: Amani Sedrat, Takieddine Chehhat, Youcef Sklab, Hanane Ariouat, Abderrazak Sebaa, Eric Chenin, Jean-Daniel Zucker, Edi Profiti
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.21067v1 Announce Type: cross Abstract: Automated plant traits recognition from herbarium images is essential for plant sciences, yet remains challenging because background elements (e.g., textual labels, mounting artifacts, and color charts) can introduce shortcut learning, leading models...
175. TracingFlow: A Simulation-Free Trajectory Inference Framework Based on Second-Order Dynamics ​
Author: Yuhao Sun, Zekun Wu, Zixun Huang, Peijie Zhou
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.GN
arXiv:2608.21070v1 Announce Type: cross Abstract: Inferring continuous system evolution from sparse temporal snapshots is a key challenge in generative modeling and single-cell omics. While Optimal Transport (OT) is popular, existing frameworks are largely restricted to first-order dynamics, assumin...
176. PromptResponse: Optimizing Prompts for LLM Coding Tasks ​
Author: Erik Thureck, Robert K"uhnen, Tim Jacobowitz
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC, cs.SE
arXiv:2608.21074v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in research workflows and software development pipelines, yet their output remains sensitive to input prompt variations. This paper presents $\unicode{x00AB}$PromptResponse$\unicode{x00BB}$, a contro...
177. Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems ​
Author: Balkrishna Giri, Md Toufique Hasan, Jussi Rasku, Muhammad Waseem, Pekka Abrahamsson
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL, cs.CR, cs.IR
arXiv:2608.21095v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Reliability Gap: high semantic relevance does not guarantee factual truth...
178. A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration ​
Author: Jiekang Feng, Zhihe Fan, Yunqi Zhu, Xinjie Yao, Yueying Zhang, Yike Gao, Ranxin Li, Guanzuo Chen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.21099v1 Announce Type: cross Abstract: Multi-modal object detection is essential for robust scene understanding in challenging conditions, including low-light and adverse environments. Recent vision foundation models (e.g., DINOv3) have exhibited strong representation capabilities, yet ad...
179. ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents ​
Author: Kai Wang, Zeming Wei, BiaoJie Zeng, Chang Jin, An Wang, Xiaokun Luan, Zhixiao Lin, Jingjing Qu, Xia Hu, Xingcheng Xu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.21101v1 Announce Type: cross Abstract: As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a single agent hijacked by a malicious third-party skill can cause data exfiltration, privilege escalation, or casca...
180. Atom Learning Model (ALM): how a real classroom got tokenised ​
Author: Philipp Bogdan
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.21106v1 Announce Type: cross Abstract: The Atom Learning Model (ALM) tokenises a school curriculum. Two secondary mathematics textbooks were read by machine into 1,934 atoms, each one thing a learner can do in a single step, ordered by 4,616 machine-written prerequisite links. Both sides ...
181. CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents ​
Author: Jiancheng Wang, Mingli Zhu, Tong Zhang, Jiaqi Ruan, Wei Wang, Siyuan Liang, Dacheng Tao
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.21114v1 Announce Type: cross Abstract: Visual world-model agents such as DreamerV3 act through a recurrent latent state rather than a single observation, which weakens frame-wise observation attacks and makes their perturbations vary sharply over time under a strict per-frame perturbation...
182. A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans ​
Author: Simon Vincent Abel, Heiko Hillenhagen, Michael G"otz, Timo Ropinski, Ayhan Can Erdur, Daniel Santak Wolf
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.21140v1 Announce Type: cross Abstract: Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report generation and structured image understanding. While modern vision-language models (VLMs) show promising pe...
183. Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence ​
Author: Yuyuan Feng, Zhishang Xiang, Chaobin Yang, Qichao Ma, Zerui Chen, Yujing Zhang, Ke Huang, Chuanjie Wu, Zhaoxu Liu, Yili Wang, Xin He, Jiapu Wang, Zijin Hong, Hao Chen, Yuanchen Bei, Kun Wang, Shengyuan Chen, Ningyu Zhang, Enyan Dai, Linhao Luo, Qingyi Pan, Qi Wang, Wenqi Fan, Guangjing Wang, Na Zou, Yangqiu Song, Xin Wang, Zechao Li, Xia Hu, Qing Li, Xiao Huang, Zhihong Zhang, Jinsong Su, Qinggang Zhang, Yi Chang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.ET
arXiv:2608.21156v1 Announce Type: cross Abstract: LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access,...
184. HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization ​
Author: Jinghao Wang, Qiqi Gu, Chenpeng Wu, Jianguo Yao, Haibing Guan, Xijun Li
Published: 8/24/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2608.21157v1 Announce Type: cross Abstract: High-performance GPU kernels underpin modern deep learning and scientific computing. As workloads become increasingly diverse and GPU hardware evolves rapidly, developing efficient methods for automated GPU kernel generation and optimization has beco...
185. AID-Guard: Stateful Authorization for Delegated Agent Effects ​
Author: Yingzhe Tong, Leyu Dai, Songhui Guo
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.21159v1 Announce Type: cross Abstract: Tool-using AI agents turn delegated tasks into provider effects, yet authorization often ends at admission while provider state, delivery, retry, and recovery evolve. A request may change before commit, or response loss may cause a replacement to cre...
186. Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds ​
Author: Lars Benedikt Kaesberg, Tianyu Yang, Florian Valentin Wunderlich, Terry Ruas, Jan Philip Wahle, Daniel Kurzawe, Bela Gipp
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.21170v1 Announce Type: cross Abstract: Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less clear is how the visual presentat...
187. SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control ​
Author: Ruihua Han, Rui Gao, Zhe Liu, Xinyi Wang, Chang Chen, Shuai Wang, Qi Hao, Jia Pan, Hengshuang Zhao
Published: 8/24/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.21175v1 Announce Type: cross Abstract: Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, simplified geometry, offline computation, or handcrafted parameters to ...
188. DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization ​
Author: Naiyuan Li, Li Dong, Diqun Yan
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.21176v1 Announce Type: cross Abstract: Automatic speech quality assessment aims to predict Mean Opinion Scores (MOS) consistent with human subjective perception and is essential for evaluating speech generation, enhancement, and communication systems. For speech signals, especially synthe...
189. Anchored Regularized Direct Least Squares (ARDLS): Integrating Established Prioritization Operators for Priority Elicitation in the Analytic Hierarchy Process ​
Author: Kevin Kam Fung Yuen
Published: 8/24/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.NA, math.NA
arXiv:2608.21187v1 Announce Type: cross Abstract: Pairwise reciprocal matrices are fundamental to the Analytic Hierarchy Process (AHP), a decision-making model. While the Direct Least Squares (DLS) method provides an intuitive mechanism for deriving priority vectors without complex transformations, ...
190. Towards Investigating Residual Hearing Loss: Quantification of Fibrosis in a Novel Cochlear OCT Dataset ​
Author: Julia Dietlmeier, Benjamin Greenberg, Wenxuan He, Teresa Wilson, Rubing Xing, Jordan Hill, Adrienne Fettig, Madeline Otto, Teyhana Rounsavill, Lina A. J. Reiss, Jingang Yi, Noel E. O'Connor, George W. S. Burwood
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.21189v1 Announce Type: cross Abstract: Objective: Cochlear implants (CIs) are bionic prostheses that restores hearing via electrical stimulation of the auditory nerve. Hybrid CIs, which use electroacoustic stimulation (EAS), combine residual low-frequency acoustic hearing with CI electric...
191. No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation ​
Author: Dimitri Staufer, David Hartmann, Ibrahim Baroud
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.21206v1 Announce Type: cross Abstract: Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's evidential status is uncontrolled, measurements may conflate memorisation, retrieval, name priors and wrong-per...
192. Curriculum-Aware Interpolate-then-Refine: Learned Physiological Time-Series Imputation under Realistic Missingness ​
Author: Yu-Chao Huang, Haochen Zhang, Nicholas Konz, Tianlong Chen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.21207v1 Announce Type: cross Abstract: Imputing physiological time series (arterial blood pressure, blood glucose, etc.) is essential for addressing the missingness that pervades clinical data. Yet modern imputation methods perform poorly in this domain: a recent benchmark found that simp...
193. Specification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software Migration ​
Author: Oleg Grynets, Oleksii Ilchuk, Dariia Zatulna, Vasyl Lyashkevych
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LO
arXiv:2608.21208v1 Announce Type: cross Abstract: This paper investigates cross-agent specification portability using Oracle-to-PostgreSQL migration as a controlled software transformation task. The study combines two experimental stages. First, a specification-first migration pipeline was evaluated...
194. Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking ​
Author: Arulnidhi Karunanidhi
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.21230v1 Announce Type: cross Abstract: Persistent memory makes false information durable: once a false statement is stored, it can be retrieved into future sessions that match it. We measure the cost of this failure mode using plainly worded false assertions generated in a single pass, wi...
195. Adapting Knowledge Graphs for Behavior Denoising in Sequential Recommendation ​
Author: Zichun Jin, Zihan Zhou, Yinan Liu, Bin Wang, Xiaochun Yang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.21243v1 Announce Type: cross Abstract: Sequential recommendation predicts the next item from a user's interaction history, but not every interaction is equally informative. Real logs combine persistent preferences with temporary needs, exploration, and incidental behavior, so some interac...
196. EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering ​
Author: Xuanyu Meng, Jiashuo Sun, Jash Rajesh Parekh, Jiawei Han
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB, cs.IR
arXiv:2608.21252v1 Announce Type: cross Abstract: Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their relationships. Existing retrieval-augmented generation (RAG) methods typically index documents as raw chunks and...
197. Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning ​
Author: Haonan Jia, Shichao Dong, Zenghui Sun, Jiawen Zheng, Ziqi Miao, Gege Shi, Qiuyu Zhao, Jinsong Lan, Xiaoyong Zhu, Bo Zheng
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.21305v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel reasoning strategies. This limitation leads to a performance gap between ...
198. TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems ​
Author: Vladimir Bataev, Lilit Grigoryan, Andrei Andrusenko, Nikolay Karpov, Vitaly Lavrukhin, Boris Ginsburg
Published: 8/24/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL, cs.LG, cs.SD
arXiv:2608.21343v1 Announce Type: cross Abstract: Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases must be recognized accurately under strict latency constraints. Although many context-biasing methods improve recognition accuracy, ...
199. AI with Authority, from Application to Silicon ​
Author: Jason Hickey
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.AR, cs.LO
arXiv:2608.21356v1 Announce Type: cross Abstract: For sixty years, machine verification has been a major cost overhead, affordable only for exceptional artifacts. Here we report that generative AI inverts this relationship: at AI speed, machine verification is not only economical but essential to pr...
200. Primal Acceleration of Newton's Method ​
Author: Nikita Doikov
Published: 8/24/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.LG
arXiv:2608.21359v1 Announce Type: cross Abstract: We develop a new direct accelerated Newton method for minimizing convex functions with Lipschitz continuous Hessian. The algorithm uses only primal variables and performs just one linear solve per iteration. With a simple predetermined choice of para...
201. Online design of dynamic networks ​
Author: Duo Wang, Andrea Araldo, Mounim El Yacoubi
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.SI, physics.soc-ph
arXiv:2410.08875v3 Announce Type: replace Abstract: Designing a network (e.g., a telecommunication or transport network) is mainly done offline, in a planning phase, prior to the operation of the network. On the other hand, a massive effort has been devoted to characterizing dynamic networks, i.e., ...
202. ACQ: A Deployed Two-Stage Framework for Automated Creative Quota Allocation in Large-Scale Online Advertising ​
Author: Ruizhi Wang, Yu Rong, Kai Liu, Bingjie Li, Qingpeng Cai, Fei Pan, Peng Jiang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2412.06167v2 Announce Type: replace Abstract: In digital advertising, demand-side platforms (DSPs) allow advertisers to create multiple ad creatives from a single photo for real-time bidding. While increasing the number of creatives can improve bidding opportunities, it cannot scale indefinite...
203. Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning ​
Author: Can Jin, Hongwu Peng, Qixin Zhang, Yujin Tang, Dimitris N. Metaxas, Tong Che
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2504.09772v3 Announce Type: replace Abstract: Test-Time Scaling has emerged as a powerful method to extend the reasoning capabilities of Large Language Models. However, single-agent TTS faces significant scalability bottlenecks, as excessively long reasoning traces lead to increased inference ...
204. Recognizing Artificial Minds: A Philosophical Defense of AI Cognition ​
Author: Herman Cappelen, Josh Dever
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2504.13988v2 Announce Type: replace Abstract: This work defends the 'Whole Hog Thesis': sophisticated Large Language Models (LLMs) like ChatGPT are full-blown linguistic and cognitive agents, possessing understanding, beliefs, desires, knowledge, and intentions. We argue against prevailing met...
205. SEISMO: Explanation-Aware, Trajectory-Conditioned LLM Agents for Sample-Efficient Molecular Optimisation ​
Author: Fabian P. Kr"uger, Andrea Hunklinger, Adrian Wolny, Tim J. Adler, Igor Tetko, Santiago David Villalba
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, q-bio.BM
arXiv:2602.00663v3 Announce Type: replace Abstract: Optimizing molecules to achieve desired properties is a central bottleneck across the chemical sciences, particularly in the pharmaceutical industry, where it underlies the discovery of new drugs. Since molecular property evaluation often relies on...
206. Can LLMs Introspect? A Reality Check ​
Author: Shashwat Singh, Tal Linzen, Shauli Ravfogel
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.26242v2 Announce Type: replace Abstract: Can large language models detect and report their own internal states? A number of recent studies have argued that they can. Drawing on lessons from human metacognition research, we argue that this conclusion may be premature. We identify two condi...
207. GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents ​
Author: Johannes Moll, Jean-Philippe Corbeil, Jiazhen Pan, Martin Hadamitzky, Daniel Rueckert, Lisa Adams, Keno Bressem
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2605.29668v3 Announce Type: replace Abstract: LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment. Prior self-improvement methods accumulate natural-language guidance without checki...
208. Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight ​
Author: Can Jin, Jiakang Li, Rui Wu, Eddy Z. Zhang, Dimitris N. Metaxas
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.00424v2 Announce Type: replace Abstract: As large language models become stronger, weak supervisors may fail to provide reliable labels, preferences, or final judgments for complex outputs, limiting both weak-to-strong generalization and scalable oversight. We study a more tractable form ...
209. Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence ​
Author: Fiona Y. Wang, Markus J. Buehler
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci, cs.CL, cs.LG, math.CT
arXiv:2606.01444v2 Announce Type: replace Abstract: Scientific discovery is not only answer generation but revision of the representational regime in which evidence, artifacts, operations, and verifiers are typed. We develop a category-theoretic account of agentic discovery for materials science. In...
210. SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment ​
Author: Hao Li, Jingkun An, Zijun Song, Pengyu Zhu, Rui Li, Hao Wang, Wendi Feng, Yesheng Liu, Lijun Li, Jin-Ge Yao, Lei Sha
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2606.02530v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax. Existing methods mitigate this by balancing dual objectives, which heavily rely on massive general-purpose data or auxiliar...
211. WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents ​
Author: Yehang Zhang, Jianchong Su, Haojian Huang, Yifan Chang, Tianhao Zhou, Xinli Xu, Yingjie Xu, Yinchuan Li, Zexi Li, Ying-Cong Chen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.18847v2 Announce Type: replace Abstract: To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions. Existing long-term memory benchmarks mainly evaluate language-centric retrieval and question answering, while em...
212. AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance ​
Author: Yang Zou, Zijian Ding, Yizhou Sun, Jason Cong
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.AR
arXiv:2606.30949v2 Announce Type: replace Abstract: High-Level Synthesis (HLS) provides a fast path from concepts to silicon, but converting real-world software into synthesizable HLS code remains challenging due to restrictive language support and the gap between software and hardware programming p...
213. The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs ​
Author: Baha Rababah, Shahzeb Qamar, Lorenz Sparrenberg, Rafet Sifa, Murat Kantarcioglu, Cuneyt Gurcan Akcora, Carson K. Leung
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.08734v2 Announce Type: replace Abstract: Post-Training Quantization has become widely used to compress large language models to make them deployable on resource-constrained devices. However, the evaluation of quantization methods mainly uses accuracy and perplexity, which cannot capture t...
214. When Words Are Safe But Actions Kill: Probing Physical Jailbreak Beyond Textual Jailbreak in Hidden-State Risk Space ​
Author: Weimeng Wang, Ziqiang Wang, Zihang Zhan, Chuanpu Fu, Qi Li, Ke Xu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2607.15218v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded jailbreak is the s...
215. Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation ​
Author: Ye Chen, Weining Zhang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.27984v2 Announce Type: replace Abstract: Agentic systems generate outputs faster than human review. We contrast two LLM evaluator specialization strategies: specialized judge weights, or rule-based deferral policies for safe judgment acceptance. On 99,952 rubric-conditioned samples, corre...
216. Fragility of Value under Imperfect Alignment ​
Author: Winter Cross, L'eo Cymbalista, Alfred Harwood, Jose Faustino
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.28881v4 Announce Type: replace Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imper...
217. MemWM: Memory-Augmented Text-Based World Model ​
Author: Yujun Wang, Tao Zhang, Jinhe Bi, Aniri, Wenxuan Ye, Boliang Liu, Sikuan Yan, Shuning Wang, Xuebing Zhou, S"oren Pirk, Hinrich Sch"utze, Yunpu Ma
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07107v2 Announce Type: replace Abstract: World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent next-state predictions can still omit task-critical facts, corrupt product attributes, or apply in...
218. TRACE-Memory: Public-Conditioned Retrieval and Utility-Aware Evidence Admission for Personalized Generation ​
Author: Jing Wang, Zhu Wang, Yifan Guo, Yulong Yang, Yunji Liang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.08446v2 Announce Type: replace Abstract: Personalized generation systems retrieve user history by request--memory relevance and inject it into the model context. Yet relevant history may concern the wrong preference aspect, duplicate public information, or provide insufficient support. We...
219. Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability ​
Author: Jeonghwan Choi, Taewon Yun, Minjeong Ban, Gyeonghun Sun, Jae-Gil Lee, Hwanjun Song
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11238v2 Announce Type: replace Abstract: Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of...
220. Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces ​
Author: Mengyu Chen, Feiyu Lu, Chun-Fu Chen, Lucas Vinh Tran, Jay Katukuri
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11354v2 Announce Type: replace Abstract: Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generativ...
221. Attributing Preprocessing Invariance in Spectral Foundation Models ​
Author: Dongjun Wei, Hongyi Wu, Yinuo Zou
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.LG
arXiv:2608.14227v2 Announce Type: replace Abstract: Preprocessing invariance is an appealing goal for spectral foundation models: a frozen model should remain useful when laboratories preprocess spectra differently. It is usually measured by training a classifier under one preprocessing pipeline and...
222. LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures ​
Author: Yunfei Zhang, Boyu Feng, Changhua Pei, Zexin Wang, Zhihuang Peng, Xinlong Liu, Hengyue Jiang, Difeng Ma, Jiayi Zhang, Yongzhou Yao, Yanan Zhao, Fei Sun, Yintong Huo, Zhaoyang Liu, Jingjing Li, Gaogang Xie, Dan Pei
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.15242v3 Announce Type: replace Abstract: When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and locali...
223. KernelArc: A Multi-Agent Framework for GPU Kernel Optimization ​
Author: Joyjit Kundu, Ben Stoffelen, Kaili Wang, Peter Vrancx, Ludovic Denoyer
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.PF
arXiv:2608.17071v2 Announce Type: replace Abstract: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard,...
224. The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations ​
Author: Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Ratna Kavuri, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18300v2 Announce Type: replace Abstract: LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has become a standard, scalable approach for accelerating and extending costly human evaluation. However, most work tr...
225. FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents ​
Author: Tianyou Wang, Chongyang Gao, Kezhen Chen, Dong Chen, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18423v2 Announce Type: replace Abstract: Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. F...
226. DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation ​
Author: Siyuan Ma, Boshi Zhang, Yutian Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Qiaojun Yu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.20114v2 Announce Type: replace Abstract: Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion f...
227. Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation ​
Author: Gijs Kassenaar, Zhao Yang, Vincent Fran\c{c}ois-Lavet
Published: 8/24/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.20256v2 Announce Type: replace Abstract: Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adaptive one, which can lead to over-computation on easy problems and insufficient computation on difficult ones. W...
228. AI-driven Prices for Externalities and Sustainability in Production Markets ​
Author: Panayiotis Danassis, Aris Filos-Ratsikas, Haipeng Chen, Milind Tambe, Boi Faltings
Published: 8/24/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.GT
arXiv:2106.06060v4 Announce Type: replace-cross Abstract: Traditional competitive markets do not account for negative externalities; indirect costs that some participants impose on others, such as the cost of over-appropriating a common-pool resource (which diminishes future stock, and thus harvest,...
229. Graphon Particle Systems, Part II: Dynamics of Distributed Stochastic Continuum Optimization ​
Author: Yan Chen, Tao Li, Xiaofeng Zong
Published: 8/24/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY, math.OC, math.PR
arXiv:2407.02765v4 Announce Type: replace-cross Abstract: We study the distributed optimization problem over a graphon with a continuum of nodes, which is regarded as the limit of the distributed networked optimization as the number of nodes goes to infinity. Each node has a private local cost funct...
230. On the Within-class Variation Issue in Alzheimer's Disease Detection ​
Author: Jiawen Kang, Dongrui Han, Lingwei Meng, Jingyan Zhou, Jinchao Li, Xixin Wu, Helen Meng
Published: 8/24/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL, cs.LG, cs.SD, q-bio.NC
arXiv:2409.16322v4 Announce Type: replace-cross Abstract: Alzheimer's Disease (AD) detection commonly employs machine learning classification models to distinguish between individuals with AD and those without. Different from conventional classification tasks, AD detection involves substantial withi...
231. Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI ​
Author: Jos'e Antonio Siqueira de Cerqueira, Mamia Agbese, Rebekah Rousi, Nannan Xi, Juho Hamari, Pekka Abrahamsson
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2411.08881v3 Announce Type: replace-cross Abstract: AI-based systems, including Large Language Models (LLMs), impact millions by supporting diverse tasks but face issues like misinformation, bias, and misuse. AI ethics is crucial as new technologies and concerns emerge, but objective, practica...
232. An Automated Pipeline for Few-Shot Bird Call Classification: A Case Study with the Tooth-Billed Pigeon ​
Author: Abhishek Jana, Moeumu Uili, James Atherton, Mark O'Brien, Joe Wood, Leandra Brickson
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, cs.SD
arXiv:2504.16276v3 Announce Type: replace-cross Abstract: This paper presents a largely automated one-shot bird call classification pipeline, incorporating targeted manual quality control steps, designed for rare species absent from large publicly available classifiers like BirdNET and Perch. While ...
233. SPD Matrix Learning for Neuroimaging Analysis: Perspectives, Methods, and Challenges ​
Author: Ce Ju, Reinmar Kobler, Antoine Collas, Motoaki Kawanabe, Cuntai Guan, Bertrand Thirion
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.IV, q-bio.NC
arXiv:2504.18882v3 Announce Type: replace-cross Abstract: Neuroimaging provides essential tools for characterizing brain activity, structure, and connectivity through modalities that capture complementary aspects of brain organization. Across these diverse modalities, a unifying perspective arises w...
234. Explaining Intrinsic Moral Self-Correction with Mechanistic Interpretability ​
Author: Yu-Ting Lee, Fu-Chieh Chang, Yu-En Shu, Hui-Ying Shih, Pei-Yuan Wu
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2505.11924v4 Announce Type: replace-cross Abstract: Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely through prompting. While effective across diverse tasks, its mechanism remains unclear. We hypothesize ...
235. WeedNet: A Foundation Model-Based Global-to-Local AI Approach for Real-Time Weed Species Identification and Classification ​
Author: Yanben Shen, Timilehin T. Ayanlade, Venkata Naresh Boddepalli, Mojdeh Saadati, Ashlyn Rairdin, Zi K. Deng, Muhammad Arbab Arshad, Aditya Balu, Daren Mueller, Asheesh K Singh, Wesley Everman, Nirav Merchant, Baskar Ganapathysubramanian, Meaghan Anderson, Soumik Sarkar, Arti Singh
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2505.18930v2 Announce Type: replace-cross Abstract: Early weed identification is crucial for effective management and control, and researchers, agronomists, and technology developers are increasingly interested in automating this process using computer vision and artificial intelligence; howev...
236. Can you see how I learn? Human observers' inferences about Reinforcement Learning agents' learning processes ​
Author: Bernhard Hilpert, Muhan Hou, Kim Baraka, Joost Broekens
Published: 8/24/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.RO
arXiv:2506.13583v2 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) agents often exhibit learning behaviors that are not intuitively interpretable by human observers, which can result in suboptimal feedback in collaborative teaching settings. Yet, how humans perceive and interpret ...
237. GeoExplain: Multimodal Reasoning based on Hierarchy of Visual Information in Street View ​
Author: Fenghua Cheng, Jinxiang Wang, Sen Wang, Zi Huang, Xue Li
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MM
arXiv:2506.16633v3 Announce Type: replace-cross Abstract: Multimodal reasoning is a process of understanding, integrating and inferring information across different data modalities. It has recently attracted surging academic attention. Although there are various tasks for evaluating multimodal reaso...
238. CulTrace: Tracing Internal Cultural Reasoning in Large Language Models ​
Author: Haeun Yu, Arnav Arora, Seogyeong Jeong, Nadav Borenstein, Siddhesh Pawar, Jisu Shin, Jiho Jin, Junho Myung, Alice Oh, Isabelle Augenstein
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2508.08879v4 Announce Type: replace-cross Abstract: The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of different cultures. Prior work has evaluated cultural awareness in LLMs by analys...
239. AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning ​
Author: Can Jin, Yang Zhou, Qixin Zhang, Hongwu Peng, Di Zhang, Zihan Dong, Marco Pavone, Ligong Han, Zhang-Wei Hong, Tong Che, Dimitris N. Metaxas
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2508.14313v4 Announce Type: replace-cross Abstract: Test-time scaling strategies for Large Language Models predominantly rely on either reinforcement learning with sparse outcome rewards or search-based methods guided by static Process Reward Models. However, outcome-based RL often suffers fro...
240. SCOPE: A Generative Approach for LLM Prompt Compression ​
Author: Tinghui Zhang, Yifan Wang, Daisy Zhe Wang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2508.15813v2 Announce Type: replace-cross Abstract: A big issue in modern LLM applications is they tend to feed long context to LLM, which results in high inference cost and latency, and may exceed the context limit. Prompt compression addresses this issue by reducing the length of input conte...
241. SKILL-RAG: Self-Knowledge Induced Learning and Filtering for Retrieval-Augmented Generation ​
Author: Tomoaki Isoda
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2509.20377v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) has significantly improved the performance of large language models (LLMs) on knowledge-intensive tasks in recent years. However, since retrieval systems may return irrelevant content, incorporating such i...
242. Perseus: Interactive Time Series Segmentation with Sparse Supervision via Stateful Memory ​
Author: Ching Chang, Ming-Chih Lo, Chiao-Tung Chan, Wen-Chih Peng, Tien-Fu Chen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2510.09930v2 Announce Type: replace-cross Abstract: Real-world systems, ranging from industrial manufacturing to wearable healthcare, generate multivariate time series with hierarchical states ranging from coarse regimes to fine-grained events. Unlike zero- or few-shot segmentation, our settin...
243. Significant Other AI: Identity, Memory, and Emotional Regulation as Long-Term Relational Intelligence ​
Author: Sung Park
Published: 8/24/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2512.00418v3 Announce Type: replace-cross Abstract: Significant Others (SOs) stabilize identity, regulate emotion, and support narrative meaning-making, yet many people today lack access to such relational anchors. Recent advances in large language models and memory-augmented AI raise the ques...
244. Fine-tuning an ECG Foundation Model to Predict Coronary CT Angiography Outcomes ​
Author: Yujie Xiao, Qinghao Zhao, Gongzheng Tang, Hao Zhang, Zhuoran Kan, Deyun Zhang, Jun Li, Guangkun Nie, Xiaocheng Fang, Haoyu Wang, Shun Huang, Tong Liu, Jian Liu, Kangyin Chen, Shenda Hong
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2512.05136v4 Announce Type: replace-cross Abstract: Coronary artery disease (CAD) remains a major global public health burden, yet scalable pre-imaging risk stratification tools are limited. In this multicenter study, we developed and validated an artificial intelligence-enabled electrocardiog...
245. MeltwaterBench: Deep learning for spatiotemporal downscaling of surface meltwater ​
Author: Bj"orn L"utjens, Patrick Alexander, Raf Antwerpen, Til Widmann, Guido Cervone, Marco Tedesco
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, physics.ao-ph, physics.data-an
arXiv:2512.12142v2 Announce Type: replace-cross Abstract: The Greenland ice sheet is melting at an accelerated rate due to processes that are not fully understood and hard to measure. The distribution of surface meltwater can help understand these processes and is observable through remote sensing, ...
246. AgentOCR: Reimagining Agent History via Optical Self-Compression ​
Author: Lang Feng, Fuchao Yang, Feng Chen, Xin Cheng, Haiyang Xu, Zhenglin Wan, Ming Yan, Bo An
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.04786v3 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) enable agentic systems trained with reinforcement learning (RL) over multi-turn interaction, but practical deployment is bottlenecked by rapidly growing textual histories that inflate token and ...
247. GroupSegment-SHAP: Shapley Value Explanations with Group-Segment Players for Multivariate Time Series ​
Author: Jinwoong Kim, Sangjin Park
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.GT
arXiv:2601.06114v2 Announce Type: replace-cross Abstract: Multivariate time-series models achieve strong predictive performance in healthcare, industry, energy, and finance, but how they combine cross-variable interactions with temporal dynamics remains unclear. SHapley Additive exPlanations (SHAP) ...
248. CFM: Language-aligned Concept Foundation Model for Vision ​
Author: Kai Wittenmayer, Sukrut Rao, Amin Parchami-Araghi, Bernt Schiele, Jonas Fischer
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2601.13798v3 Announce Type: replace-cross Abstract: Language-aligned vision foundation models perform strongly across diverse downstream tasks. Yet, their learned representations remain opaque, making interpreting their decision-making difficult. Recent work decompose these representations int...
249. Investigating Target Class Influence on Neural Network Compressibility for Energy-Autonomous Avian Monitoring ​
Author: Nina Brolich, Simon Geis, Maximilian Kasper, Alexander Barnhill, Axel Plinge, Dominik Seu{\ss}
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.17751v2 Announce Type: replace-cross Abstract: Biodiversity loss poses a significant threat to humanity, making wildlife monitoring essential for assessing ecosystem health. Avian species are ideal subjects for this due to their popularity and the ease of identifying them through their di...
250. Mind the Style: Impact of Communication Style on Human-Chatbot Interaction ​
Author: Erik Derner, Dalibor Ku\v{c}era, Aditya Gulati, Ayoub Bagheri, Nuria Oliver
Published: 8/24/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL, cs.CY
arXiv:2602.17850v2 Announce Type: replace-cross Abstract: Conversational agents increasingly mediate everyday digital interactions, yet the effects of their communication style on user experience and task success remain insufficiently understood. Addressing this gap, we report a between-subject user...
251. PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval ​
Author: Tianyi Xu, Rong Shan, Junjie Wu, Jiadeng Huang, Teng Wang, Jiachen Zhu, Wenteng Chen, Minxin Tu, Quantao Dou, Zhaoxiang Wang, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin
Published: 8/24/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CV, cs.MM
arXiv:2603.01493v2 Announce Type: replace-cross Abstract: Personal photo albums are not merely collections of static images but living, ecological archives defined by temporal continuity, social entanglement, and rich metadata, which makes the personalized photo retrieval non-trivial. However, exist...
252. Efficient Self-Evaluation for Diffusion Language Models via Sequence Regeneration ​
Author: Linhao Zhong, Linyu Wu, Wen Wang, Yuling Xi, Chenchen Jing, Jiaheng Zhang, Hao Chen, Chunhua Shen
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.02760v2 Announce Type: replace-cross Abstract: Diffusion large language models (dLLMs) have recently attracted significant attention for their ability to enhance diversity, controllability, and parallelism. However, their non-sequential, bidirectionally masked generation makes quality ass...
253. Efficient Exploration at Scale ​
Author: Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla, Xiuyuan Lu, Mehdi Jafarnia, Victor Minden, Zheng Wen, Benjamin Van Roy
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2603.17378v2 Announce Type: replace-cross Abstract: We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). Our algorithm incrementally updates reward and language models as choice data is received. The reward...
254. InverFill: One-Step Inversion for Enhanced Few-Step Diffusion Inpainting ​
Author: Duc Vu, Kien Nguyen, Trong-Tung Nguyen, Ngan Nguyen, Phong Nguyen, Khoi Nguyen, Cuong Pham, Anh Tran
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.23463v2 Announce Type: replace-cross Abstract: Recent diffusion-based models achieve photorealism in image inpainting but require many sampling steps, limiting practical use. Few-step text-to-image models offer faster generation, but naively applying them to inpainting yields poor harmoni...
255. DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis ​
Author: Cheng-You Lu, Yi-Shan Hung, Wei-Ling Chi, Hao-Ping Wang, Charlie Li-Ting Tsai, Yu-Cheng Chang, Yu-Lun Liu, Thomas Do, Chin-Teng Lin
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.13416v4 Announce Type: replace-cross Abstract: Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been developed to support comprehensive benchmarking and to facilitate progress beyond scene-specific recon...
256. ChemGraph-XANES: An Agentic Framework for XANES Simulation and Curation ​
Author: Vitor F. Grizzi, Thang Duc Pham, Luke N. Pretzie, Jiayi Xu, Murat Keceli, Cong Liu
Published: 8/24/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, physics.chem-ph
arXiv:2604.16205v3 Announce Type: replace-cross Abstract: Computational X-ray absorption near-edge structure (XANES) is widely used to interpret local coordination environments, oxidation states, and electronic structure, but large computational campaigns are often limited by workflow complexity. We...
257. AutoOR: Scalably Post-training LLMs to Autoformulate Operations Research Problems ​
Author: Sumeet Ramesh Motwani, Chuan Du, Aleksander Petrov, Christopher Davis, Philip Torr, Antonio Papania-Davis, Weishi Yan
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.16804v4 Announce Type: replace-cross Abstract: Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating complicated descriptions of these problems into solver-ready formulations requires specialized operations...
258. Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Long-Horizon Workflows ​
Author: Luay Gharzeddine, Samer Saab Jr
Published: 8/24/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2604.22820v2 Announce Type: replace-cross Abstract: Long-horizon tool-using tasks sometimes benefit from revisiting earlier subtasks, but explicit revisitation also adds routing, coordination, and token cost. We study complete cyclic subtask graphs for large language model (LLM) agents: a work...
259. RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs ​
Author: Sadia Asif, Mohammad Mohammadi Amiri
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, cs.CL, cs.CR
arXiv:2605.01913v2 Announce Type: replace-cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable to adversarial misuse. While prior work has shown that safety-relevant features are encoded i...
260. Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization ​
Author: Xin Yu, Liuchen Liao, Yiwen Zhang, Yingchen Yu, Lingzhou Xue, Qinzhen Guo
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.05040v2 Announce Type: replace-cross Abstract: On-policy distillation is an efficient alternative to reinforcement learning, offering dense token-level training signals. However, its reliance on a stronger external teacher has driven recent work on on-policy self-distillation, where the s...
261. GRALIS: Fusing Coalition and Gradient Attribution with Closed-Form Conservation Error and Finite-Sample Guarantees ​
Author: Raimondo Fanale
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2605.05480v3 Announce Type: replace-cross Abstract: The main post-hoc XAI methods for deep networks -- GradCAM, SHAP, LIME, Integrated Gradients -- originate from heterogeneous theoretical foundations and are not naturally comparable within a single representation. A recent benchmark also find...
262. S-AI-Recursive: Convergent Recursive Reasoning ​
Author: Said Slaoui
Published: 8/24/2026, 4:00:00 AM
Categories: cs.NE, cs.AI
arXiv:2605.13872v2 Announce Type: replace-cross Abstract: This article introduces S-AI-Recursive, a bio-inspired Sparse Artificial Intelligence architecture in which reasoning is implemented as a hormonally regulated closed-loop iteration rather than a single feed-forward pass. The Recursive Reasoni...
263. Component-Aware Structure-Preserving Style Transfer for Satellite Visual Sim2Real Data Construction ​
Author: Zongwu Xie, Yonglong Zhang, Yifan Yang, Yang Liu, Baoshi Cao, Guanghu Xie
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2605.19624v3 Announce Type: replace-cross Abstract: For camera-based satellite visual sensing, Sim2Real data construction requires images that approach real-domain sensor appearance while retaining the annotations inherited from simulation. Real sensor images of satellite targets with reliable...
264. Behavior-Consistent Deep Reinforcement Learning ​
Author: Marcel Hussing, Liv G. d'Aliberti, Claas Voelcker, Benjamin Eysenbach, Eric Eaton
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.21214v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) often exhibits high variance across training runs, leading to unreliable performance and posing a major challenge to deployment in real-world domains. In this work, we address the challenge of cross-run policy dive...
265. Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS) ​
Author: Samer Awad, Javier Conde, Carlos Arriaga, Tairan Fu, Javier Coronado-Bl'azquez, Pedro Reviriego
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.27268v2 Announce Type: replace-cross Abstract: Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabularies. While previous research has focused on model knowledge and training data, we investigate the ...
266. Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS ​
Author: Deokjin Seo, Gangin Park, Kihyun Nam
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, eess.AS
arXiv:2605.30748v3 Announce Type: replace-cross Abstract: We present Chatterbox-Flash, a zero-shot text-to-speech model obtained by fine-tuning a pretrained autoregressive TTS decoder into a block-diffusion decoder, enabling parallel token generation within each block while retaining block-by-block ...
267. Audio Interaction Model ​
Author: Zhifei Xie, Zihang Liu, Ze An, Xiaobin Hu, Yue Liao, Ziyang Ma, Dongchao Yang, Mingbao Lin, Deheng Ye, Shuicheng Yan, Chunyan Miao
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL, cs.MM, eess.AS
arXiv:2606.05121v2 Announce Type: replace-cross Abstract: Audio is continuous and interactive, yet most Large Audio Language Models (LALMs) remain offline and streaming systems usually specialize in ASR or spoken dialogue. We formalize the Audio Interaction Model, an always-on perceive--decide--resp...
268. INFUSER: Influence-Guided Self-Evolution Improves Reasoning ​
Author: Siyu Chen, Miao Lu, Beining Wu, Heejune Sheen, Fengzhuo Zhang, Shuangning Li, Zhiyuan Li, Jose Blanchet, Tianhao Wang, Zhuoran Yang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.GT, stat.ML
arXiv:2606.09052v4 Announce Type: replace-cross Abstract: Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or,...
269. Valid Inference with Synthetic Data via Task Exchangeability ​
Author: Lezhi Tan, Tijana Zrnic
Published: 8/24/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, cs.LG, stat.ML
arXiv:2606.13629v2 Announce Type: replace-cross Abstract: There is a proliferation of work arguing for the use of synthetic data in scientific research. For example, social scientists are arguing for the use of LLM-generated "silicon samples" in pilot studies; AI evaluations increasingly rely on "LL...
270. Crypto x AI, AI x Crypto: A Survey ​
Author: Sarah Allen, Pranay Anchuri, James Austgen, Maryam Bahrani, Samuel Breckenridge, Aaron Buchwald, Christian Cachin, Andr'es F'abrega, Jared Fernandez, James Hsin-yu Chiang, Marwa Mouallem, Roi Bar-Zur, Neil DeSilva, Ittay Eyal, Giulia Fanti, Ari Juels, Andrew Miller, Christian Sillaber, Dani Vilardell, Pramod Viswanath, Wenhao Wang, Matt Weinberg, Sen Yang, Jianzhu Yao, Fan Zhang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2606.13892v2 Announce Type: replace-cross Abstract: The intersection of crypto x AI is spawning papers, products, online posts, and companies. All the surrounding buzz, though, obscures what exactly has been done, what the opportunities and challenges are, and what open questions deserve atten...
271. The Metanym Game: An LLM Benchmark Without Ground Truth That Rises With the Models It Measures ​
Author: David Nordfors
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2606.21008v3 Announce Type: replace-cross Abstract: We present evidence that analogy is at the core of LLM intelligence. In our benchmark, LLMs compete in generating sets of analogous statements and rate each other's sets on their own understandings of factual correctness, beauty, intelligence...
272. Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models ​
Author: Renwei Meng, Bowen Zhang, Jian Wang, Xican Wang, Haoyi Wu, Xuanyan Qiu, Shengan Yang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.26101v2 Announce Type: replace-cross Abstract: Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior. We present a contamination-aware, ...
273. MatMMExtract: An Open-Source Pipeline for Panel-Level Extraction of Grounded Image-Text Pairs from Materials Science Literature ​
Author: Subham Ghosh, Shubham Tiwari, Mohammad Ibrahim, Abhishek Tewari
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cond-mat.mtrl-sci, cs.AI
arXiv:2606.29667v2 Announce Type: replace-cross Abstract: The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inaccessible to AI at scale. The core difficulty is structural: most scientific figures are compound, with a...
274. The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry ​
Author: Yuan Yuan
Published: 8/24/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, math.DG
arXiv:2607.02368v2 Announce Type: replace-cross Abstract: Evaluations of LLM personas via psychometric questionnaires typically rely on aggregate scores, discarding within-instance correlation structure. We test whether this geometric structure is intrinsic or frame-dependent. Constructing within-in...
275. What a World Model Represents Is Three Questions ​
Author: Donna Vakalis
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.06640v2 Announce Type: replace-cross Abstract: World models learn task-relevant information through many routes: observation reconstruction, recurrent state, temporal filtering, and explicit task supervision. Different routes can make different variables available. The same variable can a...
276. MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents ​
Author: Kaixin Ma, Di Feng, Alexander Metz, Jiarui Lu, Eshan Verma, Afshin Dehghan
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.11818v2 Announce Type: replace-cross Abstract: We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, mu...
277. Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts ​
Author: Haodi Fan, Zucong Lan
Published: 8/24/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.18970v3 Announce Type: replace-cross Abstract: Agent Skills have become persistent behavioral artifacts across independent AI agent systems. They combine natural-language task specifications with metadata and optional references, scripts, assets, hooks, package manifests, tests, and compa...
278. A Distributional Robustness Margin For Pathology Foundation Models ​
Author: Cl'ement Grisi, Jeroen van der Laak, Geert Litjens
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.25497v4 Announce Type: replace-cross Abstract: Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut learning that undermines generalisation across institutions. The Robustness Index (RI) was proposed to asse...
279. Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation ​
Author: Fan Wu, Cuiyun Gao, Yiming Huang, Yang Xiao, Yujia Chen, Qing Liao
Published: 8/24/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2607.28645v2 Announce Type: replace-cross Abstract: Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildable codebase with shared components and working navigation. This project...
280. Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery ​
Author: Jiayu Gu, Yiwei Wang, Jie Zhang, Guojun Cao, Keshen Lyu, Song Zhou, Yimeng Chen, Haorui Wang, Qingmin Feng, Shenchao Shi, Hongkuan Shi, Qiuyu Yu, Qiang Xie, Huan Zhao, Wenbin Chen, Caihua Xiong, Chidan Wan, Jing Samantha Pan, Xiong Cai, Han Ding
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02471v2 Announce Type: replace-cross Abstract: In laparoscopy, surgeon gaze tracks where the instruments will act; easing this demand through visual attention modeling requires dense labels of those interaction loci. These encode tacit knowledge: experts converge on consensus loci yet str...
281. SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation ​
Author: Wen Wang, Jiahua Bao, Tu Yongsiqi, Yihao Liu, Haotian Zhou, Haoxuan Ma, Mengyu Zhou, Wenkui Fan, Junwei He, Xiaoxi Jiang, Guanjun Jiang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.03092v2 Announce Type: replace-cross Abstract: We aim to improve model performance in multi-reward reinforcement learning training process. Existing Group reward-Decoupled Normalization Policy Optimization (GDPO) has mitigated the issue of reward signals masking one another during direct ...
282. Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers ​
Author: Xiangyu Meng, Peng Chen, Mingzhen Li, Jianmin Wang, Sen Wang, Guangming Tan, Weile Jia, Mohamed Wahib, Tao Luo, Xun Wang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2608.07078v2 Announce Type: replace-cross Abstract: Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale. Among existing approaches, LightDock leverages Glowworm Swarm Optimization (GSO) for accuracy, yet su...
283. Defining Decentralization: An Ontological Perspective ​
Author: Jakub Kacper Szel\k{a}g, Aydin Abadi, Mohammad Naseri
Published: 8/24/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG, cs.LO, cs.SY, eess.SY
arXiv:2608.09748v2 Announce Type: replace-cross Abstract: Decentralization as a concept in computer science has existed for over half a century. Despite its fundamental role across domains such as security, distributed computing, artificial intelligence, cloud infrastructures, and Internet of Things...
284. ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation ​
Author: Ge Wang
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.10398v2 Announce Type: replace-cross Abstract: ELVAE places an input-dependent normal--inverse-gamma (NIG) hierarchy at each VAE latent coordinate, separating location uncertainty $u_{\mathrm{epi}}=\beta/[\nu(\alpha-1)]$ from conditional variability $u_{\mathrm{var}}=\beta/(\alpha-1)$. Th...
285. Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians ​
Author: Timothy Heightman, Elena Orlova, Philip Mantrov, Aleksei Ustimenko
Published: 8/24/2026, 4:00:00 AM
Categories: quant-ph, cond-mat.dis-nn, cond-mat.str-el, cs.AI
arXiv:2608.11911v3 Announce Type: replace-cross Abstract: A central promise of useful quantum advantage is the ability to compute ground states of Hamiltonian systems beyond the reach of classical simulation methods. Here we demonstrate that this problem can be effectively amortized across an arbitr...
286. DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data ​
Author: Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina, Kenneth Enevoldsen, Lukas Galke Poech
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.13517v2 Announce Type: replace-cross Abstract: Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language mod...
287. Information Geometry of Message Passing ​
Author: Mykola Lukashchuk, Kyrylo Yemets, Alex Ledbetter, .{I}smail \c{S}en"oz
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.15922v2 Announce Type: replace-cross Abstract: We show that the natural-gradient stationary condition of variational inference has an edge-local form on a Forney-style factor graph. We start from the Bethe free energy and constrain a selected edge marginal to an exponential family. At a s...
288. SuTRA : Structurally-Unified Tokenization with Root Awareness ​
Author: Vaibhav Rathore, Siddhant Gole, Dadhichi Telwadkar, Rooshil Bhatia, Maulik Ruparel, Siddharth Sureka, Neha Bhargava
Published: 8/24/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.18087v2 Announce Type: replace-cross Abstract: Existing subword tokenizers optimize statistical compression but ignore morphological structure, particularly the relationship between roots and affixes. This is harmful for morphologically rich Indic languages, where basic units are complex ...
289. CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning ​
Author: Julien Merand, Boris Meden, Liming Chen, Mathieu Grossard
Published: 8/24/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.19776v2 Announce Type: replace-cross Abstract: Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it should be grasped to support downstream functional tasks. However, conditioning grasp synthesis on spe...
290. Separating Covariate Shift from Mechanism Change with Two Discriminators: CJSD, a Conditional Discrepancy with an Exact Covariate-Concept Decomposition ​
Author: Kentaro Oda
Published: 8/24/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.19885v2 Announce Type: replace-cross Abstract: After the inputs X are known, how much additional information does the label Y carry about which dataset a sample came from? That single quantity -- estimable as the difference of two discriminators' held-out cross-entropies, D_CJS = CE(Z|X) ...