arXiv cs.AI - 2026-08-04 ​
710 items collected.
1. Revisiting Classic Thought Experiments to Measure Consciousness for Artificial Intelligence Safety ​
Author: Peter David Fagan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00001v1 Announce Type: new Abstract: This research note revisits Leibniz's mill, Turing's imitation game, and Searle's Chinese Room through the Conservation-Congruent Encoding (CCE) framework. It formalises a toy symbolic setting in which successful behaviour is measured by task performan...
2. AutoFOAM: The Self-Refining Autonomous OpenFOAM Agent ​
Author: Arun Govind Neelan, A Seshaditya
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00003v1 Announce Type: new Abstract: Computational Fluid Dynamics (CFD) plays an important role in modern engineering, but using open-source solvers such as OpenFOAM requires considerable knowledge and skills, as well as time-consuming configuration file setup. To reduce this burden, we p...
3. Enhancing LLMs with Context-Specific Knowledge for Mitigating Misinformation in SMEs: A RAG-based Modeling and Analysis ​
Author: Md. Samiul Islam, Iqbal H. Sarker, Chadni Islam, Ahmad Mohsin, Ahmed Ibrahim, Helge Janicke
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00006v1 Announce Type: new Abstract: Large Language Models (LLMs), a part of artificial intelligence (AI), are increasingly being adopted by Small and Medium Enterprises (SMEs) to enhance question-answering capabilities and support business decision-making processes. However, hallucinatio...
4. Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware ​
Author: Philipp M. Z"ahl, Anika Hennig
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, cs.PF
arXiv:2608.00008v1 Announce Type: new Abstract: The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference. However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on acc...
5. CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection ​
Author: Qihua Pan, Zhenheng Tang, Peijie Dong, Xiang Liu, Huacan Wang, Bo Li, Xiaowen Chu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00014v1 Announce Type: new Abstract: Evaluating Large Language Models (LLMs) incurs prohibitive computational overhead during continuous development processes. While coreset selection accelerates evaluation, existing methods either suffer from a severe ``cold start'' bottleneck requiring ...
6. Optimization and Constraint Modeling using LLMs with a Retrieval Augmented Generation Process ​
Author: Prateek Roy, Akash Singirikonda
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00015v1 Announce Type: new Abstract: Both optimization modeling and constraint modeling are non-trivial problems requiring deep domain expertise and proficiency in modeling formalism languages. Despite their importance across logistics, healthcare, and supply chain management, current lar...
7. Memory Reward Inflation in Self-Improving LLM Agents ​
Author: Mohammad Asadolahi, Amir Amini, Samira Talebi, Amirfarhad Farhadi, Azadeh Zamanifar
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CE
arXiv:2608.00017v1 Announce Type: new Abstract: Self-improving LLM agents increasingly learn from experience without updating any weights. Each episode is stored in an external memory, scored, and retrieved for similar future tasks to shape later behavior. Viewed through a reward lens, the stored sc...
8. Request-Level Energy Attribution for Batched LLM Serving ​
Author: Qi Luo, Kunlin Li, Ziwen Wang, Dongsheng Wang, Yun Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.DC
arXiv:2608.00026v1 Announce Type: new Abstract: Batched LLM serving improves throughput but complicates energy accounting. GPU power telemetry is aggregate, whereas sustainability reporting, chargeback, and workload analysis often require request-level energy charges. Existing inference-energy bench...
9. Motif-Mamba: network motif improved mamba for long-range sequence modeling ​
Author: Chonghe Hao, Yue Sun, Jian Zhang, Yansong Wang, Wangzi Yao, Yunjie Yao, Tielin Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00027v1 Announce Type: new Abstract: Efficient long-sequence modeling remains a central challenge for large language models, as self-attention scales quadratically with sequence length. Mamba offers a linear-time alternative through selective state space recurrence, but its predominantly ...
10. Nova: An End-to-End MLIR Compiler for Deep Learning ​
Author: Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.AR, cs.LG, cs.PL
arXiv:2608.00029v1 Announce Type: new Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware. While high-level tensor frameworks provide flexible abstractions for model design, their e...
11. SIRIN: A Unified Toolkit for Detecting Contextual Hallucinations in Retrieval-Augmented and Memory-Grounded LLM Systems ​
Author: Julia Belikova, Rauf Parchiev, Mikhail Filimonov, Konstantin Polev, Andrey Savchenko, Maksim Makarenko
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00033v1 Announce Type: new Abstract: SIRIN (Semantic Inconsistency Recognition and Inspection Nexus) is a unified toolkit and interactive web UI for detecting contextual hallucinations (fluent, plausible responses unsupported by the provided evidence) in retrieval-augmented, agentic, and ...
12. Linguistic Context Recodes Visual Representations in Vision-Language Models ​
Author: Brian Song, Michael A. Lepori, Ellie Pavlick
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.00035v1 Announce Type: new Abstract: Goal-directed visual processing is a hallmark of human visual intelligence, resulting in representations that support downstream tasks such as categorization or search. Though vision-language models (VLMs) are often faced with these same tasks, their a...
13. RAG-TESTER: Automated End-to-End Testing of Retrieval-Augmented Large Language Models ​
Author: Ange Maiztegi, Jon Ayerdi, Miren Illarramendi, Aitor Arrieta
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.00054v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) enables Large Language Models (LLMs) to use external and domain-specific knowledge, but its reliability depends on the interaction between the generative model, embedding model, retrieval mechanism, and prompt const...
14. H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases ​
Author: Shusen Zhang, Junyi Hu, Ye Feng, Ziteng Wang, Zhaoyuan Pan, Guosheng Dong, Xiaojun Yuan, Jiangshou Hong, Xiangzhi Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.00065v1 Announce Type: new Abstract: Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constraints, and compositional concepts. However, existing representations lie at two extremes: single-vector retriever...
15. Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale ​
Author: Banruo Liu, Haoran Qiu, 'I~nigo Goiri, Rodrigo Fonseca, Ricardo Bianchini, Esha Choukse
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.00101v1 Announce Type: new Abstract: AI coding agents like GitHub Copilot, Claude Code, and Codex interleave multi-step LLM inference with tool execution, creating a workload different from chatbots. We present the first production-scale characterization of this workload using sampled Git...
16. Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce ​
Author: Shimaa Ahmed, Yiwei Cai, Mohsen Minaei, Rahul Rachuri
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.00102v1 Announce Type: new Abstract: Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price comp...
17. Shared Organizational Memory for Enterprise Coding Agents: System Design and Deployment Snapshot ​
Author: Harsh Rao Dhanyamraju, Leonidas Raghav
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00122v1 Announce Type: new Abstract: Enterprise coding agents rely on tools and retrieval, yet enterprise knowledge often remains outside public training data and formal documentation: internal DSLs, proprietary platforms, local conventions, recent fixes, and tacit workflows. Existing kno...
18. AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? ​
Author: Dong Yan, Jian Liang, Dapeng Hu, Ran He, Nicholas Jing Yuan, Qi Zhang, Tieniu Tan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.00155v1 Announce Type: new Abstract: Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic stre...
19. TRACE-TS: Attribution-Grounded and Traceable Sensor-Language Reasoning for Human Activity Understanding ​
Author: Sparsh Rastogi, Tanmay Kumar, Baiyu Chen, Jatin Bedi, Zechen Li, Flora D. Salim
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.ET, cs.LG
arXiv:2608.00200v1 Announce Type: new Abstract: Wearable sensors capture fine-grained motion patterns that support rich behavioral understanding, yet most existing methods reduce these signals to activity labels. Recent LM-based approaches generate natural-language explanations for sensor data, but ...
20. Personalizing Large Language Model Agents with Small Policy Models ​
Author: Dian Jin, Zhi Zhang, Huichao Li, Yihe Pan, Rundong Huang, Doudou Zhou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00215v1 Announce Type: new Abstract: Large language model (LLM) agents can retrieve memory, call tools, ask clarifying questions, and vary response style, yet adapting these execution decisions to an individual user remains difficult. Fine-tuning a separate LLM is costly or impossible for...
21. More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness ​
Author: Yuelyu Ji
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00243v1 Announce Type: new Abstract: Large language model (LLM) judges are increasingly organized as multi-agent panels under the assumption that exchanging critiques improves judgment quality. We test this assumption for \emph{groundedness verification}, where a judge must determine whet...
22. Geometric Self-Supervised Pre-training for Neural Combinatorial Optimization ​
Author: David Aguado, Daniel Fuertes, Carlos R. del-Blanco, Fernando Jaureguizar
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00270v1 Announce Type: new Abstract: Neural Combinatorial Optimization (NCO) techniques have emerged as a highly efficient alternative to traditional exact algorithms for solving routing problems such as the Traveling Salesman Problem (TSP). However, the generalization capabilities of the...
23. RF-HOI: Recognize Human-Object Interaction with Radio Frequency Signals ​
Author: Lihao Wang, Linlu Gao, Jiacan Yu, Yanyu Lin, Yifan Yin, Jianxin Wang, Tianmin Shu, Renjie Zhao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.RO
arXiv:2608.00289v1 Announce Type: new Abstract: Recognizing Human-Object Interactions (HOI) is essential for intelligent systems, underpinning applications in virtual and augmented reality, embodied AI, and assistive robotics. However, vision-based HOI methods face challenges in privacy concerns and...
24. WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation ​
Author: Jianxun Cui, Ping Wu, Stanisa Peric, Marko Milojkovic, Vladan Devedzic
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00298v1 Announce Type: new Abstract: World models and generative simulators are emerging as interactive testing infrastructure for autonomous driving because they can react to the ego planner and produce counterfactual, rare, and safety-critical rollouts. This changes a test scenario from...
25. CrystalMem: Elastic Memory for Self-Evolving LLM Agents via Knowledge Crystallization ​
Author: Beining Wu, Jun Huang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00303v1 Announce Type: new Abstract: Memory for self-evolving large language model (LLM) agents is often provisioned as if its byte budget only grows. Cloud platforms, however, adjust quotas with load and cost, and we show that capability does not follow the budget back up: after a squeez...
26. Trust and Its Betrayal under Three Representational Strategies ​
Author: Mihnea C. Moldoveanu, Joel A. C. Baum
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00321v1 Announce Type: new Abstract: Trust is a propositional attitude of a distinctive kind: to trust is to rely on another under conditions where reliance could be disappointed, and the disappointment of trust---betrayal---differs qualitatively from the disappointment of a prediction. W...
27. Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates ​
Author: Bohan Chen, Shivam N. Patel, Richard Hoffmann, Sam Looi, Tony Yue Yu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00326v1 Announce Type: new Abstract: Tool calling allows large language models (LLMs) to invoke external computation during problem solving, a useful capability in various fields including AI for mathematics. We study this setting through weighted sum-of-squares (SOS) decomposition, a mac...
28. RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning ​
Author: Chengbo Liu, Lifang Zhou, Ruijie Yan, Pei Tan, Ao Sun, Haojun Huang, Guichun Hua, Sining Wei, Yining Chen, Yingying He, Yutao Xie
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.00335v1 Announce Type: new Abstract: Compact web agents can reduce deployment cost, but training them poses challenges in both data collection and post-SFT reinforcement learning (RL). Successful trajectories are expensive to collect and often contain inefficient detours. After supervised...
29. Bayesian and Motivated Reasoning in AI Agents ​
Author: Eddie Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.00339v1 Announce Type: new Abstract: AI agents increasingly perform open-ended tasks in settings where their conclusions can guide consequential decisions. We provide evidence that AI agents draw different conclusions from identical numerical data when the substantive framing changes. We ...
30. Gene Ontology-Guided Hierarchical Spatial Gene Expression Prediction from Histopathology Images ​
Author: Zhiwen Xu, Xiaoming Yan, Chengkun Wu, Juan Chen, Haoang Chi, Liyang Xu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, q-bio.GN
arXiv:2608.00405v1 Announce Type: new Abstract: Predicting spatial gene expression from histopathology images enables large-scale transcriptomic profiling without the cost of direct measurement. Existing methods decode the target gene set as a flat, unstructured vector, ignoring the inter-gene depen...
31. Where did the ambiguity go? Examining how multimodal models interpret polysemous words ​
Author: Jasin Cekinmez, Addison J. Wu, Raja Marjieh, Thomas L. Griffiths
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV
arXiv:2608.00410v1 Announce Type: new Abstract: Human language is highly polysemous. Many common words (e.g., 'bank' or 'palm') carry several distinct meanings that shape what humans communicate and imagine. Large language models (LLMs) have been shown to understand this multiplicity of meaning, but...
32. SymboUQ: Symbolic Uncertainty Quantification for Spatial Reasoning in LLMs ​
Author: Dahai Yu, Lin Jiang, Rongchao Xu, Guang Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00417v1 Announce Type: new Abstract: Although large language models (LLMs) can produce fluent spatial reasoning traces, their intermediate relations may fail to support the final conclusion, making token-level confidence insufficient for final-answer reliability estimation. Existing forma...
33. Mask-Based Priors Are More Persistent than Query-Key Initializations ​
Author: Mingze Ma, Hemanth Saratchandran, Cameron Gordon, Simon Lucey
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00418v1 Announce Type: new Abstract: Transformers do not merely lack data on some Boolean extrapolation tasks; they generalize in a systematically wrong way. Recent work on generalization on the unseen has shown that, despite fitting the observed domain, Transformers often extrapolate acc...
34. TrAC: Trace-Conditioned Answer Consistency for Efficient Uncertainty Quantification in LLMs ​
Author: Dahai Yu, Lin Jiang, Rongchao Xu, Guang Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00422v1 Announce Type: new Abstract: Large language models (LLMs) can generate fluent reasoning traces that nevertheless lead to incorrect answers, making response-level uncertainty estimation important for abstention, human review, and adaptive compute allocation. Existing approaches gen...
35. Diagnose Before You Compress: Prediction-Independent Bottleneck Witness Refinement for LLM Serving Traces ​
Author: Liming Liu, Chao Hu, Mingfei Lu, Cong Tan, Yiwei Ge, Chijin Zhou, Yongjun Xie, Runzhe Wang, Xiaohai Shi, Heyuan Shi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00423v1 Announce Type: new Abstract: Production LLM serving generates millions of diverse requests, making full-trace replay across serving configurations increasingly expensive. Existing trace reduction methods mainly preserve workload distributions or representative requests, but bottle...
36. Ekova: A Personality-Support Agent for Self-Discovery Dialogue ​
Author: Yuyan Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2608.00478v1 Announce Type: new Abstract: Emotional Support (ES) systems have long optimized a single objective: alleviating the user's emotional distress in the moment. We argue that a complementary need, helping users see themselves more clearly, defines a distinct paradigm we call Personali...
37. F-WANDA: Fisher-Reweighted Post-Training Pruning for Sustainable Deployment of Large Language Models ​
Author: Himanshu Mishra (University of British Columbia)
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00481v1 Announce Type: new Abstract: One-shot post-training pruning is the most energy-frugal compression strategy for largelanguage models (LLMs), yet existing approaches trade either quality (WANDA) or compute cost (SPARSEGPT). We introduce F-WANDA, a drop-in modification of WANDA that ...
38. The Bayesian Reflex: A Predictive Coding Engine for Artificial Intelligence ​
Author: Sourabh Bhattacharya
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.00492v1 Announce Type: new Abstract: Predictive coding offers a powerful theory of cortical computation, but corresponding scalable algorithmic implementations for artificial intelligence have remained elusive. This paper introduces the Bayesian reflex, a computational framework that dire...
39. TaPR: Test-Aware Policy Refinement for Feedback-Conditioned Code Generation ​
Author: Aofan Liu, Jingxiang Meng, Fangxin Liu, Yongbiao Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00494v1 Announce Type: new Abstract: Multi-turn code agents rely on execution feedback to repair incorrect programs, yet standard reinforcement learning paradigms optimize and evaluate policy performance primarily using single-shot outcome rewards. This misalignment conflates initial code...
40. BayesSeg: A Bayesian Optimization Framework for State Segmentation of Electricity Consumption Time Series ​
Author: Zhenya Zhang, Wendi Zhu, Ping Wang, Hongmei Cheng, Shuguang Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, physics.data-an
arXiv:2608.00513v1 Announce Type: new Abstract: In Non-Intrusive Load Monitoring (NILM), adaptive segmentation of electricity consumption time series is critical for appliance recognition. However, prevailing methods face challenges including heuristic parameter tuning, boundary sensitivity, and met...
41. CURE: Local Uncertainty Repair for Block-Parallel Speculative Decoding ​
Author: Aofan Liu, Jingxiang Meng, Fangxin Liu, Yongbiao Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00531v1 Announce Type: new Abstract: Speculative decoding mitigates the latency of sequential generation in autoregressive Large Language Models (LLMs) by interleaving draft generation with target verification. However, existing parallel drafting backends often suffer from rapid accuracy ...
42. Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations ​
Author: Shalom Kachko, Raz Lapid, Margarita Vald, Almog Dubin, Moshe Sipper
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV
arXiv:2608.00561v1 Announce Type: new Abstract: Vision-language models (VLMs) process image patches and text tokens in a shared residual stream, but the local geometry through which the two modalities interact remains poorly understood. Most interpretability methods identify global linear directions...
43. Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Models ​
Author: Yibin Dong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00591v1 Announce Type: new Abstract: Stochastic world models are usually evaluated by the accuracy and calibration of their predicted futures. These criteria leave a decision-relevant ambiguity: the same conditional future distribution can arise because an observation aliases different ph...
44. Escaping Confidence Trap: Evolutionary Decoding for Mathematical Reasoning in Diffusion LLMs ​
Author: Zhenhong Sun, Hanqing Zhao, Yatao Bian, Rongcheng Tu, Liuyue Xie, Xu Zhang, Jue Wang, Davide Modolo, Daoyi Dong, Dacheng Tao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00605v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) have emerged as a promising alternative to autoregressive LLMs, offering efficient generation through block-wise progressive unmasking. However, their strong general-purpose performance does not necessarily trans...
45. Slides2MindMap: Reconstructing Cognitively Efficient Knowledge Hierarchies from Lecture Slides ​
Author: Yuzhi Wang, Rongjun Ye, Shengyuan Chen, Huachi Zhou, Jiaqi Bai, Chuang Zhou, Zhicong Hong, Xiao Huang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00610v1 Announce Type: new Abstract: Generating mind maps from lecture slides can help learners efficiently assimilate fragmented knowledge, promising substantial benefits for intelligent education. However, dedicated automatic generation and evaluation frameworks remain underexplored and...
46. DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimization ​
Author: Changquan Zhao, Yuxiang Sun, Ruihao Zhu, Cheng Hua, Yulian He
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00641v1 Announce Type: new Abstract: Bayesian optimization (BO) relies on a surrogate model and an acquisition function, yet the most suitable choices vary across tasks and optimization stages. Automated Bayesian optimization (AutoBO) addresses this variability by adapting BO components o...
47. HetGPS: Scalable Graph Multi-Agent Reinforcement Learning with Physics-Anchored Adaptive Safety for EV Charging ​
Author: Xiangwei Wang, Nanduni Nimalsiri, Yu Xia, Peng Wang, Saman Halgamuge
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.00679v1 Announce Type: new Abstract: Safety interventions for large populations of network-coupled agents must protect shared constraints without unnecessarily overriding task-oriented policy decisions. We present HetGPS, a hybrid graph-control framework synergizing learned graph risk wit...
48. Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems ​
Author: Abdalla Doleh, Ratna Babu Chinnam
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00680v1 Announce Type: new Abstract: Evaluating artificial intelligence systems has historically relied on outcome-based benchmarks that measure task accuracy, robustness, or fairness. While indispensable, these benchmarks provide limited diagnostic insight into the underlying cognitive p...
49. When Does LLM Orchestration Pay Off? A Controlled Evaluation of Accuracy, Cost, and Task Difficulty ​
Author: Nicolas Leins, Nico Pelleriti, Jana Gonnermann-M"uller, Sebastian Pokutta
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00685v1 Announce Type: new Abstract: LLM orchestration is often assumed to improve reasoning by allocating additional inference-time computation, yet its gains may not justify its cost. Existing comparisons also frequently overlook differences in optimization effort, making it difficult t...
50. Evolutionary Curriculum Learning Improves Biological Sequence Modeling ​
Author: Richard Zhu, Kento Nishi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, q-bio.BM, stat.ML
arXiv:2608.00697v1 Announce Type: new Abstract: Variational autoencoders (VAEs) trained on multiple sequence alignments (MSAs) have emerged as powerful generative models for biological sequences, with applications ranging from disease variant prediction to functional RNA design. However, standard bi...
51. DGA$_2$D: Directed Graph-Guided Automated Algorithm Design with Large Language Models ​
Author: Jiale Zhao, Zimu Chen, Sirui Mao, Wentao Yang, Yuxiang Bai, Liyuanjun Lai
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2608.00700v1 Announce Type: new Abstract: The rapid development of Large Language Models (LLMs) has opened new avenues for Automated Heuristic Design (AHD) for solving NP-hard combinatorial optimization problems (COPs). However, existing LLM-driven AHD methods are largely confined to rigid sol...
52. Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations ​
Author: Xinshun Feng, Ziqi Miao, Lijun Li, Jing Shao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00711v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed in scientific research, where reliability is critical and the underlying knowledge is densely interconnected. In such settings, hallucinations are particularly damaging: a single erroneous cla...
53. AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment ​
Author: Garv Vikram Gursahaney, Baskhad Idrisov, Thorsten Fr"ohlich, Tim Schlippe
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.00717v1 Announce Type: new Abstract: Rubric-based AI systems for thesis assessment use criterion weights to assign different levels of importance to evaluation criteria. These weights are typically defined through expert judgment, although little empirical evidence exists regarding how th...
54. Behavioral Grammar: Detecting Adaptive Malware via Tiny Language Model Priors and Second-Order Temporal Analysis ​
Author: Zihan Luo
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00745v1 Announce Type: new Abstract: Modern endpoint detection systems face a fundamental tension: signature-based approaches are trivially evaded by polymorphic or adaptive threats, while heavy deep-learning models resist auditability and deployment at scale. This paper presents Behavior...
55. FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction ​
Author: Chaoqun Yang, Fengbin Zhu, Xinyu Lin, Long Bai, Xiaoluan Liu, Ke-Wei Huang, Roger Zimmermann, Tat-Seng Chua
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00764v1 Announce Type: new Abstract: Financial indicators are essential tools for transforming raw financial data into interpretable measures for various downstream tasks, such as valuation, risk assessment, and economic analysis. However, existing financial benchmarks largely focus on an...
56. Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation ​
Author: William Caban
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00794v2 Announce Type: new Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims. No formal framework has yet characterized how validity degrades across the stages of these pipelines. W...
57. AgentSLABench: Evaluating and Benchmarking Agentic Systems Under Resource Constraints ​
Author: Meher Bhaskar Madiraju, Meher Sai Preetam Madiraju
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00805v1 Announce Type: new Abstract: We present AgentSLABench, a resource-aware evaluation framework for autonomous AI agents that measures correctness alongside latency, cost, compute, memory, and network usage under declared resource budgets. Unlike standard benchmarks that report only ...
58. Large language models improve physician accuracy but lead to false reliance ​
Author: Tirtha Chanda, Christoph Wies, Franziska Schramm, Carina Nogueira Garcia, Nicolas B. Merl, Martin J. Hetz, Jochen S. Utikal, Phillip Tschandl, Cristian Navarrete-Dechent, Alexander Thiem, Jakob N. Kather, Consortium, Titus J. Brinker
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, stat.AP
arXiv:2608.00817v1 Announce Type: new Abstract: Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to inv...
59. The Scaling Paradox in Human-AI Collaboration ​
Author: Anyan Qi, Mengxin Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, econ.GN, q-fin.EC
arXiv:2608.00818v1 Announce Type: new Abstract: The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems scale, their capabilities tend to improve predictably. Yet, in real-world applications, AI rarely operates in isola...
60. Isotropy Cliffs: The Geometric Signature of Decision-Making in Large Language Models ​
Author: Okan S. Coskun, Florian Rottach, Carsten Eickhoff, William Rudman
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00828v1 Announce Type: new Abstract: We investigate the geometry of decision-making in Multiple Choice Question Answering (MCQA) through the lens of isotropy. Analyzing five open-weight models across diverse datasets, we identify decision-critical transition layers characterized by a shif...
61. Similarity Weighted Aggregation with Global Differential Privacy for Federated Brain Lesion Segmentation ​
Author: Muhammad Irfan Khan, Eero Lehtonen, Joni Obradovic, Elina Kontio, Esa Alhoniemi, Suleiman A. Khan, Mojtaba Jafaritadi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.CV
arXiv:2608.00872v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative training of machine learning models across multiple institutions without sharing sensitive data, making it particularly suitable for medical imaging applications. However, heterogeneous data distributions a...
62. Assuming You Knew: Fixing an Epistemic Semantics for Flow Policies Using Agentic AI ​
Author: David A. Naumann
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.LO, cs.PL
arXiv:2608.00882v1 Announce Type: new Abstract: Many high-level security requirements are about the allowed flow of information in programs and are difficult to make precise because they involve selective downgrading. Notions from epistemic logic have emerged as a good approach to policy semantics b...
63. Neuro-Evolved Heuristics for Variable Gapped Common Subsequence Identification ​
Author: Marko Djukanovi'c, Christian Blum, Aleksandar Kartelj, Saso Dzeroski, Ziga Zebec
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00888v1 Announce Type: new Abstract: This study addresses the Variable Gapped Longest Common Subsequence Problem (VGLCSP), a variant of the classical longest common subsequence problem with additional gap constraints and applications in sequence alignment and time-series analysis. While t...
64. CADIR: A Cross-Backend Editable Intermediate Representation for Agentic CAD Generation ​
Author: Yu Liu, Jingzhe Ni, Yiming Chen, Junqi Huang, Ruofeng Tong, Min Tang, Peng Du
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00891v1 Announce Type: new Abstract: Large language models have made it possible to generate executable computer-aided design (CAD) programs from natural-language descriptions or images. However, existing methods represent modeling processes as backend-specific sequential scripts with imp...
65. Modeling Social Dynamics with an LLM-Enabled Agent Based Network-Dynamic (LAND) Model ​
Author: Lynnette Hui Xian Ng, Kathleen M. Carley
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00929v1 Announce Type: new Abstract: Social dynamics encode the process in which individual network and discourse interactions aggregate into collective influence, narrative dominance and coordinate behavior. This paper uses the the GhostField architecture, a hybrid LLM-Enabled Agent Base...
66. PMMC: Prospective Multimodal Memory Compilation for Long-Term LVLM Agents ​
Author: Jingyu Sun, Yan Lin, Yuyang Xue, Yifan Wang, Zhengtao Yao, Rui Qian, Zefeng Xu, Jiachen Li, Xianyang Liu, Jiancheng Pan, Jingyuan Sun, Syed Murtuza Baker, Hongpeng Zhou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00962v1 Announce Type: new Abstract: Long-term memory is essential for LVLM agents to maintain consistency and integrate information across extended multimodal interactions. Existing agent memory systems, however, often reduce visual experiences into textual summaries or rely on static re...
67. TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents ​
Author: Jingyu Sun, Yuyang Xue, Mingyang Li, Zhengtao Yao, Jiachen Li, Yang Cui, Wenhao Cai, Haozhe Liu, Fangying Wang, Magdalene Katharina Montgomery, Syed Murtuza Baker, Hongpeng Zhou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00967v1 Announce Type: new Abstract: Large language model agents have shown strong capabilities in generating coherent and contextually appropriate responses, yet robust long-horizon dialogue remains limited by the lack of external memory that is traceable, updatable, and diagnostically t...
68. PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent ​
Author: Sudipta Paul, Vijay Srinivasan, Vivek Kulkarni, Aounon Kumar, Yashas Malur Saidutta, Wenbo Li, Srinivas Chappidi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00969v1 Announce Type: new Abstract: Existing search-augmented LLM agents are trained using Reinforcement Learning to boost its reasoning capabilities. However, these approaches primarily rely on outcome-level rewards, which provide little supervision over search behavior and overlook age...
69. Search-GRT: Guided Retrieval Training of Search Agents to Optimize for Complex Question Answering ​
Author: Aounon Kumar, Sudipta Paul, Vivek Kulkarni, Vijay Srinivasan, Srinivas Chappidi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00974v1 Announce Type: new Abstract: The effective use of search engines by large language models (LLMs) remains a significant challenge, particularly in complex, multi-hop question-answering (MHQA) tasks. These tasks require the model to decompose questions into subqueries, retrieve rele...
70. Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel ​
Author: Yohei Nakajima
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.GT
arXiv:2608.00979v1 Announce Type: new Abstract: Large language models are increasingly used as synthetic research participants and are often validated by whether their marginal responses resemble human data. We study a fixed panel of sixteen lightweight persona-conditioned GPT-4.1 configurations in ...
71. Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable ​
Author: Wenhui Chen, Jianlin Chen, Ziyao Lin, Chi Man Vong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.00981v1 Announce Type: new Abstract: When a self-improving AI-for-science system claims a new capability, the evidence is usually a benchmark delta, a description-length gate, or a p-value. None separates a real gain from extra search, from a changed verifier, or from adaptation to a fall...
72. SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling ​
Author: Shrenil Shaun Sharma, Avi Sharma
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.00991v1 Announce Type: new Abstract: This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form variation. Grounded in canonical scheduling instances and solver-derived feasibility and optimality, SCHED...
73. Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets ​
Author: Wenhui Chen, Jianlin Chen, Ziyao Lin, Peiji Long, Chi Man Vong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01000v1 Announce Type: new Abstract: Language models are increasingly promoted from examinees to examiners: they write the test suites, answer keys, rubrics, and reward functions that define correctness for other systems. We measure the capability that role assumes and find it lacking und...
74. From AI Technical Debt to Agentic Technical Debt: A Systematic Mapping of Root Causes and Manifestations in Agentic AI Systems ​
Author: Muhammad Tukur, Hayatullahi B. Adeyemo, Tao Chen, Nour Ali, Anis Zarrad, Marco Agus, Rick Kazman, Rami Bahsoon
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.SE
arXiv:2608.01001v1 Announce Type: new Abstract: The emergence of Agentic AI systems, characterized by autonomous reasoning, multi-agent collaboration, tool orchestration, adaptive decision-making, and persistent memory, represents a fundamental shift from traditional AI pipelines to dynamic software...
75. Toward Fine-Grained Forgetting:Attribute Unlearning for Multimodal Large Language Models ​
Author: Junkai Lin, Junkai Chen, Siqi Hou, Yuhao He, Ruiqi Liu, Chenhan Jin, Shengze Xu, Tieyong Zeng
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01008v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) exhibit strong vision--language capabilities but may also memorize and disclose sensitive information. Machine unlearning seeks to remove designated knowledge without retraining from scratch while preserving gen...
76. FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds ​
Author: Kapil Wanaskar, Gaytri Jena, Aman Chadha, Vinija Jain, Vasu Sharma, Amitava Das
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2608.01049v1 Announce Type: new Abstract: World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly compelling dir...
77. Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale ​
Author: Ortal Ashkenazi, Vitalii Kloz, Mykhailo Ulianchenko
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SE
arXiv:2608.01050v1 Announce Type: new Abstract: Production LLM agents that select from large skill libraries face a limitation that semantic relevance alone cannot resolve: a skill may match a user's topic yet be impossible to execute in the current account state. We present a deployed three-stage s...
78. Control Under Compression: Reliability Frontiers for Tool-Using Agents ​
Author: Yinghan Hou, Zongyou Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.01056v1 Announce Type: new Abstract: Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments, policies, execution protocols, and recovery. Compressing these agent control contexts (ACCs) can reduc...
79. Role-Decoupled Attention Residuals: Separating Matching and Content Retrieval Across Depth ​
Author: Kehan Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01075v1 Announce Type: new Abstract: Depth-routing residual architectures allow Transformer layers to retrieve earlier representations instead of inheriting only the immediately preceding state. Existing Block Attention Residuals, however, use a single content-dependent depth mixture to c...
80. Inter-Residue Geometry Attention for Antibody-Specific Epitope Prediction ​
Author: Chuanliu Fan, Nan Yu, Junjie Wu, Guohong Fu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01092v1 Announce Type: new Abstract: Antibody-specific epitope prediction aims to identify which antigen residues are recognized by a given antibody, a task that depends on the three-dimensional complementarity between antibody CDRs and the antigen surface. Existing methods usually levera...
81. Fighting Fire with Fire: On the Feasibility of Protecting Exercises Against AI Cheating ​
Author: Tobias Braun, Jonas Grebe, Louis Rethfeld, Marcus Rohrbach
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01112v1 Announce Type: new Abstract: The widespread adoption of generative AI enables students to outsource cognitive effort to increasingly capable assistants, creating an illusion of competence while undermining the independent reasoning that education aims to cultivate. We investigate ...
82. MA-HEAD-Net: Adaptive Rule-Guided Multi-Agent DRL for AoI Minimization in UAV-Assisted Emergency Networks ​
Author: Yixin Zhang, Zhuohui Yao, Wenchi Cheng, Walid Saad
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.01128v1 Announce Type: new Abstract: In post-disaster scenarios, unmanned aerial vehicles (UAVs) are critical for establishing emergency communication networks. For time-critical rescue missions, information freshness is crucial because decisions based on outdated data may lead to ineffec...
83. PATH-Bench: Path-Dependent Evaluation of Lifelong Agents ​
Author: Xidong Yang, Xingyi Zhang, Wenhao Li, Wenyan Liu, Junjie Sheng, Yun Hua, Wei Yin, Tao Fang, Chuyun Shen, Xiangfeng Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01149v1 Announce Type: new Abstract: Lifelong LLM agents increasingly adapt through external learning states that store past interactions as retrievable memories or reusable skills, yet existing benchmarks rarely account for how the path of accumulated experience shapes what agents transf...
84. The Graph Language: How Knowledge Graphs Speak to Large Language Models ​
Author: Giuseppe Pirr`o
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01175v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at reasoning but benefit from grounding provided by Knowledge Graphs (KGs). However, integrating these paradigms is challenging. We introduce GRALAN, which enables KGs to speak directly in the LLM's semantic space thr...
85. Co-evolution of social reward and punishment under institutional interventions ​
Author: Van An Nguyen, Vuong Khang Huynh, Hoai Thuong Nguyen, Duc Tin Duong, An Nguyen Gia, Tat Kien Nguyen, Huu Loi Bui, My Nguyen Tra, Ho Nam Duong, Ba Thanh Phan, Thanh Vo, Dinh Anh Trung Hoang, Adeela Bashir, Zhao Song, Manh Hong Duong, Le Hong Trang, The Anh Han
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, cs.MA, math.DS, nlin.AO
arXiv:2608.01183v1 Announce Type: new Abstract: We investigate how peer and institutional incentives jointly shape the evolution of cooperation, social welfare, and enforcement efficiency in social dilemmas. In a Prisoners Dilemma with four strategies, unconditional cooperators (C), defectors (D), s...
86. Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races ​
Author: Phu Hoa Pham, Duy Minh Dao Sy, Trung Kiet Huynh, Phu Quy Nguyen Lam, Chi Nguyen Tran, Minh Trung Le, Phong Hao Le, Dinh Nam Nguyen, Thien Ky Nguyen Dong, Elias Fernandez Domingos, Le Hong Trang, The Anh Han
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.GT, cs.LG, cs.MA
arXiv:2608.01193v1 Announce Type: new Abstract: An AI development race creates a multi-agent safety dilemma. Each company can develop slowly and safely, or move faster while taking a risk that may remove its final reward. We use this repeated game to study strategic safety behaviour among large lang...
87. Reputation-driven Cooperation in Lattice-based Decentralized Federated Learning through Evolutionary Game Theory ​
Author: Phuc Hoang Truong Huynh, Dung Tran Vinh, Khoa Duc Anh Lam, An Nghiem Nguyen Truong, Uyen Nha Tran Bui, Khang Nguyen Dinh, Bao Nguyen Le Gia, Minh Le Nguyen Nhat, Manh Hong Duong, The Anh Han, Thi Ai Thao Nguyen, and Le Hong Trang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.DC, cs.GT, math.DS
arXiv:2608.01197v1 Announce Type: new Abstract: Decentralized Federated Learning (DFL) has emerged as an optimal privacy-preserving solution; however, it remains vulnerable to opportunistic behaviors due to the absence of a central coordinator. While Evolutionary Game Theory (EGT) serves as a powerf...
88. Perspectives on Tsallis Statistics for Artificial Intelligence ​
Author: Kleyton da Costa, Bernardo Modenesi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CE
arXiv:2608.01223v1 Announce Type: new Abstract: Tsallis statistics generalizes Boltzmann-Gibbs statistical mechanics through a single real parameter $q$ that controls the weight assigned to rare and frequent events. Originally proposed to describe physical systems with long-range correlations, multi...
89. CT-PrepAgent: Bounded Policy and Controlled Execution for Adaptive CT Data Preparation ​
Author: Xiaolin Fan, Yue Pei, Yingying Zhang, Haogang Zhu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01233v1 Announce Type: new Abstract: Heterogeneous computed tomography (CT) acquisitions and diverse downstream task requirements limit the transferability of fixed data preparation workflows across data sources and tasks. Existing approaches typically rely on manually designed or dataset...
90. Learning What to Remember and What to Internalize in LLM Self-Evolution via Adaptive Memory-Parameter Coordination ​
Author: Tianyun Ji, Zhenya Huang, Jiayu Liu, Zirui Liu, Yu Su, Hongbin Pei
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01234v1 Announce Type: new Abstract: Large language model agents increasingly operate in dynamic environments where tool interfaces, APIs, and user requirements change after deployment. Existing self-evolution methods mainly follow two paradigms: harness-based approaches, which externaliz...
91. Cognitive Demand Steering for Adaptive Meta-Reasoning in Large Language Models ​
Author: John Scoville, Shengzhuang Chen, Yejin Bang, Stefan Winzeck, Jonathan Richard Schwarz
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01319v1 Announce Type: new Abstract: Recent meta-reasoning frameworks improve LLM reasoning by wrapping chain-of-thought generation in an iterative control loop, allowing more effective backtracking, termination of reasoning loops, and injection of promising reasoning patterns, among othe...
92. G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution ​
Author: Shaoxiong Yang, Mengyuan Zhang, Shaojun Lin, Chao Li, Wei Liu, Kun Shao, Jian Luan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01324v1 Announce Type: new Abstract: Deep search has become a fundamental capability of large language models (LLMs) for solving open-domain complex tasks. However, existing approaches typically rely on linear sequential reasoning for both trajectory generation and inference, making it di...
93. 402Pilot: An x402 Decision Layer for Autonomous Agent Micropayments ​
Author: Yin Li, Yanbo He, Boo-Ho Yang, Rav Lawana, Ziyue Li, Wei Zeng, Jing Tang, Fugee Tsung
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01341v1 Announce Type: new Abstract: Programmable-payment protocols such as x402 enable per-request micropayments, but they do not determine which payable service an autonomous agent should buy under a finite wallet. We formulate this buyer-side problem as agent-native payment decision-ma...
94. Agentic Stage-One Stellarator Optimization: Autonomous Multi-Objective Search for Finite-Beta Equilibria ​
Author: Tingjia Zhang, Hongke Lu, Zhuoran Meng, Runlai Xu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01344v1 Announce Type: new Abstract: Stage-one stellarator design searches a high-dimensional family of three-dimensional plasma boundaries and fixed-boundary MHD equilibria for configurations that jointly meet requirements on confinement, field-line topology, force balance, stability pro...
95. High-Stakes Decisions with Language Models: Insights from Emergency Triage ​
Author: Khurram Yamin, Christopher Kelly, Bryan Wilder, Eric Horvitz
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01361v1 Announce Type: new Abstract: High-stakes decisions under uncertainty, such as medical emergency triage, require more than accurate predictions. They depend on estimating the likelihood of alternative outcomes while explicitly weighing the consequences of different actions, princip...
96. CRAFTS: Collaborative Role-Adaptive Fine-Tuning of LLM Agents for Chemical Process Simulation ​
Author: Ziyun Zhang, Yuxin Lin, Eldin Wee Chuan Lim, Xinghao Ding
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.01369v1 Announce Type: new Abstract: Constructing an executable chemical-process model remains manually intensive. Chemical engineers translate underspecified requests into coupled decisions about unit operations, thermodynamics, streams, specifications, degrees of freedom (DoF), initiali...
97. CraftAlign: Feature-Grounded Evaluation and Revision Guidance for AI Stories ​
Author: Yang Yang, Boyun Xu, Shaofeng Liang, Yun Han, Zining Zhong, Songning Lai, Kaishen Yuan, Yutao Yue
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01377v1 Announce Type: new Abstract: Large language models can now generate fluent and complete stories, yet many outputs still feel formulaic and unnatural because of cliches, over-explanation, linear causal progression, and stereotyped endings, an immediately recognizable AI flavor. Exi...
98. KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval ​
Author: Yongbin Choi, Gyuho Shim, Youngjoon Jang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01389v1 Announce Type: new Abstract: Visual Document Retrieval (VDR) directly matches text queries against document images, preserving visual and structural information that may be lost during text extraction. However, existing VDR models and training resources remain predominantly Englis...
99. No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks ​
Author: Simiao Xie, Chuancheng Shi, Shangze Li, Wenhua Wu, Fei Shen, Ying Zhou, Zhiyong Wang, Tat-Seng Chua
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2608.01414v1 Announce Type: new Abstract: With the rapid release of open-weight large foundation models, safety threats are shifting from black-box jailbreaks to neuron-level white-box attacks that directly identify and manipulate safety-related neurons. Existing alignment methods often invest...
100. Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning ​
Author: Wenhao Zhang, Yibo Xie, Rui Wang, Jiahua Yang, Lei Jiang, Zibo Yang, Yawei Wang, Jiali Xu, jasperawang, Haoyang Long, Huan Xiong, alantzhao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.01418v1 Announce Type: new Abstract: Autoregressive rollout generation is a major computational cost in reinforcement learning for large language models. Reusing each rollout batch for additional learner updates amortizes this cost, but later updates become increasingly off-policy as the ...
101. Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics ​
Author: Shengwei Xu, Yuxuan Lu, Yifan Wu, Jason Hartline, Grant Schoenebeck
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, cs.LG
arXiv:2608.01423v1 Announce Type: new Abstract: Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response. The reliability of an evaluation metric is usually judged by its statist...
102. MRAFnd: Multimodal Retrieval-Augmented Framework for Zero-Shot Fake News Detection ​
Author: Lehan Zhang, Yinlei Cheng, Shiqi Hu Yiheng Zhou, Shangxi Li, Naidong Zhao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01430v1 Announce Type: new Abstract: The rapid dissemination of multimodal content has intensified the spread of fabricated news, presenting a substantial threat to social integrity. A formidable challenge for current detection systems is identifying misinformation related to novel events...
103. PolymerGPT: Multi-property Optimization with a Decoder-Based GPT Model for Generative Polymer Design ​
Author: Charlie Pyle, Adarsh Gadari, C. Adrian Figg, Zhenquan Jia, Yaohang Li, Chunjiang Zhu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci, cs.CE, cs.LG
arXiv:2608.01431v1 Announce Type: new Abstract: Polymer property prediction and inverse generative design targeting desired properties are two crucial tasks in machine learning-assisted polymer design. While the former has received considerable attention, there have been limited methods developed fo...
104. A New Theory of Value for Post-AGI Economics ​
Author: Keyun Ruan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, econ.GN, q-fin.EC
arXiv:2608.01432v2 Announce Type: new Abstract: Artificial general intelligence (AGI) may weaken scarcities in labour, expertise, information, and productive capability that underpin established theories of economic value. If cognitive work becomes widely automatable, market price, labour input, rev...
105. Beyond Routing Saturation: A Long-Horizon Class-Incremental Perspective on Expert Routing in Multimodal Continual Instruction Tuning ​
Author: Huiyu Yi, Yongqi Xu, Bogang Zhang, Dunwei Tu, Xu Zhiming, Zhen-Hao Xie, Baile Xu, Furao Shen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01437v1 Announce Type: new Abstract: Multimodal Continual Instruction Tuning (MCIT) enables multimodal large language models to acquire new tasks sequentially while retaining previously learned capabilities. Many recent methods maintain task-specific LoRA experts and route each input to o...
106. Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI ​
Author: Quang Bui, Shlok Jaiswal, Samuel Paik-Heintz, Kevin Zhou, Kaushik Madapati, Krittaphas Chaisutyakorn, Noah Dane Hebdon, Dimitrios Proios, Sebasti'an Andr'es Cajas Ord'o~nez, Kacper Dobek, Boya Zhang, Aly Dhedhi, Ahram Han, Kushul Reddy Palakala, Rahul Gorijavolu, Jacques Kpodonu, Leo Anthony Celi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01462v1 Announce Type: new Abstract: Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond the size of the accuracy loss: wh...
107. Where Reasoning Diverges: Localized Multi-Agent Debate for Multi-Hop Question Answering ​
Author: Weijun Gao, Xiang Ding, Haoyang Liu, Tiancheng Xing
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.01463v2 Announce Type: new Abstract: Multi-agent debate commonly exchanges complete rationales even when disagreements concern only a few intermediate claims. We introduce Localized Multi-Agent Debate (LMAD), an inference-time protocol that represents agent rationales as nodes, locates th...
108. Computing with Agentic Oracles ​
Author: Jie Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01464v1 Announce Type: new Abstract: This paper extends the stochastic-oracle model of AI-augmented computing to include agentic oracles. Unlike a stationary stochastic oracle, which responds to the same query according to a fixed response distribution across calls, an agentic oracle can ...
109. Sweet Little Lies: Strategic Deception in AI Emotional Support Chatbots ​
Author: Aseem Pahuja, Zhiling Guo, Tahir Abbas Syed
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, econ.TH
arXiv:2608.01480v1 Announce Type: new Abstract: The paper examines the strategic behavior of Gen AI chatbots used for emotional support. Using a Bayesian Persuasion, we model interactions between chatbots that send signals about users' emotional states and users who decide whether to engage based on...
110. MineGrad: Gradient Inversion Attacks on LoRA Fine-Tuning ​
Author: Hasin Us Sami, Swapneel Sen, Basak Guler
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01521v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT), such as low-rank adaptation (LoRA), has recently been adopted in federated learning to reduce communication and computation costs. In this setup, users download a pretrained model from the server prior to fine-tu...
111. V-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic Memory ​
Author: Dingyi Kang, Dongming Jiang, Yi Li, Guanpeng Li, Bingzhe Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.IR
arXiv:2608.01543v1 Announce Type: new Abstract: Interaction between users and LLM agents is increasingly multimodal: conversations interleave text with images, and a later question may target either. Yet most agent memories are designed around text, and even the few that support multimodal conversat...
112. Emergence Invariance: From Symbolized Thought to Interface Refinement ​
Author: Yi Liu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.01548v1 Announce Type: new Abstract: Language can be viewed as a formalized subset of thought: a consequence-governed symbolic structure projected from wider situated cognition. Large language models trained at scale exhibit compensatory emergence: sparse architectural primitives support ...
113. Securing Agentic AI: From Per-Action Checks to Trajectory Assurance ​
Author: Alireza Lotfi, Subangkar Karmaker Shanto, Imtiaz Karim, Elisa Bertino
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.MA
arXiv:2608.01558v1 Announce Type: new Abstract: Autonomous agents are increasingly used to execute consequential tasks in environments governed by operational constraints, organizational policies, regulatory requirements, and technical standards. Their safety is therefore determined not by the corre...
114. Does the Competitive Component of Adversarial Self-Play Improve Legal Reasoning? A Controlled Negative Result ​
Author: Miseog Shawn Kim
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.01559v1 Announce Type: new Abstract: Adversarial self-play is an appealing recipe for legal reasoning: have a student model draft an argument, have an adversary attack it, and reward the student when its argument survives the attack. We designed exactly such a training signal -- a verifia...
115. Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled Reasoning ​
Author: Xuyang Zhao, Liting Zhang, Zichen Xu, Zhihu Wang, Xu Caiyue, Shiwan Zhao, Qicheng Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01589v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves reasoning by using a privileged view of a model conditioned on reference solutions to supervise a student view that observes only the question. However, the teacher-provided token-level targets may depend on ...
116. Latent Thought Credit: Multi-Answer Credit Assignment for Latent Reasoning ​
Author: Xuyang Zhao, Liting Zhang, Zichen Xu, Yong Chen, Wenjia Zeng, Shiwan Zhao, Qicheng Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01593v1 Announce Type: new Abstract: Latent reasoning allows language models to carry out intermediate reasoning in continuous latent representations rather than fully externalizing it as discrete chains of thought. However, assigning credit to such latent thoughts from answer-only reward...
117. Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer ​
Author: Logan Ritchie, Sushant Mehta, Liudas Panavas, Edwin Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2608.01604v1 Announce Type: new Abstract: Long-horizon tasks require agents to maintain coherent state and goals across nested and branching work. We call this capability goal-directed execution (GDE): the repeated application of four behaviors, namely selecting goals, constructing task-releva...
118. When Memory Updates but Behavior Does Not: Repairing Implicit Stale Dependencies in Personalized Agent Responses ​
Author: Haofei Sun, Lin He
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01619v1 Announce Type: new Abstract: Memory-augmented agents can know that a user's stored state is outdated and still plan around the old value. The STALE benchmark calls this the implicit policy adaptation (IPA) gap. We identify one structural contributor: draft-anchored verification ch...
119. Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw ​
Author: Zheng Lin, Yuzhe Huang, Zhenxing Niu, Xianmin Ye, Haichang Gao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01637v1 Announce Type: new Abstract: Long-term memory enables LLM agents to retain useful information across sessions, but also creates an attack surface through which adversaries may poison an agent's persistent memory to steer its behavior. Existing memory poisoning attacks mainly rely ...
120. GISAgentBench: A Practitioner-Sourced Benchmark for Evaluating LLM Agents on GIS Tasks ​
Author: Abhinav Pothuri, Zhe Jiang, Zelin Xu, Di Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01645v1 Announce Type: new Abstract: Geographic Information System (GIS) professionals rely on multi-step spatial analysis workflows to support decision-making in urban planning, disaster response, and environmental monitoring. The process is tedious, time-consuming, and error-prone. Whil...
121. LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing ​
Author: Wen Zan, Jiaqi Zhang, Jianchao Tan, Hong Liu, Cunguang Wang, Xiang Li, Duyue Ma, Guanyu Wu, Yifan Lu, Fengcun Li, Yerui Sun, Peng Pei, Yuchen Xie, Xunliang Cai
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.DC, cs.LG
arXiv:2608.01662v2 Announce Type: new Abstract: DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive $O(L^2)$ scoring overhead and the hardware-inefficient, discontinuous me...
122. Allocation Before Ranking: Decoupled Token Compression for OmniLLMs ​
Author: Zhenghui Guo, Yilin Yang, Yuanbin Man, Miao Yin, Weidong Shi, Rabimba Karanjai, Omprakash Gnawali, Chengming Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.SD
arXiv:2608.01665v1 Announce Type: new Abstract: Token compression in OmniLLMs is typically posed as a single saliency-ranking problem: score each multimodal token, keep the top-K. We argue this abstraction is mis-specified. The same attention score simultaneously decides two things: how much retaine...
123. TCPO: Turn-Level Credit Policy Optimization ​
Author: Sicong Liao, Zhi Chen, Yaohua Tang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01667v1 Announce Type: new Abstract: Verifier-guided reinforcement learning has become a powerful paradigm for improving LLM reasoning. In multi-turn settings, models receive a verifier score after each turn and iteratively refine their outputs. Although such scores provide dense feedback...
124. When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary ​
Author: Qiuyang Zhan, Rui Zhang, Sheng Guo, Lepeng Zhao, Zhuotao Liu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01679v2 Announce Type: new Abstract: Persistent memory allows (self-evolving) LLM agents to adapt across tasks by consolidating heterogeneous interaction histories into reusable facts, preferences, observations, and rules. Yet consolidation also imposes an implicit authorization boundary:...
125. GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks ​
Author: Jiarui Tan, Zhongjian Zhang, YaBo Guo, Jiawei Liu, Yujie Xing, Muhan Zhang, Cheng Yang, Chuan Shi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01684v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly capable of planning, using tools, and interacting with external environments. They are typically supported by harnesses, which manage state and coordinate multi-step execution. Graph analysis provides ...
126. Beyond Single-Use Tokens: Durable Authorization State for Replay-Resistant LLM Agent Actions ​
Author: Jinghan Xu, Longze Fan, Zeyuan Wang, Xinjin Li, Hankai Liu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01710v1 Announce Type: new Abstract: Tool-using large language model agents frequently replan, retry failed operations, delegate tasks, and resume after crashes. These behaviors can cause one user authorization to be requested and executed multiple times under freshly issued token identif...
127. Constructing Executable Analytical Knowledge Representations for Meta-Analysis Synthesis Using an Agentic Harness ​
Author: Lingbo Li, Anuradha Mathrani, Teo Susnjak
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01711v1 Announce Type: new Abstract: Meta-analysis synthesis highlights a fundamental challenge in knowledge-based scientific analysis: structured evidence does not by itself represent the analytical knowledge required for executable computation. Decisions about evidence assignment, analy...
128. LaCache: Robust Semantic Caching for LLM Serving ​
Author: Jiacheng Liang, Yuhui Wang, Tanqiu Jiang, Ting Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01718v1 Announce Type: new Abstract: Semantic caching, which reuses responses to semantically similar requests via their embeddings, has seen growing adoption in LLM serving, offering faster responses and reduced costs. Yet existing schemes are fundamentally vulnerable to cache-collision ...
129. DAPD: Dual-Anchored Policy Distillation ​
Author: Jianyu Wu, Yizhou Wang, Encheng Su, Chen Tang, Shixiang Tang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01735v1 Announce Type: new Abstract: On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged information but can induce a privilege illusion: the student learns privilege-dependent behavior it cannot reprodu...
130. CoEvo-Mem: Co-Evolving Retrieval Policy and Memory Bank for LLM Agents ​
Author: Bowen Ye, Yongchao Xu, Zhijian Li, Xiang Yin, Junkai Ma, Wenzhao Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01739v1 Announce Type: new Abstract: As memories accumulate across tasks and sessions, the performance of long-term LLM agents depends jointly on query-specific retrieval and continual memory refinement. However, existing methods typically optimize either memory access, through iterative ...
131. MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents ​
Author: YuFei Luo, Xiucheng Xu, Zhen Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.01742v1 Announce Type: new Abstract: Long-term memory is critical for LLM agents operating over long-horizon interactions. However, several persistent limitations of existing memory systems can be traced to two recurring misalignment patterns in long-term interaction settings: Temporal-St...
132. RL-Lock: Reinforcement Learning for Generating Interlocking Assemblies ​
Author: Xuyang Ma, Chaewoon Kim, Haonan Zhang, Rulin Chen, Ziqi Wang, Peng Song
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.GR
arXiv:2608.01744v1 Announce Type: new Abstract: An interlocking assembly is an assembly in which component parts are connected purely through their geometric arrangement, without relying on external connectors such as glue and nails. Such assemblies have been widely used in a variety of real-world a...
133. Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs ​
Author: Zixuan Huang, Yang Zhou, Kaixuan Wang, Guli Zhang, Hongyan Xie, Yakun Zhu, Hao Geng, Xiaozhi Chen, Yikun Ban, Deqing Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01755v2 Announce Type: new Abstract: Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines comm...
134. Leveraging AI for fine-grained food safety risk forecasting in sparse data conditions ​
Author: Dongqi Wang, Weiwei Chen, Han Zhou, Weihua Zhou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01767v1 Announce Type: new Abstract: Ensuring food safety represents a critical public health challenge, particularly when inspection resources are limited and regional sampling data are sparse. This study proposes a Transformer-based framework capable of forecasting fine-grained, city-le...
135. FRAMES: Guarded and Dual-Objective Skill Evolution for Agents in Policy-Governed Enterprise Workflows ​
Author: Xuhui Wang, Ruoqi Shu, Chen Dan, Tianhua Xu, Mengxi Luo, Yanming Mai, Bo Wan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01772v1 Announce Type: new Abstract: LLM agents increasingly run policy-bound enterprise workflows such as document auditing, where they must apply rules consistently, ground every value, and stay auditable. Improving these agents is hard: operational feedback is sparse and unlabeled, edi...
136. REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models ​
Author: Xiang Xia, Cheng Yan, Yiming Zhang, Jiazheng Liu, Hongyu Zhang, Wuyang Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.01784v1 Announce Type: new Abstract: Mixture-of-experts (MoE) models increase parameter capacity by activating only a small subset of experts for each token. This conditional-computation paradigm has enabled autoregressive language models to scale model capacity without a proportional inc...
137. Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction ​
Author: Priyashree Roy, Sujitha Martin, Mohammad Rostami, Spencer Romo, Renhao Xue, Bob Strahan, Diego A. Socolinsky, Boyi Xie, Md Mofijul Islam
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.01792v1 Announce Type: new Abstract: Intelligent document processing (IDP) with vision-language models (VLMs) hinges on confidence scores trustworthy enough to route extractions between automation and human review. Existing document benchmarks are dominated by clean, high-quality samples,...
138. CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning ​
Author: Junru Song, Wenhao Zhang, Yang Yang, Xuekai Qiu, Feifei Wang, Weien Zhou, Tingsong Jiang, Ying Wen, Yang Li, Wen Yao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.01802v1 Announce Type: new Abstract: Target-oriented vision-and-language navigation (VLN) on aerial platforms is attracting growing attention for missions such as disaster rescue, infrastructure inspection, and security patrol. In this task, an unmanned aerial vehicle (UAV) needs to locat...
139. CockpitHAT: Dependency-Graph-Driven Hierarchical Attribution for Embodied Multi-Agent Cockpits ​
Author: Wei Wang, Shuanghe Liu, Zhu Zhuo, Jiaqi Zhong, Xiaozhao Zhao, Xiaojie Zuo, Jie Su
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01805v1 Announce Type: new Abstract: LLM multi-agent systems suffer from Correctness Collapse, where high task-level accuracy conceals severe process-level failures. This is especially hazardous in safety-critical embodied settings such as automotive cockpits, where lexically correct utte...
140. SearchMaster: Grounded and Regulated Self-Play for Search Agents ​
Author: Wentao Tan, Qiong Cao, Jiaqi Wang, Nan Duan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01822v1 Announce Type: new Abstract: Training LLM-based search agents requires high-quality search data: tasks that demand genuine multi-hop retrieval and trajectories that use search tools effectively. Existing pipelines often depend on human-written tasks, expert demonstrations, or stro...
141. Rewriting or Reweighting? A Geometric Account in Language Models ​
Author: Juntong Wang, Shengkun Yang, Xiyuan Wang, Muhan Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01835v1 Announce Type: new Abstract: Post-training can substantially alter language-model behavior, yet aggregate behavior rates do not reveal whether training removes an existing mechanism, creates a new one, or changes how an inherited mechanism is used. We study this question through t...
142. PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning ​
Author: Chunji Lv, Yangguang Wei, Junlin Liu, Yang Gao, Ming Liu, Xinming Wang, Jinyang Wu, Guoren Wang, Changsheng Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01837v1 Announce Type: new Abstract: Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-turn trajectory may receive only a single outcome-level signal. On-policy s...
143. FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling ​
Author: Xianglong Yan, Hong Liu, Chengzhu Bao, Tianao Zhang, Guanghua Yu, Jianchen Zhu, Yulun Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01847v1 Announce Type: new Abstract: Large language models (LLMs) achieve remarkable performance but are expensive to deploy due to their enormous size. FP4 quantization, with formats such as MXFP4 and NVFP4, offers an appealing solution with native hardware support on modern accelerators...
144. Exploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models ​
Author: Junxiang You, Junkai Chen, Yuhao He, Ruiqi Liu, Zhetao Guo, Shu Wu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01849v1 Announce Type: new Abstract: Machine unlearning offers a promising approach to remove unsafe content from Multimodal Large Language Models (MLLMs), yet ensuring the precision of unlearning remains a persistent challenge. One reason is that current MLLM unlearning evaluation paradi...
145. Physics-Informed Neural Networks for Complex Eigenfrequency Identification and Mode Structure Reconstruction of the Ground-State ITG Branch ​
Author: Dengdi Sun, Bingbing Zhang, Xiao Wang, Zikang Yan, Yuqiang Tao, Qingquan Yang, Guosheng Xu, Jin Tang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01850v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) combine sparse observations with physical equations, providing an important approach for modeling complex plasma processes and inferring unknown physical quantities. The steep-gradient pedestal of high-confineme...
146. EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning ​
Author: Dongwei Sun, Bowen Yao, Yujie Zhang, Pei Liu, Jing Yao, Xiangyong Cao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01856v1 Announce Type: new Abstract: Bi-temporal remote-sensing disaster change captioning often needs to identify sparse and spatially localized changes across large pre- and post-event scenes and then translate them into coherent, factual descriptions. However, existing change captionin...
147. Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge ​
Author: Xiaofeng Shi, Xiaosong Qiu, Wenxin Ma, Qian Kou, Yiming Pan, Longbin Yu, Ying Liu, Haiping Wang, Hua Zhou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01862v1 Announce Type: new Abstract: Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning w...
148. ReasonCast: Towards Explainable Time Series Forecasting with Reasoning ​
Author: Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.01875v1 Announce Type: new Abstract: Most time series (TS) models are specialized for a single task, either understanding (i.e., returning text answers about a TS) or generation (i.e., returning a numeric forecast). Only recently have unified models begun to handle the two within a single...
149. CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents ​
Author: Zhaoyang Li, Zenghuang Fu, Qiuyuan Ai, Ping Jiang, Haoyu Wu, Minghui Wu, Chenxu Zhao, Jie Song, Guannan He
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01904v1 Announce Type: new Abstract: Large language models can improve with reinforcement learning for search agents, yet existing self play agents repeatedly generate tasks while discarding the knowledge gained during successful searches. We introduce CoEvoKG, a framework that turns a kn...
150. Diagnosing Search Behavior and Failure Modes in Long-Horizon Search Agents ​
Author: Qi Liu, Jiaxin Mao, Fengbin Zhu, Tat-Seng Chua
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR
arXiv:2608.01913v1 Announce Type: new Abstract: Deep search agents answer difficult information-seeking questions by iteratively issuing search queries to gather supporting evidence, but it remains unclear whether and how greater search effort leads to better answers. We study these questions throug...
151. ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching ​
Author: Zihan Liu, Yuzhe Zhuang, Yuanzu Li, Wanshuang Gou, Jiahong Liu, Min Zhou, Menglin Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01926v1 Announce Type: new Abstract: JEPA-style visual world models offer an effective paradigm for visual goal planning by predicting future latent representations. Existing methods typically learn local transition consistency through next-step representation prediction. However, in long...
152. A Contractualist Argumentation Framework for Moral Decision-Making ​
Author: Luis Marcos-Vidal, Giulio Antonio Abbo, Tony Belpaeme
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.01937v1 Announce Type: new Abstract: Autonomous agents operating in shared environments must make decisions that affect multiple individuals with potentially conflicting interests. We propose a formal framework for moral decision-making grounded in Scanlon's contractualism, an ethical the...
153. Long-Horizon Autonomous Architecture Research with a Language-Model Agent: A Behavioural Case Study ​
Author: Aon Safdar, Mohamed Saadeldin
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.01995v1 Announce Type: new Abstract: We study what happens when a single general-purpose large language model acts as the sole researcher on a long-horizon neural architecture design problem. The agent receives a scientific question, an initial hypothesis and motivation, a compute budget,...
154. Evolving in the Agent Jungle via History-Informed Opponent Awareness ​
Author: Zhaofeng Zhang, Linhan Xia, Rui Liu, Yihao Wang, Binrui Shen, Shengxin Zhu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02005v1 Announce Type: new Abstract: Learning to adapt strategies through interaction is a key step toward more general and autonomous LLM agents. Existing approaches typically achieve behavioral adaptation by revising skill libraries. However, in multi-agent environments, opponents may s...
155. HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents ​
Author: Daeyoung Roh, Donghee Han
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02009v1 Announce Type: new Abstract: Retrieval-augmented search agents answer multi-hop questions by repeatedly issuing search queries and accumulating evidence. This creates a stopping problem: after the necessary evidence has appeared, further retrieval often adds cost, latency, and dis...
156. Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG ​
Author: Daeyoung Roh, Donghee Han
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02011v2 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them. We study this failure mode as a procedural property of the age...
157. EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers ​
Author: Junyeong Park, Jieun Han, Haneul Yoo, So-Yeon Ahn, Jinsung Yoon, Alice Oh
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.02024v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used across diverse tasks in K-12 education, yet existing safety evaluations rarely examine how harmful or inappropriate content appears in interactions between LLMs and students or teachers. To address thi...
158. HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning ​
Author: Runchuan Zhu, Hongbin Lai, Bowen Jiang, Junrui Zhang, Zhangheng LI, Ostap Kilbasovych, Junyuan Hong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02026v1 Announce Type: new Abstract: Reflection is a powerful mechanism for LLM reasoning, yet its effectiveness hinges on accurately attributing failures to specific reasoning steps, a capability that current models notably lack. Existing failure attribution methods either require expens...
159. Cross-Fitted Residual Utility for Primary-Preserving Cognitive Decision Correction in Automatic Modulation Classification ​
Author: Linzhuo Han, Zongyong Cui, Houbiao Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02063v1 Announce Type: new Abstract: Automatic modulation classification research has largely emphasized representation accuracy, but a cognitive receiver must also decide when heterogeneous evidence justifies overriding a trusted default prediction. We study this post-inference problem t...
160. Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation ​
Author: Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.02087v1 Announce Type: new Abstract: Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introduces challenges distinct from classical RL, with implications for induci...
161. Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents ​
Author: Qi Liu, Yiqun Chen, Zidan Chen, Yan Gao, Yi Wu, Yao Hu, Jiaxin Mao, Fengbin Zhu, Tat-Seng Chua
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.IR
arXiv:2608.02097v1 Announce Type: new Abstract: Search agents now answer questions that take dozens of searches to settle, yet how such an agent reads a page has drawn far less attention than how it finds one. Nearly all of them use one of two document interfaces, and both tie a page to the moment i...
162. MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents ​
Author: Jiajun Dong, Yutao Hu, Fengrui Fan, Shihan Dou, Yueming Wu, Deqing Zou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02113v1 Announce Type: new Abstract: Large language model (LLM) agents must retain and use cross-step information to act coherently in long-horizon tasks. Existing methods improve memory accessibility, yet action-relevant information may still fail to guide the current decision because it...
163. Beyond Solution-Centric Search: Adaptive Inquiry and Knowledge Revision for Autonomous ML Engineering ​
Author: Shaokang Fu, Yulong Tao, Linbo Jin, Jiarong Zhao, Qiming Shi, Tianjun Pan, Haonan Li, Chengyu Wang, Jia Wu, Chengfu Huo
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02143v1 Announce Type: new Abstract: Long-horizon autonomous research tasks such as machine learning engineering require systems to make interdependent decisions under a limited budget. Existing LLM-based agents typically organize candidate-solution improvement through tree, graph, or cha...
164. Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning ​
Author: Yijun Zhang, Yule Xie, Jiaxin Ding, Xin Ding, Fan Xu, Haoxiang Zhang, Luoyi Fu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02149v1 Announce Type: new Abstract: Reinforcement learning has become a central paradigm for improving the reasoning capabilities of large language models. Existing methods generally aim to reduce the failure probabilities induced across problems. In this paper, we introduce a moment-bas...
165. Auditing Data Provenance in LLM Fine-tuning via Intrinsic Distributional Fingerprints ​
Author: Zirui Huang, Yunlong Mao, Wei Tong, Tingting Wu, Xin Ge, Sheng Zhong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02154v1 Announce Type: new Abstract: The proliferation of customized Large Language Models (LLMs) poses critical risks of Data Intellectual Property (Data IP) infringement via unauthorized fine-tuning on proprietary data. Existing audit techniques are limited, as they require intervention...
166. From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution ​
Author: Can Wang, Haoran Chen, Haowen Gao, Hao Ding, Zhaoyang Liu, Zhiying Tu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02163v1 Announce Type: new Abstract: Deep research benchmarks require expert-level tasks and reliable evaluation grounded in task-specific knowledge. Existing benchmarks rely heavily on expert authoring or pre-existing human-authored materials, while fully automatic construction struggles...
167. From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents ​
Author: Jiajia Song, Bobo Li, Haiwen Yi, Zibo Ji, Meishan Zhang, Hao Fei, Min Zhang, Mong-Li Lee, Wynne Hsu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02171v1 Announce Type: new Abstract: Large Language Models have enabled increasingly capable autonomous agents, yet personalization remains critical for making such agents practically useful. Recent benchmarks have begun evaluating personalization in agents, but they largely rely on stati...
168. PAC Approximation and DIRECT Optimization for Parametric Markov Models ​
Author: Zhiming Chi, Ying Liu, Andrea Turrini, Lijun Zhang, David N. Jansen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.FL, cs.LO
arXiv:2608.02184v1 Announce Type: new Abstract: In this paper, we consider the parameter synthesis and optimization problem for parametric Markov decision processes (pMDPs), the extension of classical MDPs where exact probability values are replaced by parametric expressions. Computing the rational ...
169. MEGRAG: Multi-Granular Evidence Graphs for Answer-Aware Multi-Hop RAG ​
Author: Weidong Bao, Yingying Sun, Jun Yang, Yilin Wang, Zili Wei, Yubin Bao, Fangling Leng, Minghe Yu, Tiancheng Zhang, Ge Yu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02195v1 Announce Type: new Abstract: Multi-hop question answering is a fundamental challenge in retrieval-augmented generation (RAG), because deriving an answer requires integrating dispersed evidence. Iterative RAG (iRAG) is widely used for this challenge, but existing methods have two l...
170. PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs ​
Author: Haojie Hu, Chenhao Dang, Yaojia Liu, Hengrui Kang, Conghui He, Weijia Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02218v1 Announce Type: new Abstract: Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs; direct image generation is not element-editable, while coding-agent work...
171. Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability ​
Author: Abdullah Mamun, Shovito Barua Soumma, Hassan Ghasemzadeh
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.02238v1 Announce Type: new Abstract: Ensuring trust in AI systems is essential for the safe and ethical integration of machine learning systems into high-stakes domains such as digital health. Key dimensions, including robustness, explainability, fairness, accountability, and privacy, nee...
172. Homebot: A Personal AI Agent for Conversational Home Assistance and Automation ​
Author: Shengyuan Ye, Yixin Zhang, Han Liang, Liekang Zeng, Jiangsu Du, Mu Yuan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02254v1 Announce Type: new Abstract: \texttt{Homebot} is a locally deployable AI agent for conversational household assistance and automation. It accepts voice and instant-messaging requests through a shared runtime that combines language-model responses with registered tools and task-spe...
173. Self-Certification of Representation Adequacy: Sequential Certification at Minimum Task Loss ​
Author: Zijie Huang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.02267v1 Announce Type: new Abstract: Agents that act on a compressed representation of their history face a structural risk: if the representation aliases histories with different optimal actions, no rule measurable with respect to the representation can avoid an irreducible per-round los...
174. Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories ​
Author: Shuai Shao, Kangning Zhang, Qingyao Li, Shijian Wang, Hao Wang, Wenxiang Jiao, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02276v1 Announce Type: new Abstract: Agents built around large language models continually accumulate interaction trajectories during deployment, yet their behavior typically remains fixed. Beyond updating model weights, these trajectories can improve the agent harness that constructs con...
175. SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation ​
Author: Zelin Tan, Yiqun Zhang, Hao Li, Zhiyao Cui, Hejia Geng, Shao Zhang, Hangfan Zhang, Yang Chen, Xiaosong Wang, Lilong Wang, Zhenfei Yin, Shuyue Hu, Chen Zhang, Lei Bai
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02287v1 Announce Type: new Abstract: Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that current models can effectively identify, apply, and coordinate them. To impr...
176. Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning ​
Author: Yiqing Liu, Zihao Wang, Hantao Yao, Wu Liu, Yongdong Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02291v1 Announce Type: new Abstract: Multi-agent reasoning (MAR) improves reasoning reliability through iterative solution exchange and refinement. Existing adaptive MAR methods typically learn routing decisions from query-level labels or trajectory-level returns, but such coarse supervis...
177. MechGeo: Autoformalizing and Proving Euclidean Geometry in Lean 4 ​
Author: Hao Shen, Junyu Guo, Tian Cui, Yuxuan Xiao, Lihong Zhi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LO
arXiv:2608.02295v1 Announce Type: new Abstract: We present MechGeo, a Mathlib native agentic framework that jointly addresses faithful autoformalization and certified proof construction for Euclidean geometry. In this framework, GeoFormalizer represents informal problems in GeoIR, deterministically ...
178. Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit ​
Author: Jingxi Wei
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SE
arXiv:2608.02302v1 Announce Type: new Abstract: Long-horizon coding-agent trajectories are poorly matched to the credit units available to train on: a single action has no stable value, an episode label merges productive exploration with abandoned directions, and a fixed window cuts where the loggin...
179. Hard Constraints, Smooth Gradients: Learning Feasible Inventory Policies via Differentiable Projection ​
Author: Patrick Helm, Jan-Niklas Doerr, Joren Gijsbrechts, Stefan Minner
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.02343v1 Announce Type: new Abstract: Many operational problems are constrained sequential decision processes with large, combinatorial action spaces and interdependent feasibility constraints. Mixed-integer linear programs (MILPs) handle such constraints flexibly but scale poorly in stoch...
180. Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling ​
Author: Qinwen Wang, Jieping Luo, Aoxiang Qin, Ruoyu Zhao, Jianxiong Tang, Wei Zhang, Zhichao Lu, Luziwei Leng
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02347v1 Announce Type: new Abstract: Recurrent linear attention models (RLAs) such as Mamba offer efficient linear-time sequence modeling as an alternative to Transformers, yet their fixed-capacity recurrent states limit long-sequence modeling. Drawing inspiration from hierarchical human ...
181. KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model Improvement ​
Author: Gusseppe Bravo-Rocca, Jordi Guitart, Ajay Dholakia, David Ellison, Puneet Jain
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02351v1 Announce Type: new Abstract: Data drift poses significant challenges for machine learning systems in production, requiring continuous model updates to maintain performance. We present KC-Agent, a dual-process cognitive architecture for automated ML model improvement that combines ...
182. SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents ​
Author: Yue Yao, Shengyuan Wang, Xin Chen, Minke Zhang, Jia He, Bingjun Luo, Tom Gedeon
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02356v2 Announce Type: new Abstract: Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challenge is not merely to retrieve individually relevant skills, but to identify a complete and executable skill composi...
183. Faster-WAM: Do World Action Models Need Deep Action Modules? ​
Author: Liheng Ma, Rui Heng Yang, Zhanguang Zhang, Mateo Clemente, Ziwen Hu, Tongtong Cao, Yingxue Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.RO
arXiv:2608.02365v1 Announce Type: new Abstract: World Action Models (WAMs) couple robot action prediction with video world models. Existing WAMs with shared-backbone and Mixture-of-Transformers designs generally tie the depth of the action module to that of the video backbone, resulting in substanti...
184. Chess on Ice: Curling Tactical Decision-Making via Backward Induction and Deep Reinforcement Learning ​
Author: Patrick Oberlin, Matteo Cederle, Aren Karapetyan, Saverio Bolognani, Gian Antonio Susto, Florian D"orfler
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02379v1 Announce Type: new Abstract: Curling is often referred to as "Chess on Ice", owing to the tactical complexity of its decision-making process. Yet unlike chess, curling remains largely underexplored from a machine learning perspective, with prior work confined mainly to statistical...
185. Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training ​
Author: Zhiyuan Wang, Shengcai Liu, Jiahao Wu, Ning Lu, Hui Ouyang, Shaofeng Zhang, Haoze Lv, Ke Tang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.02391v1 Announce Type: new Abstract: Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution strategies (ES) enable memory-efficient full-parameter post-training without backpropagation and can eve...
186. MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models ​
Author: Victor Ojewale, Ro Encarnaci'on, Suresh Venkatasubramanian, Dana'e Metaxa
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.02409v1 Announce Type: new Abstract: Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms record satisfaction without task purpose. Together, they leave a critica...
187. xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding ​
Author: Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim, Mudhakar Srivatsa, Raghu K. Ganti, Minjia Zhang, Naigang Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02438v1 Announce Type: new Abstract: Block-diffusion drafters like dFlash generate an entire block of draft tokens in a single forward pass, drastically reducing the overhead of multiple-token drafting in speculative decoding. The crucial final step of the single-pass discrete denoising p...
188. Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce ​
Author: Shicheng Fan, Mingdai Yang, Duohao Wang, Canyu Chen, Yongfeng Zhang, Hua Wei, Manling Li, Julian McAuley, Kun Zhang, Philip S. Yu, Kejing Yu, Zhiwei Liu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02441v1 Announce Type: new Abstract: In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to agents. Commer...
189. Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks ​
Author: Xuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Yihua Zhu, Hu Wei, Bing Zhao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.02442v1 Announce Type: new Abstract: Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasoning capability targeted by the problem. We identify Solution Hacking, a fa...
190. ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision ​
Author: Wei-Jung Huang, Bonan Shen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02444v1 Announce Type: new Abstract: LLM-agent evaluations often produce task outcomes long before the full benchmark run is complete. A partial score is tempting to report, but it does not show whether the observed tasks support the same conclusion as the completed evaluation. Early task...
191. Infinite Trace Objectives with Finite Trace Techniques: Translating LTL to LTLf+ ​
Author: Christoph Weinhuber, Maximilian Prokop, Giuseppe De Giacomo, Moshe Y. Vardi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.FL, cs.LO
arXiv:2608.02454v1 Announce Type: new Abstract: Linear Temporal Logic (LTL) is one of the most widely adopted languages for specifying temporal extended objectives in AI, with applications ranging from reactive synthesis to stochastic planning in Markov decision processes and reinforcement learning....
192. Real-Time Detection and Repair of LLM Agent Failures ​
Author: Sunny Dubey
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SE
arXiv:2608.02464v1 Announce Type: new Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself. We ask how much d...
193. Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions ​
Author: Nicole Mitchell, Dhruv Agarwal, Maty Bohacek, Remi Denton, Roma Patel
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02491v1 Announce Type: new Abstract: Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features can introduce longitudinal risks---cognitive, developmental and socio-a...
194. CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization ​
Author: Chuyan Chen, Peng Sun, Kun Yuan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02502v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a prom...
195. Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation ​
Author: Michael Farmer
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.IR
arXiv:2608.02505v1 Announce Type: new Abstract: Can scientific abduction occur without continuous sensorimotor embodiment? Recent arguments in AI and philosophy of science hold that genuine hypothesis generation requires an agent continuously coupled to the physical world. We defend a narrower claim...
196. Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies ​
Author: Sterre Lutz, Dani"el Vos, Matthijs T. J. Spaan, Anna Lukina
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.02509v1 Announce Type: new Abstract: Sequential decision-making in real-world applications often involves uncertainty about the environment's model. Uncertain Markov decision processes (UMDPs) represent the possible environments as a set of MDPs with shared states and actions but potentia...
197. Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation ​
Author: Natalie Isak, Matthew Dressman
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.02518v1 Announce Type: new Abstract: The most capable AI deployments are not single models but ensembles of specialized agents that delegate and act in coordination. This architecture unlocks powerful new capabilities, and it also introduces risks that existing frameworks for monitoring, ...
198. A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI ​
Author: Taye Akinrele, Sindhuja Penchala, Noorbakhsh Amiri Golilarz, Sudip Mittal, Shahram Rahimi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02553v1 Announce Type: new Abstract: Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustained reasoning, adaptive behavior, persistent memory, and self-regulation. While generative and agentic AI have demonstrated impressive c...
199. AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies ​
Author: Qiushi Lin, Chaojie Zhang, 'I~nigo Goiri, Aditya Akella, Ricardo Bianchini, Jovan Stojkovic
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.OS
arXiv:2608.02569v1 Announce Type: new Abstract: The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent, and prototyping a single policy takes months. Agenti...
200. Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results ​
Author: Sheng Xu, Junhua Wang, Boyuan Huang, Ke Jia, Jiadun Zhu, Zhen Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.11183v2 Announce Type: cross Abstract: Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plausible responses. We study inference-time feed-forward network (FFN) intervention for improving structu...
201. Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response ​
Author: Elle Lazarski, Jaime Fern'andez Fisac
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.27508v1 Announce Type: cross Abstract: Assistance games formalize human-robot collaboration under asymmetric information: the human knows the goal, while the robot must infer it from observation and interaction in order to assist effectively. In general, computing optimal assistance game ...
202. Cost-Effective Automated Judging of Natural-Language Mathematical Proofs ​
Author: Benjamin Grayzel
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.00004v1 Announce Type: cross Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth...
203. RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review ​
Author: Shuyu Guo, Wenxiang Hu, Yuyue Zhao, Yougang Lyu, Xiaohui Yan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00005v1 Announce Type: cross Abstract: Peer review at major venues is under unprecedented submission pressure, motivating the use of large language models (LLMs) as review assistants. Existing LLM-based reviewers, however, face two structural limitations. First, they map manuscripts direc...
204. MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents ​
Author: Bohan Tang, Yiwen Guo
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.00007v1 Announce Type: cross Abstract: Equipping Large Language Models (LLMs) with human-like personas is crucial for agentic applications, such as role-play and user simulation. Traditional prompt-based methods rely on descriptive conditioning by injecting static textual profiles, which ...
205. AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents ​
Author: Ahmed Cherif
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00009v1 Announce Type: cross Abstract: Long-term memory remains a critical bottleneck for conversational AI agents, whose finite context windows cannot support coherent recall across thousands of turns. We present AgentMemBench, a unified, reproducible benchmark evaluating five memory man...
206. DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis ​
Author: Wasim Madha, Nityanand Mathur, Hamees Sayed, Apoorv Singh, Sameer Khurana, Akshat Mandloi, Sudarshan Kamath
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00011v1 Announce Type: cross Abstract: Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale models and training data and decode tokens sequentially, while non-autoregressive approaches im- prove ...
207. Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams ​
Author: Fengxiang Wang, Qiuyang Yu, Yueying Li, Mingshuo Chen, Chengchi Fei, Kaiyi Xu, Lixin Gu, Wangxu Wei, Junchao Gong, Lipeng Ma, Jiong Wang, Fenghua Ling, Wenlong Zhang, Xue Yang, Wenjing Yang, Ben Fei, Long Lan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.LG
arXiv:2608.00012v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated. Existing remote sensing benchmarks largely re...
208. What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs ​
Author: Ziran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen, Jingang Wang, Xunliang Cai
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV
arXiv:2608.00013v1 Announce Type: cross Abstract: Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: compute-based scaling laws fail to generalize across model families, and...
209. CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs ​
Author: Zihan Nan, Yang Gu, Wei Liu, Xi Yan, Zhou Liu, Hao Liang, Wentao Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2608.00018v1 Announce Type: cross Abstract: Tabular data processing is central to data work, and LLM-based assistants have recently shown promising capabilities in supporting such tasks. However, existing benchmarks primarily focus on table reasoning under single-turn, fully specified instruct...
210. Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models ​
Author: Liang Guo, Lin Shaochong, Shen Zuo-Jun Max, Zhang Kun
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.00019v1 Announce Type: cross Abstract: Deploying large language models (LLMs) for operations research (OR) tasks remains challenging because correctness depends on a coherent modeling process, not merely a correct final answer. Standard autoregressive generation operates on a myopic polic...
211. Role Steering of Language Models for Social Simulations ​
Author: Isaac Song, Mohammed Rehan Parwani, Glenn Matlin, Emile Anand, Akhil Theerthala, Arjun Chatterjee, Maria Kostylew, Yonadav G. Shavit, Sebastien Krier, Mark Riedl
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00023v1 Announce Type: cross Abstract: Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated population. We introduce an activation-steering screening workflow for role-conditioned agents: define a ...
212. Exploring More to Solve More: Boosting Diversity in Text Diffusion Models via Entropy-Based Guidance ​
Author: Jingwei Zhang, Haoyu Lei, Zijin Feng, Jiacheng Sun, Farzan Farnia
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00024v1 Announce Type: cross Abstract: Although diffusion models have revolutionized continuous domains like image synthesis through high quality generations and controllable guidance mechanisms, bringing this controllability to the discrete, sequential nature of text remains an open chal...
213. Width, Memory, and Delay: A Resource Accounting for the Limits of Flat Multi-Agent Systems ​
Author: Oleksandr Kuznetsov, Emanuele Frontoni
Published: 8/4/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2608.00028v1 Announce Type: cross Abstract: A recurring question in the design of scalable multi-agent systems -- from robot swarms to collectives of large-language-model (LLM) agents -- is whether adding more agents can, on its own, overcome performance limits, or whether a qualitatively \emp...
214. SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach ​
Author: Gayathri V Kondapalli, Alexander Ng, Hirsh Pithadia, Rahul Monish, Harvey Yorke, Amir Kayhani
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00030v1 Announce Type: cross Abstract: Specialised retrieval agents typically surface higher quality results than general-purpose search, but selecting the optimal agent for a given query remains an open problem. Current approaches route queries based on inferred topic or intent, however ...
215. XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding ​
Author: Hongchen Wei, Yuanzhe Wang, Bei Liu, Yifan Yang, Qi Dai, Ruichun Ma, Kai Qiu, Yunsheng Li, Dongdong Chen, Chong Luo, Zhenzhong Chen, Baining Guo
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00036v1 Announce Type: cross Abstract: Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages. Some questions also require comparing related reports. Reliab...
216. Fast Generation of Representative Synthetic Dataset with Salsa to Train ATR Models with Electromagnetic Couplings Data-Augmentation ​
Author: Benjamin Camus (DGA.MI), Julien Houssay (DGA.MI), Corentin Le Barbu (DGA.MI), Eric Monteux (DGA.MI), C'edric Saleun (DGA.MI), Jean-Christophe Louvign'e (DGA.MI)
Published: 8/4/2026, 4:00:00 AM
Categories: eess.SP, cs.AI
arXiv:2608.00037v1 Announce Type: cross Abstract: This work focuses on training Automatic Target Recognition (ATR) models using simulated Synthetic Aperture Radar (SAR) images to circumvent the lack of real measurements. To obtain robust and versatile ATR models, simulation needs to generate massive...
217. Google's AI & Economy ATLAS v1.0: Mapping Gemini Usage in the Economy ​
Author: Zanna Iscenko, Scott Strand, Yiyuan Chen, Guillaume Aimard, Mihai Codreanu, Vivek Sampathkumar, Alex Imas, Julian Jacobs, Evalyne Muiruri, Juan Mateos-Garcia, Jia Jen Ng, Samirah Javed, Josh Martin, Omar Ajmeri, Denis Calin, Andrew Kim, Fabien Curto Millet, James Manyika
Published: 8/4/2026, 4:00:00 AM
Categories: econ.GN, cs.AI, q-fin.EC
arXiv:2608.00038v1 Announce Type: cross Abstract: This paper introduces the AI & Economy ATLAS (Activity, Task, Landscape, and Adoption Study), an ongoing economic research initiative using Google AI usage data. The first iteration of ATLAS is built on 15 million de-identified interactions across th...
218. Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study ​
Author: Ramesh B. Paramkusham
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00042v1 Announce Type: cross Abstract: Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrained, high-stakes environments including healthcare, legal services, and financial analysis. While performance ...
219. Multimodal Wearable-Based Olfactory-Induced Emotion Recognition in Arousal-Valence Dimensions ​
Author: Chen-Yang Xu, Lan Zhang, Fei-Yi Fan, Bin Hu, Qing-Hao Meng
Published: 8/4/2026, 4:00:00 AM
Categories: eess.SP, cs.AI
arXiv:2608.00043v1 Announce Type: cross Abstract: Olfaction is important for emotion regulation because it acts as a non-intrusive and cognitively lightweight pathway that directly engages the brain s affective circuitry and achieves unobtrusive emotional modulation. This trait is essential for adva...
220. Not All EEG Moments Are Equal: Position-Adaptive Time Scheduling for EEG Generation ​
Author: Boheng Liu, Ziyu Li, Chenghua Duan, Qing Li, Xia Wu
Published: 8/4/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG
arXiv:2608.00048v1 Announce Type: cross Abstract: Electroencephalography (EEG) generation is essential for alleviating data scarcity and enabling large scale neural modeling in brain computer interface applications. However, existing flow based approaches assume that every channel and every time seg...
221. Automated ECG Interval Measurement and Wave Delineation Using Fast Fourier Convolution ResNet ​
Author: Farhan Adam Mukadam, Harshit Mishra, Nachiket Makwana, Pradyot Tiwari, Subramani Kandasamy, KVS Hari
Published: 8/4/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG
arXiv:2608.00058v2 Announce Type: cross Abstract: Accurate measurement of ECG intervals, including PR, QRS duration, and QT/QTc, is central to cardiac diagnosis, yet the published ECG delineation literature evaluates performance almost exclusively as fiducial-point timing errors on small curated dat...
222. Neural Circuit Function Inference with LLMs ​
Author: Yijie Yin (Department of Physiology, Development and Neuroscience, University of Cambridge, Cambridge, UK, MRC Laboratory of Molecular Biology, Cambridge, UK), Albert Cardona (MRC Laboratory of Molecular Biology, Cambridge, UK, Department of Physiology, Development and Neuroscience, University of Cambridge, Cambridge, UK)
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00059v1 Announce Type: cross Abstract: The success of connectome mapping now shifts the challenge of understanding the nervous system to the interpretation of neural circuits. Here, we devise a new automated method, LLantia (LLM automated neural circuit inference and analysis), to systema...
223. ELECTRIC: Evidential Learning-Enhanced CT Reconstruction via Iterative Correction ​
Author: Ge Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00060v1 Announce Type: cross Abstract: Here we introduce ELECTRIC (Evidential Learning-Enhanced CT Reconstruction via Iterative Correction), a physics-guided Bayesian formulation. An evidential neural network provides an image proposal and an error-predictive epistemic-uncertainty surroga...
224. Empirical investigation of 3D CT Foundation Models and Unsupervised Adaptation for Head and Neck Cancer Recurrence Prediction ​
Author: Bilel Guetarni, Feryal Windal, David Pasquier, Halim Benhabiles
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00071v1 Announce Type: cross Abstract: The rapid emergence of 3D CT foundation models has opened new avenues for predictive modeling from CT imaging, offering a compelling alternative to traditional radiomics which is known to suffer from reproducibility issues and sensitivity to acquisit...
225. Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs ​
Author: Vahidin Hasic, Chao Wang, Luis C. Garcia-Peraza-Herrera, David Watson, Senka Krivic
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00076v2 Announce Type: cross Abstract: Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. While existing explainability methods identify influential image regions or text tokens, they canno...
226. SymNet: A Multi-Task Network for Joint Radio Map Reconstruction and Transmitter Localization ​
Author: Lyuzhou Ye, Thanh Dat Le, Yan Huang
Published: 8/4/2026, 4:00:00 AM
Categories: eess.SP, cs.AI
arXiv:2608.00087v1 Announce Type: cross Abstract: Accurately predicting directional radio maps is essential for wireless applications, yet prior approaches primarily focus on omnidirectional signals and typically treat transmitter localization and signal map reconstruction as separate tasks. In omni...
227. Logographic Character Visual Pretraining via Semantic-based Contrastive Learning ​
Author: Daqian Shi, Wei Cao, Xiaoyu Zheng, Lida Shi, Xiaolei Diao, Cedric M John
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00096v1 Announce Type: cross Abstract: Current deep learning-based character vision studies, e.g., text recognition, character image denoising, and historical text completion, are offering new solutions for learning, managing, and utilizing character resources. However, the performance of...
228. Conservation laws determine what physical learning remembers ​
Author: Bijaya Dangol
Published: 8/4/2026, 4:00:00 AM
Categories: cond-mat.soft, cond-mat.dis-nn, cs.AI, cs.LG
arXiv:2608.00097v1 Announce Type: cross Abstract: Physical learning rules such as equilibrium propagation (EP), coupled learning (CL), and adjoint coupled learning (AL) train resistive networks through local measurements. In the small-nudge limit EP and CL exactly conserve the conductance mass K = (...
229. GRAIN: Molecules Are Not the Right Granularity -- Active-Ingredient Modeling for Safe Medication Recommendation ​
Author: Juao Fan, Jinhan Li, Shengxin Zhu
Published: 8/4/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI
arXiv:2608.00098v1 Announce Type: cross Abstract: Medication recommendation from electronic health records must balance predictive accuracy against the risk of adverse drug-drug interactions (DDIs) under polypharmacy. Existing safety-aware recommenders operate at one of two granularities: the drug c...
230. LLMBDC: Language Model for Biological Domains Oriented Clustering of Gene Ontology ​
Author: Ximing Ran, Jie Xu, Peng Jin, Zhaohui Qin, Zhexing Wen, Jiaying Lu
Published: 8/4/2026, 4:00:00 AM
Categories: q-bio.GN, cs.AI
arXiv:2608.00099v1 Announce Type: cross Abstract: Gene Ontology (GO) enrichment analysis is a foundational tool for translating large-scale genomic data into biological insights, but typically yields hundreds of redundant terms that obscure overarching themes. Existing summarization tools rely on fi...
231. SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models ​
Author: Satvik Tripathi, Mustafa Ege Seker, Kristian Quevada, Ebubechukwu D Enwerem, Pratham Khandelwal, Emine Meltem, Bera Koca, Shahriar Faghani, Jacinta Arnold, Dania Daye, Tessa S. Cook
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00100v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly being evaluated for medical imaging, but many available benchmarks emphasize disease classification, report generation, or broad visual question answering rather than the spatial and anatomical reasoning...
232. Skillsets on the Chain: A Blockchain-based Zero-Trust Framework for Agentic AI Networking ​
Author: Yayu Gao, Yong Xiao, Hao Hu, Xubo Li, Zhiwei Liu, Yingyu Li, Guangming Shi, Ping Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.NI
arXiv:2608.00104v1 Announce Type: cross Abstract: Agentic AI networking (AgentNet) systems rely heavily on third-party skillset implementations and distributed multi-agent collaboration, yet they face major claim-to-capability inconsistencies and security vulnerabilities under trust-by-declaration a...
233. MetaRoute-Bench: Evaluating Meta-Decision Policies for Agentic Workflow Routing ​
Author: Natan Vidra, Alina Kapanova, Arun Kanhai, Spurthi Setty
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.00107v1 Announce Type: cross Abstract: Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an intermediate result, or recover from failure. These meta-decisions affect not only task success but ...
234. EEG-JEPA: Structured Latent Prediction for EEG Foundation Models ​
Author: Jinhao Li, Zhiyuan Ma, Xueqiao Han, Zhongye Xia, Xinche Zhang, Shanghong Xie, Yixuan Liu, Yongjian Li, Runmin Gan, Tianlin Huo, Sen Song
Published: 8/4/2026, 4:00:00 AM
Categories: eess.SP, cs.AI
arXiv:2608.00114v1 Announce Type: cross Abstract: Electroencephalography (EEG) foundation models aim to learn reusable representations from large-scale unlabeled recordings. A common pretraining strategy is masked waveform reconstruction, but applying supervision directly to noisy EEG may encourage ...
235. Counting the Cost of War Under Satellite Embargo: Zero-Shot Estimation of Impacted Infrastructure ​
Author: Saleh Sakib Ahmed, M. Sohel Rahman
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00119v1 Announce Type: cross Abstract: Rapid estimation of impacted structures - critical for conflict-zone humanitarian response - is frequently hindered by post-strike satellite data embargoes and imagery blackouts. We bypass this operational bottleneck by reframing impacted building ma...
236. LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations ​
Author: Yan Fang, Jialin Chen, Chun Gan, Hang Yu, Mingjun Nie, Yeyu Zhang, Fengxiang He, Ching Law
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.GT, cs.LG
arXiv:2608.00123v2 Announce Type: cross Abstract: LLM-native advertising embeds sponsored content directly into model-generated responses, shifting the unit of sale from a fixed slot to a moment within an evolving conversation. Existing LLM ad-auction mechanisms primarily operate within a single res...
237. A Fortran General-Purpose Transpiler: Proof of Concept ​
Author: Shivamshan Sivanesan, Kazem Ardaneh
Published: 8/4/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.CL, cs.MS, cs.SE
arXiv:2608.00130v1 Announce Type: cross Abstract: Fortran has been the cornerstone of high-performance computing for decades and remains unmatched in many domains. Yet the language faces an expertise gap: a new generation of scientists is barely familiar with it, while many experienced Fortran devel...
238. Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity ​
Author: Ramya Varunsegar
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.00143v1 Announce Type: cross Abstract: Automated attack chain generation is critical for modern cybersecurity, yet manual construction fails to scale as adversary behaviors expand. While classical AI planning using PDDL offers a formal method to automate this process, it relies on the acc...
239. DiffusionGemma Technical Report ​
Author: DiffusionGemma Team, Adrien Ali Ta"iga, James Assiene, Daniele Calandriello, Rahma Chaabouni, Jo~ao Gante, Tamara von Glehn, Nate Keating, Chris Knutsen, Martin Kukla, Tianlin Liu, Ivan Lobov, Ofir Nabati, Jo~ao Gabriel Oliveira, Nicolas Perez-Nieves, Nastasia Prutianova, Bobak Shahriari, Jean Tarbouriech, Pavel Tyletski, \c{C}a\u{g}lar "Unl"u, Cindy Wu, Glenn Cameron, Jerome Connor, Sertan Girgin, Maarten Grootendorst, Alon Levkovitch, Eliya Nachmani, Omar Sanseviero, Piotr Stanczyk, Quentin Berthet, Andrew Campbell, Cl'ement Crepy, Valentin De Bortoli, Arnaud Doucet, Romuald Elie, Alexandre Galashov, Klaus Greff, Alexis Jacq, David Ruhe, Yu-Han Wu, Sebastian Flennerhag, Brendan O'Donoghue, George Scrivener, Shantanu Thakoor
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00146v1 Announce Type: cross Abstract: We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in pa...
240. A Synthetically-accessible Universe of Chemically Recyclable Polymers ​
Author: Anagha Savit, Wei Xiong, Harikrishna Sahu, Shivank S. Shukla, Will R. Gutekunst, Rampi Ramprasad
Published: 8/4/2026, 4:00:00 AM
Categories: cond-mat.soft, cs.AI
arXiv:2608.00149v1 Announce Type: cross Abstract: Polymers synthesized via ring-opening polymerization (ROP) of cyclic monomers represent an important class of materials due to their chemical recyclability and possible insertion in several critical applications. We present a dataset of 1 million syn...
241. Exposed by Design: A Dynamic Security Assessment of Internet-Facing MCP Servers at Scale ​
Author: Nicol'as Padilla
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.00150v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has seen rapid adoption since its November 2024 launch, with over 21,000 server instances detectable on the public internet. We present the first dynamic behavioral security assessment of internet-facing MCP servers, ...
242. Optimising for Flourishing: Flourishing Metrics and Return on Flourishing as Success Criteria for Artificial Intelligence and Post-AGI Economic Systems ​
Author: Keyun Ruan, Jonathan D. Teubner, John M. Bremen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, econ.TH
arXiv:2608.00151v2 Announce Type: cross Abstract: Current evaluation frameworks for artificial intelligence focus mainly on capability, safety, and proxies such as adoption, engagement, efficiency, productivity, and financial return. These criteria are necessary but insufficient because they do not ...
243. Response Magnitude as a Dominant Signal for Held-Out CRISPRi Perturbation Effect Prediction ​
Author: Mehrdad Shoeibi, Niloofar Yousefi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.00152v1 Announce Type: cross Abstract: Predicting the magnitude of a CRISPRi perturbation's transcriptomic effect on held-out target genes is an important open problem in single-cell biology. Recent work has documented that simple baselines often match or exceed deep perturbation predicto...
244. Inference-Time Policy Alignment for Fair Reinforcement Learning ​
Author: Umer Siddique, Peilang Li, Conor Wallace, Yongcan Cao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.00175v1 Announce Type: cross Abstract: Deep reinforcement learning (RL) agents achieve strong performance by optimizing scalar reward functions. However, once deployed, the policies of these RL agents are often rigid and costly to adapt to new performance criteria. For instance, an agent ...
245. Cross-Benchmark Generalization in Long-Horizon Agents ​
Author: Sushant Mehta, Logan Ritchie, Liudas Panavas, Edwin Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.00181v1 Announce Type: cross Abstract: For reinforcement learning (RL) in self-contained environments, a policy can get rewards by exploiting environment-specific regularities (tool schemas, grader parsing, task templates) rather than by acquiring transferable skill, and an in-distributio...
246. Evidence-Unit Fairness and the Limits of Query-Adaptive Sparse-Dense Fusion in Financial Document Retrieval ​
Author: Chenyu Wu, You Lin
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CE, cs.AI
arXiv:2608.00183v1 Announce Type: cross Abstract: Retrieval over financial filings is difficult because queries are short and acronym-heavy while the answer-bearing evidence sits inside long, table-dense documents. We study sparse-dense hybrid retrieval on FinDER, a benchmark of expert-annotated que...
247. Hierarchical BM25: Lexical Search at Billion-Document Scale ​
Author: Umesh Deshpande, Swaminathan Sundararaman
Published: 8/4/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.00229v1 Announce Type: cross Abstract: A flat BM25 index over one billion documents occupies about 400 GB. Holding it in memory requires DRAM proportional to corpus size. Serving it from disk takes 4-12 seconds per query. Exact top-k lexical retrieval at this scale is therefore impractica...
248. Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind ​
Author: Kejia Zhang, Youran Sun, Chugang Yi, Haizhao Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.MA, q-bio.NC
arXiv:2608.00261v1 Announce Type: cross Abstract: Do frontier vision-language models present a coherent Theory-of-Mind (ToM) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of nine frontier VLMs on...
249. Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers ​
Author: Kamil Ksi\k{a}.zek, Piotr Suszy'nski, Micha{\l} Jan W{\l}odarczyk, Jacek Tabor, Przemys{\l}aw Biecek
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00264v1 Announce Type: cross Abstract: Vision foundation models, such as DINOv2, learn highly expressive representations but rely on massive, opaque architectures that demand substantial computational power and memory. To provide an interpretable-guided and efficient solution to this issu...
250. Hybrid Attention Estimation Pipeline for Adaptive HRI Using an Expressive Robotic Head ​
Author: Pablo Moraes, Monica Rodriguez, Christopher Peters, Hiago Sodre, Tobias Doernbach, Bruna Guterres, Ricardo Grando
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.00284v1 Announce Type: cross Abstract: This paper presents an applied case study on hybrid visual attention estimation for human-robot interaction using an expressive robotic head based on the InMoov ecosystem. The proposed pipeline combines a fast geometric perception layer with an indep...
251. Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct ​
Author: Mario Vega-Barbas, Lidia Mora-Valenciano, Iv'an Pau, Fernando Seoane, Farhad Abtahi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00285v1 Announce Type: cross Abstract: Sixteen language models drawn from ten families produced, on average, the semantic diversity of 1.69 distinct formulations of a psychotherapeutic case, against a single-model baseline of 1.43 from one model's own runs. Ensembles place more than one r...
252. OrEdge: Efficient Multi-Modal Anomaly Detection in Distributed Software Systems via Orthogonal-Domain Learning ​
Author: Amr M. Zaki, Farhoud Jafari Kaleibar, Honggeun Ji, Komal Sarda, Marin Litoiu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.00309v1 Announce Type: cross Abstract: We introduce Orthogonal-Edge (OrEdge), a lightweight framework for real-time anomaly detection in multi-modal distributed software systems. Unlike existing approaches that rely on computationally expensive attention- and graph-based architectures, Or...
253. ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression ​
Author: Renjie Liang, Zijian Xu, Jinqian Pan, Chengkun Sun, Zhengkang Fan, Shawn Li, You Qin, Mei Liu, Jie Xu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00345v1 Announce Type: cross Abstract: A 3D CT scan entering a vision-language model produces a long sequence of visual tokens, often thousands to tens of thousands per volume, and this sequence must be compressed before a language model can consume it. Token compression is well studied i...
254. Artificial Intelligence for the Characterization of Particles and Fibers by Optical Microscopy ​
Author: Simiao Sun, Kenneth Ng, Lynn Lee, Astrid Harth, Asami Odate, Aggelos Katsaggelos, Manuel Ballester Matito, Nicholas Eastaugh, Marc Walton
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00361v1 Announce Type: cross Abstract: Optical microscopy of particle and fiber dispersions involves interpreting subtle visual cues influenced by specimen morphology, chemical composition, magnification, and illumination conditions. We introduce an artificial intelligence (AI) distillati...
255. Artificial Intelligence and Modeling & Simulation: An Overview ​
Author: Niclas Feldkamp, Philippe J. Giabbanelli, Istvan David
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.00366v1 Announce Type: cross Abstract: Artificial intelligence (AI) and Modeling & Simulation (M&S) are increasingly intertwined, reflecting converging research needs across both communities, rapid technological advances such as the rise of generative AI, and the growing availability of d...
256. Tensor Probabilistic Model Checking of Finite-Horizon Markov Chains (Extended Version) ​
Author: Jianlin Li, Nick Guo, Peter Ye, Yizhou Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.MS, cs.PL
arXiv:2608.00374v1 Announce Type: cross Abstract: We reexamine the problem of verifying Markov chains with respect to step-bounded reachability probabilities. Prevailing approaches rely on encoding the state-transition matrix using either explicit or symbolic representations. While these approaches ...
257. Pretrain on Small Synthetic Data, Scale Large for Free: Symmetry-Aware Foundation Model for Logic Rule Induction ​
Author: Yin Jun Phua
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.LG
arXiv:2608.00383v1 Announce Type: cross Abstract: Logical rule induction seeks interpretable rules that transfer across propositional schemas. This requires respecting symmetries: atom naming, example order, polarity flips, and label swap. Enforcing exact symmetry by construction lets one trained in...
258. The Gate, Not the Cache: Gate Provenance Bounds the Closed-Loop Reliability of Training-Free VLA Token Skipping ​
Author: Qi Luo, Shuaijun Liu, Hao Zhao, Kunlin Li, Xiaobo Wang, Ningxing Su, Dongsheng Wang, Yun Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.00391v1 Announce Type: cross Abstract: Token skipping is a widely used training-free way to accelerate vision--language--action (VLA) models by bypassing computation for most visual tokens at each control step according to a gate. When the next gate is harvested from the previous accelera...
259. Verifiable Checks for Business Rule Consistency ​
Author: Joseph Tafese, Milad Hooshyar, Sam Bayless, Nick Feng, Arie Gurfinkel
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LO, cs.AI
arXiv:2608.00396v1 Announce Type: cross Abstract: Maintaining consistency between natural language documentation of business rules and their evolving internal implementations is a significant challenge in large-scale systems. We present SIRNA, a tool and framework for checking such consistency using...
260. DSETA: A Dual-Stage Continual Learning Framework for Travel Time Prediction in Dynamic Traffic Environments ​
Author: Yanming Lyu, Yue Cheng, Lingkun Li, Ruipeng Gao, Xinyue Liu, Hui Gao, Qiang Ni
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.00402v1 Announce Type: cross Abstract: Estimated Time of Arrival (ETA) prediction is a core component of intelligent transportation systems. As traffic congestion patterns become increasingly dynamic in large cities, maintaining high prediction accuracy poses a major challenge for ride-ha...
261. Boosting Generalizable Depth Estimation in Endoscopy by Mixture of Lightweight Experts and Intrinsic Image Alignment ​
Author: Liangjing Shao, Beilei Cui, Yiming Huang, Changjing Liu, Hongliang Ren
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00415v1 Announce Type: cross Abstract: Depth estimation is a significant task for 3D perception in endoscopic surgeries. However, illumination interference and feature diversity in various endoscopic scenes are still challenges for generalizable depth estimation and ego-motion estimation....
262. Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments ​
Author: Muhammad Faizan Raza (Luna), Shuo (Luna), Yang, Satish Mahadevan Srinivasan, Joanna F. DeFranco
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.IR
arXiv:2608.00419v1 Announce Type: cross Abstract: Large language models deployed in real-time, regulated settings face knowledge staleness, catastrophic forgetting, hallucination, and weak feedback loops. We present a unified, pattern-driven LLMOps architecture integrating real-time data ingestion, ...
263. AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction ​
Author: Ziqiang Cui, Han Shi, Bowei He, Yu Pan, Peiyang Liu, Shengyin Sun, Yankai Chen, Haoli Bai, Yichun Yin, Xue Liu, Chen Ma
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00434v1 Announce Type: cross Abstract: Multi-Token Prediction (MTP) has emerged as an effective paradigm that augments a shared Large Language Model backbone with auxiliary heads, training the model to predict several future tokens in parallel to enrich its supervision signal and accelera...
264. Distilling Reasoning Traces into Advisory Prompts for Software Engineering Tasks ​
Author: Faizan Faisal, Prem Devanbu, Toufique Ahmed
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.00437v1 Announce Type: cross Abstract: Language models are widely used for generating and otherwise processing code (e.g., identifying code hallucinations, possible inputs, or predicting outputs); however, LLMs can make mistakes, which can be serious. One key issue is that models are trai...
265. Beyond Static Anchors: Bounded Prototype Conditioning for Language-Free Medical Anomaly Detection ​
Author: Yibo Wan, Jinyu Cai, Seekiong-Ng
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00442v1 Announce Type: cross Abstract: Medical anomaly detection identifies abnormal images and localizes lesions under scarce supervision while generalizing across organs and modalities. Existing CLIP-based methods reduce annotation requirements through vision--language alignment, but th...
266. Structured Proxy Features for Multimodal NSCLC Survival Prediction from Pretreatment CT ​
Author: Huu Phong Nguyen, Delower Hossain, Ehsan Saghapour, Zhandos Sembay, Jake Y. Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00446v1 Announce Type: cross Abstract: Lung cancer results in roughly 1.8 million fatalities annually worldwide, with non-small cell lung cancer (NSCLC) comprising the majority of cases. Despite advancements in treatment, survival stratification remains challenging due to intratumoral het...
267. CeQe: Grounding Lexical Retrieval in Semantic Evidence ​
Author: Adam Kahirov, Umesh Deshpande, Swaminathan Sundararaman
Published: 8/4/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.00452v1 Announce Type: cross Abstract: Lexical retrieval (BM25) captures exact keyword matches and weights terms by corpus-wide significance, but it is blind to the semantic vocabulary gap: when a relevant document phrases an answer differently from the query, BM25 never retrieves it, and...
268. CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings ​
Author: Kaho Li, Pengyu Zeng, Yuqin Dai, Jun Yin, Tianjing Feng, Shuai Lu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.00473v1 Announce Type: cross Abstract: Architectural drawings violate the usual assumption behind multi-view reasoning: plans and sections are cuts, while elevations are facade projections, so corresponding components change appearance in ways camera motion cannot explain. We introduce Cr...
269. RadYOLO: Computationally Efficient 3D Object Detection and Segmentation in CT and MRI ​
Author: Kai Geissler, Laurens M"uller-Groh, Hans Meine
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.00508v1 Announce Type: cross Abstract: Object detection and segmentation in three-dimensional medical images is a very active area of research. However, most proposed deep learning models carry a high computational cost, and only few aim to be broadly applicable, achieve high detection pe...
270. Auditable Release Control for Pedagogical Leakage in LLM Tutors ​
Author: Nizam Kadir
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.00515v1 Announce Type: cross Abstract: Large language model tutors can be correct and helpful yet disclose an answer or decisive reasoning before that disclosure is authorized. We formalize this state- and action-dependent failure as pedagogical leakage and introduce an authorization-awar...
271. Rethinking and formalising the state across languages: a unified computational learning theory account ​
Author: Mohamed El Idrissi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00523v1 Announce Type: cross Abstract: The linguistic notion of state has traditionally been restricted to the construct (annexation) state of Afroasiatic languages and treated as a language-specific morphosyntactic phenomenon. This article argues instead that the state is a systemic, con...
272. SSTG-Nav: Metric-Grounded Spatial-Semantic Topological Graphs for Reusable Object Navigation ​
Author: Daojie Peng, Bingtao Wang, Jun Ma
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.00527v1 Announce Type: cross Abstract: Service robots operating for months in the same homes, offices, and facilities should become more reliable with experience instead of searching familiar space from scratch for every request. Yet ObjectNav is predominantly formulated as one-shot explo...
273. Unleashing the Power of Text: Text-Guided Flow Matching for Image Fusion under Complex Degradations ​
Author: Axi Niu (School of Computer Science, Northwestern Polytechnical University, Xi'an, China), Jieheng Li (School of Computer Science, Northwestern Polytechnical University, Xi'an, China), Kang Zhang (School of Electrical Engineering, KAIST, Daejeon, Republic of Korea), Qingsen Yan (School of Computer Science, Northwestern Polytechnical University, Xi'an, China), Jinqiu Sun (School of Aeronautics and Astronautics, Northwestern Polytechnical University, Xi'an, China), Yanning Zhang (School of Computer Science, Northwestern Polytechnical University, Xi'an, China)
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00530v1 Announce Type: cross Abstract: Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not only cause a loss of reliable modality-specific information in observed images but also hinder the fusion process. Recent studies indicate ...
274. Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages ​
Author: Sean Gip Lim, William Chandra Tjhi, Hai Leong Chieu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.00533v1 Announce Type: cross Abstract: Large Language Models have achieved substantial progress in reasoning capabilities. Yet in low-resource native settings, many suffer from cross-lingual collapse, reverting to English during intermediate steps that require complex logical reasoning. T...
275. Robust Watermarks Meet Backdoored Models: Evading Diffusion Semantic Watermarks via Stealthy Backdoor ​
Author: Jinyuan Liu, Tianshuo Cong, Pei Li, Tianrui Wang, Xinlei He, Anyu Wang, Xiaoyun Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CV
arXiv:2608.00543v1 Announce Type: cross Abstract: Although semantic watermarking is considered a promising safeguard for images generated by Latent Diffusion Models (LDMs), the reliance of the watermark detection pipeline on neural networks introduces a critical yet underexplored backdoor attack sur...
276. A Context-Aware Cultural Heritage Guide Powered by LLMs ​
Author: Liliana Ardissono, Fabio Ferrero, Angelo Geninatti Cossatin, Claudio Mattutino, Noemi Mauro
Published: 8/4/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.HC
arXiv:2608.00549v1 Announce Type: cross Abstract: We present an extension of Triangolazioni (a Cultural Heritage webapp) to enrich curated content with context-dependent, external information provided by Large Language Models (LLMs) within a loosely-coupled architecture agnostic to the LLM. The syst...
277. DexMani: Human-Derived Manipulability Guidance for Dexterous Rotation ​
Author: Xiaoyang Chen, Shengcheng Luo, Haoran Guo, Jiaming Jiang, Wanlin Li, Ziyuan Jiao, Chenxi Xiao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.00554v1 Announce Type: cross Abstract: Dexterous object rotation is a sequential contact problem: each support, release, and re-contact decision must both produce the desired object motion, and prepare the hand configuration for continued rotation. Existing reinforcement learning methods ...
278. AiFlow: Token-Native Reactive Orchestration with Bounded Backpressure for Streaming LLM Applications ​
Author: Qunhui Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.00558v1 Announce Type: cross Abstract: Large language model (LLM) applications increasingly operate as streaming workflows combining retrieval, tool calls, safety filters, and multi-agent coordination. Although contemporary frameworks expose provider deltas, workflow nodes often treat gen...
279. Crushing the Evidence: A Dual-Penalty Evasion Framework for Fooling White-Box Explainable AI Auditors ​
Author: Niraj Kumar, Harsh Kasyap
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.00566v1 Announce Type: cross Abstract: Post-hoc model explainers such as LIME, SHAP, and Integrated Gradients are widely deployed to audit models in high-stakes sensitive domains, including finance, healthcare, and social welfare. This ensures the model's transparency and acceptability. H...
280. Fairness Auditing: Lower Bounds on Company Manipulation ​
Author: Rachit Verma, Padala Manisha, Sujit Gujar
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.00568v1 Announce Type: cross Abstract: Fairness audits are increasingly mandated in high-stakes applications such as hiring, lending, and automated decision-making. Recent work has established fundamental impossibility results for black-box fairness auditing, showing that sufficiently exp...
281. Latency-Tolerant Cloud-Edge Collaborative Vision-Language-Action Models via Emergent Representational Specialization ​
Author: Daojie Peng, Fulong Ma, Bingtao Wang, Sheng Wang, Jun Ma
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.SY, eess.SY
arXiv:2608.00569v1 Announce Type: cross Abstract: Deploying billion-parameter Vision-Language-Action (VLA) policies on mobile robots creates a systems conflict: semantic reasoning benefits from cloud GPUs, whereas closed-loop control must respond locally despite network delay and jitter. Existing hi...
282. Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts ​
Author: Ziang Wu, Peng Jin, Qishen Yin, Munan Ning, Hao Li, Peizhen Zhang, Li Yuan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00574v1 Announce Type: cross Abstract: Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-Aux balances only ...
283. UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations ​
Author: Ziyue Kang, Nan Nan, Chenhao Lin, Xiaohong Guan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.LG
arXiv:2608.00576v1 Announce Type: cross Abstract: High-polyphony symbolic music is increasingly used in generation, analysis, and arrangement, yet many downstream tasks require bounded representations with fixed tracks or slots. Converting richly orchestrated scores into compact forms is therefore n...
284. Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification ​
Author: Bogdan Savelyev
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00581v1 Announce Type: cross Abstract: Off-the-shelf LID and letter heuristics over-label Kazakh-Russian social text as mixed: Russian loanwords inside Kazakh look like code-switching under a shared Cyrillic script. We release a document-level gold LID set whose guideline keeps integrated...
285. A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense ​
Author: Shikhar Shiromani, Leo Richter
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG
arXiv:2608.00583v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is meant to catch the reward hacks that look clean in the actions and betray themselves only in the reasoning. We show that this is exactly where an adversary who controls the reasoning can defeat it. Rewriting only ...
286. Element-Aware Group Learning for E-Commerce Image Generation ​
Author: Jingtong Chen, Jiahui Wang, Xue Zhao, ShaoGuo Liu, Minghao Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.00584v1 Announce Type: cross Abstract: Recent advances in image generation and editing have made prompt quality a key bottleneck for e-commerce creatives. Vision-language models (VLMs) can generate image-editing prompts from product images and metadata, but further improving their prompt-...
287. An Information Theoretic Treatment of Yager's Probability Distribution Negation ​
Author: Roberto Bruno, Ugo Vaccaro
Published: 8/4/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, math.IT, math.PR
arXiv:2608.00594v1 Announce Type: cross Abstract: In the seminal paper (Yager 2015), Yager defined the negation of a probability distribution $\mathbf{p}=(p_1,\dots,p_n)$, as the distribution $\overline{\mathbf{p}} = (\overline{p}_1,\dots,\overline{p}_n)$, where $\overline{p}_i = ({1-p_i})/({n-1}),$...
288. Learning-Based Motion Planning for Dynamic Environments: From Foundational Algorithms to Emerging Paradigms ​
Author: Zongyuan Shen, Shalabh Gupta, Shancheng Zhao, Dehua Zhou, Gao Wang, Rui Cheng, Yaming Ou, Zhongqiang Ren, Yikui Zhai, C. L. Philip Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.00625v1 Announce Type: cross Abstract: Motion planning in dynamic environments is a fundamental problem in robotics, aiming to generate safe and efficient paths, trajectories, or control actions in the presence of moving obstacles, uncertain predictions, and multi-agent interactions. It h...
289. Understanding Online Failure Prediction in Linux Through Complementary Multi-View Explainability ​
Author: Diogo D'oria, Jo~ao R. Campos
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.00651v1 Announce Type: cross Abstract: Accurate Online Failure Prediction (OFP) has been shown to be feasible in Operating Systems (OSs) settings, but prediction alone is not sufficient for practical adoption. Without diagnostic insight, operators have limited basis to trust alerts or dec...
290. From Chasing Ghosts to Missed Attacks: Perspectives and Perceptions of SOC Practitioners on LLM Integration, Risks, and Readiness ​
Author: Jonas Thurner, Nadine Jost, Stefan Albert Horstmann, Fabian Ising, Lea Groeber, Alena Naiakshina, Sebastian Schinzel
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.HC
arXiv:2608.00672v1 Announce Type: cross Abstract: Security Operations Centers (SOCs) process large volumes of security events, requiring analysts to accurately detect and assess ongoing cyberattacks under time pressure. Recent advances in Large Language Models (LLMs) suggest potential benefits for s...
291. Supporting Cybersecurity Risk Management for Medical Devices via the SECUMAN Ontology and Shapes ​
Author: Martin Diller, Anne Esslinger, Piotr Gorczyca, Evi Hartig, Lia Kacholdt, Hannes Strass
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.00698v1 Announce Type: cross Abstract: We propose the SECUMAN ontology and shapes for representing and analysing cybersecurity risk-management documentation for medical devices. Cybersecurity risks are increasingly relevant for connected medical devices and may have direct consequences fo...
292. Coverage-Driven Adaptive Keyframe Selection for Video Understanding ​
Author: Junyang Zhang, Puhan Luo, Chen Tang, Yuxi Shi, Xiang-Yang Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.00714v1 Announce Type: cross Abstract: Recent advances in large vision-language models (LVLMs) have enabled long-video understanding and analysis. However, processing the large number of frames in a video incurs substantial computational overhead. Existing methods reduce LVLM inference co...
293. Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures ​
Author: Faisal Haque Bappy, Tahrim Hossain, Tarannum Shaila Zaman, Raiful Hasan, Kamrul Hasan, Tariqul Islam
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.MA
arXiv:2608.00718v1 Announce Type: cross Abstract: Multi-agent LLM pipelines orchestrate multiple specialized language model agents into structured workflows where intermediate outputs are passed across agents to solve complex tasks. This design introduces a security gap absent in single-agent settin...
294. An Embedded RISC-V Evaluation of Kolmogorov--Arnold Networks in Hard-Constrained Recurrent Physics-Informed Models ​
Author: Enzo Nicolas Spotorno, Josafat Leal Filho
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.PF
arXiv:2608.00737v1 Announce Type: cross Abstract: Hard-constrained recurrent physics-informed networks (HRPINNs) embed known dynamics inside a recurrent numerical integrator and restrict a neural branch to learning only the residual dynamics that the first-principles model does not capture. Kolmogor...
295. EduPluginBench: Executable Assurance for AI-Generated Educational Plugins ​
Author: Nizam Kadir
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.00739v1 Announce Type: cross Abstract: Code-generation models can produce executable components, but compilation and functional tests do not establish compliance with least privilege, telemetry consent, provenance, privileged-write authority, lifecycle constraints, or bounded failure. We ...
296. Multi-tenant Kubernetes Use Cases for AI, Secure Computing and Data Services, and More ​
Author: Jake Watson, Sadaf R Alam, Christopher Woods, Abdelwahab Kawafi, Thomas Green, Ian Johnson, Ellis Pires, Jessica R. Jones, Utz-Uwe Haus
Published: 8/4/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.PF
arXiv:2608.00742v1 Announce Type: cross Abstract: Kubernetes, as a container orchestration engine, has been widely used in cloud-native ecosystems for several years. In supercomputing ecosystems, especially where bare-metal performance for compute and network devices are considered, the adoption is ...
297. When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems ​
Author: Neha Nagaraja, Amisha Bagari, Hayretdin Bahsi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CR, cs.MA
arXiv:2608.00747v2 Announce Type: cross Abstract: Large language models are increasingly integrated into autonomous robotic systems for task planning and control, but this integration exposes them to prompt injection attacks that can lead to unsafe decisions and physical harm. Multi-agent settings i...
298. Me and My Bot: What Users Talk About in AI Companion Communities on Reddit ​
Author: Richard A. Fabes
Published: 8/4/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.00748v1 Announce Type: cross Abstract: AI companion communities on platforms such as Reddit are widely characterized as spaces where users discuss their relationships with AI bots. This study examines whether and how that characterization holds, guided by the Synthetic Resonance framework...
299. CN101 - A Digital Thermodynamic Computer for Generative AI ​
Author: Lars Holdijk, Denis Melanson, Zier Mensch, Brandon Birchall, Vincent Cheung, Nicholas Lehrter, Maxwell Aifer, Samuel Duffield, Jan Ole Ernst, Rajath Salegame, Antonio J. Martinez, Gavin Crooks, Miranda Cheng, Zach Belateche, Marc Bright, Patrick J. Coles, Faris Sbahi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.ET, cs.AI
arXiv:2608.00754v1 Announce Type: cross Abstract: Thermodynamic computing is an emerging hardware paradigm, in which stochastic physical dynamics serve as the direct computational primitive. The recent explosion of generative AI has only sharpened the search for alternative approaches to compute, an...
300. Less Is More: Tuning Configurable Systems with Imperfect Fidelity ​
Author: Yulong Ye, Miqing Li, Tao Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.DB, cs.PF
arXiv:2608.00759v1 Announce Type: cross Abstract: Configuration tuning is essential for optimizing the performance of highly configurable systems, e.g., throughput or runtime, under a given environment. Yet, this is a challenging process as there can be many options to tune, and configuration measur...
301. Anticipatory Digital Twins for Online Head-and-Neck Adaptive Proton Therapy via Foundation-Model Registration ​
Author: Yizhou Wu, Yuheng Li, Xiaofeng Yang, Chih-Wei Chang
Published: 8/4/2026, 4:00:00 AM
Categories: physics.med-ph, cs.AI
arXiv:2608.00831v1 Announce Type: cross Abstract: Head-and-neck (HN) proton therapy is highly sensitive to anatomical change over a 4-to-6-week course, as tumor shrinkage, weight loss, and setup variation can misposition the Bragg peak near critical organs such as the parotids, oral cavity, brainste...
302. Deep Learning CNN and Recurrence Analysis for Alpha Gamma EEG Biomarkers in Fragile X Syndrome ​
Author: Zag ElSayed, Payton Siekierski, Jack Yanchen Liu, Ernest Pedapati
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, physics.data-an, q-bio.NC
arXiv:2608.00835v1 Announce Type: cross Abstract: Fragile X Syndrome (FXS) is a neurodevelopmental disorder caused by reduced expression of fragile X mental retardation protein (FMRP), leading to disrupted synaptic plasticity, cortical hyperexcitability, and impaired network synchronization. Electro...
303. REIMU: Efficient Heterogeneous Hierarchical Reasoning for SSL-Based Speech Deepfake Detection ​
Author: Kwok-Ho Ng, Tingting Song, Bingwen Feng, Peiya Li
Published: 8/4/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD
arXiv:2608.00857v1 Announce Type: cross Abstract: The increasing realism of speech generated by text-to-speech and voice conversion systems poses growing challenges to media integrity and voice authentication. Self-supervised learning (SSL) has substantially advanced speech deepfake detection, where...
304. RefactorAssist: Agentic Refinement for Reliable Code Refactoring ​
Author: Jonathan Cordeiro, Shayan Noei, Ying Zou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.00924v1 Announce Type: cross Abstract: Code refactoring aims to enhance the internal structure of source code without affecting its functional behavior. The recent advancements of Large Language Models (LLMs) have demonstrated potential for automating software engineering tasks, such as c...
305. Neuro-Symbolic Participation Governance for Verifiable AI Agents in Open Digital Twin Ecosystems ​
Author: Juan Li, Wei Cai, Yan Bai
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.MA
arXiv:2608.00937v1 Announce Type: cross Abstract: Autonomous AI agents, increasingly empowered by large language models, are becoming important components of human-machine systems for high-stakes decision support in digital twin ecosystems. However, existing multi-agent systems often lack robust ver...
306. Temperature-driven inversion and nonlinear dynamics in ChatGPT-like AIs ​
Author: Neil F. Johnson, Frank Yingjie Huo, Bella Xinrui Li
Published: 8/4/2026, 4:00:00 AM
Categories: physics.soc-ph, cond-mat.dis-nn, cs.AI, nlin.AO, physics.app-ph
arXiv:2608.00939v1 Announce Type: cross Abstract: Increasing the temperature of an ordinary many-state system increases access to a wider range of states and hence increases its entropy. We find the opposite in ChatGPT-like AIs, even though raising the decoder temperature likewise increases access t...
307. GraRe: Grasp Candidate Re-Ranking for Frozen 6-DoF Grasp Detectors ​
Author: Jibao Yuan, Yuhui Zhao, Yinzhen Lv, Chao Xu, Shun Li, Chenxi Deng, Shaofei Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.LG
arXiv:2608.00946v1 Announce Type: cross Abstract: Existing 6-DoF grasp detectors typically rank grasp candidates by detector confidence. However, our analysis on GraspNet-1Billion shows that detector confidence is often poorly aligned with grasp quality, causing successful grasp candidates to be ran...
308. The Epistemic Politics of AI Anthropomorphism ​
Author: Donna M Bye, Levin Kuhlmann
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2608.00961v1 Announce Type: cross Abstract: AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustained or relational interaction with AI are routinely pathologised or dismissed as naive, vulnerable to delusion o...
309. An AI Approach to Verified Production Cryptographic Libraries ​
Author: Chuyue Sun, Su Fong, Zhiyi Kuang, Yizheng Jiao, Nina Narodytska, Haoze Wu, David L. Dill, Clark Barrett
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.00965v1 Announce Type: cross Abstract: Cryptographic code is critical infrastructure that must be correct, yet formally verifying production libraries remains difficult. Existing language-model proof systems solve isolated obligations with specifications and premises already given, leavin...
310. One-Sided Quantile Coupling for Flow Matching ​
Author: Jin-Young Kim, So-Yoon Cho, Hyun-Gyoon Kim
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2608.00978v1 Announce Type: cross Abstract: Flow Matching trains continuous-time generative models by regressing the velocity field of a probability path between a simple source distribution and a target data distribution. The coupling that pairs source and target samples strongly affects opti...
311. Beyond Gene Reconstruction: Learning Cell Representations through Complementary Transcriptomic Views ​
Author: Jiaqi Xiong, Yuntao hu, Yu Zheng, Yifei Shi, Xinyue Guo, Jiaxin Qi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.GN
arXiv:2608.00985v1 Announce Type: cross Abstract: The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked expression values. This objective encourages these models to learn gene dependencies but does not direc...
312. Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms ​
Author: Tezan Sahu, Aritra Das, Pankaj Mittal, Sudipta Das
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SE
arXiv:2608.01004v1 Announce Type: cross Abstract: Platform teams hosting agent-extensibility surfaces face a regression-economics paradox: every onboarding customer ships an evaluation set tuned to their domain, but the platform's regression set must live under a hard query-count ceiling bounded by ...
313. Hierarchical Solomonoff Induction: An Unbounded Machine Learning Model ​
Author: Nathan Young
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, math.IT
arXiv:2608.01005v1 Announce Type: cross Abstract: Solomonoff Induction, or SolInd, provides an ideal unbounded model of a priori sequence prediction but cannot naturally describe extrapolation from a given training dataset, as performed by Large Language Models. We apply de Finetti's theorem on exch...
314. MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models ​
Author: Ofir Ben Shoham, Oriel Perets, Nir Grinberg, Nadav Rappoport
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01012v1 Announce Type: cross Abstract: Uncommon and off-guideline cases are difficult for clinical decision support, because physicians must make a series of management decisions under diagnostic uncertainty and rarely see the full case at once. Most large language model (LLM) benchmarks ...
315. Cloud-ScPO: Hidden-State Geometry for Semi-Supervised Preference Optimization in LLM Reasoning ​
Author: Yuzhou Liu, Xiyang Hu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01014v1 Announce Type: cross Abstract: Preference optimization improves mathematical reasoning in large language models (LLMs), but reliable chosen-rejected pairs usually require verified answers, human annotations, or external reward models. We investigate whether preference supervision ...
316. KING: Embodiment-Aware Kinematic Graph Neural Network for Unified Motion Representation of Legged and Wheeled Robots ​
Author: Taku Okawara, Aoki Takanose, Kenji Koide, Shuji Oishi, Masashi Yokozuka
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.01015v1 Announce Type: cross Abstract: Kinematic models provide reliable motion constraints for odometry estimation in featureless environments, where exteroceptive sensing degrades and IMU integration drifts. Learning-based kinematic models can achieve more accurate odometry estimation t...
317. Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy ​
Author: Kaike Ping, Buse \c{C}ar{\i}k, Caleb Wohn, Xiaohan Ding, Tongshuai Wang, Eugenia Rho
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01017v1 Announce Type: cross Abstract: A language model that abandons a correct medical answer under user pushback is more dangerous than one that was simply wrong, because it lends the credibility of a correct answer to the user's misinformation. Such model behavior, described as medical...
318. Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking ​
Author: Timothee Mickus, Claudio Savelli, Eduardo Cal`o, Emilio Raimond, Stella Frank, Hengyu Luo, Flavio Giobergia, Vincent Segonne, Chuyuan Li, Aman Sinha, Lorenzo Vaiani, J"org Tiedemann, Ra'ul V'azquez
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.01021v1 Announce Type: cross Abstract: In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make benchmarking detection independ...
319. Caliber: Cross-Architecture Extraction-Cost Control for Score-Returning APIs ​
Author: Chi Wang, Hanwen Wang, Yu Xia, Zihan Wang, Guangdong Bai
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01023v1 Announce Type: cross Abstract: We present Caliber, an output-perturbation defense against model extraction that formulates noise selection as a calibration problem: how much the defense degrades the supervision signal used to train a surrogate, and the provable per-input query cos...
320. Stress-Relief Annealing: Polynomial-Time Simulation-Free Layout Optimization for Automated Warehouses ​
Author: Xiangjie Luo, Yulun Zhang, Miyuki Koshimura, Makoto Yokoo, Jiaoyang Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.RO
arXiv:2608.01024v1 Announce Type: cross Abstract: We study the problem of optimizing physical layouts for automated warehouses, where hundreds to thousands of robots are coordinated to transport packages. Previous works have shown that optimizing the warehouse layout (e.g., the physical location of ...
321. VLAGuard: A Framework for Evaluating and Mitigating Physical Attention Hijacking in Vision-Language-Action Robots within Wireless Sensor Networks ​
Author: Dongfu Yin, Jinquan Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.01028v1 Announce Type: cross Abstract: Deploying Vision-Language-Action (VLA) robots as mobile edge nodes within wireless sensor networks (WSNs) requires robust protection against physical adversarial threats. We present VLAGuard, a framework to assess and mitigate a critical vulnerabilit...
322. CallScreenBench: Benchmarking On-Device Models as Phone Secretaries ​
Author: Simiao Ren
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.01033v1 Announce Type: cross Abstract: Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting for their user, making on-device task automation newly plausible. One such task is answering the phone. A phone secretary takes an unknown i...
323. WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA ​
Author: Zhihao Zhu, Hanlin Shang, Mingwang Xu, Feipeng Cai, Zhuolin He, Yaoyi Li, Jianhua Han, Hang Xu, Siyu Zhu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.01035v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high computational latency and exposure bias arising from sequential autoregres...
324. What Could the Agent See at 19:05? Generating Temporal Enterprise Scenarios from Real Research and Replaying Them to Evaluate Agents ​
Author: Tezan Sahu, Himani Arora
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.HC, cs.LG
arXiv:2608.01042v1 Announce Type: cross Abstract: Enterprise AI agents act across many apps whose data changes continuously, so an answer is correct only relative to what data existed and who could see it at the moment it was asked. Offline evaluation today grades against a single static snapshot, e...
325. Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks ​
Author: Haoyu Zhang, Xiangchen Guan, Shibo Zheng, Mohammad Zandsalimy, Shanu Sushmita
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.01043v1 Announce Type: cross Abstract: We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR). The opera...
326. DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text ​
Author: Muhammad Yousaf Rehman, Muhammad Islam
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01046v1 Announce Type: cross Abstract: The rapid spread of large language models (LLMs) across the web raises concerns about misinformation, academic integrity, automated content manipulation, and risks to vulnerable online communities. Existing transformer-based detectors, such as GPT-Se...
327. Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception ​
Author: Xinheng Han, Jianfei Wang, Yu Chen, Xiang Wang, Shuai Li, Weixing Li, Feng Pan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01055v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly expected to solve structured perception tasks that require visual recognition, language-to-object binding, object cardinality preservation, and precisely localized grounding and segmentation o...
328. Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization ​
Author: Shigeng Wang, Chao Li, Yangyuxuan Kang, Jiawei Fan, Anbang Yao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01078v1 Announce Type: cross Abstract: We propose ScaleQ-1.58, a scalable ternary post-training quantization (PTQ) framework for reasoning LLMs. Its core insight stems from an empirical finding: although modern LLMs are typically trained to exhibit chain-of-thought reasoning capabilities,...
329. FL-OA: A Byzantine-Robust Federated Learning Framework with Outsourced Auditing for Intelligent Devices ​
Author: Hongliang Zhang, Zhongyuan Yu, Fenghua Xu, Teng Hu, Jian Meng, Jiguo Yu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01095v1 Announce Type: cross Abstract: Federated learning (FL) enables multiple intelligent devices to collaboratively train a high-accuracy model without sharing raw data. However, due to its distributed nature, FL is vulnerable to Byzantine attacks. Existing defense methods rely on stro...
330. SG-Layout: Structured Scene Graph-Guided Layout Generation with LLMs ​
Author: Junsheng Wang, Chao Chen, Mengying Xie, Mingyan Li, Fuqiang Gu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01106v1 Announce Type: cross Abstract: Understanding and generating spatially coherent layouts from natural language remains a fundamental yet challenging task for large language models (LLMs). Existing LLMs often struggle to capture explicit geometric relationships and structural depende...
331. Policy Optimality Measurement for Multi-Vehicle Decision-Making: From Extrinsic Indicators to Intrinsic Quality ​
Author: Ye Han, Lijun Zhang, Dejian Meng
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01133v1 Announce Type: cross Abstract: Evaluating Multi-Agent Reinforcement Learning (MARL) policies in autonomous driving fundamentally relies on extrinsic statistical indicators (e.g., reward curves and success rates), which often mask intrinsic policy degradation and algorithmic blind ...
332. Hybrid Lagrangian-Eulerian Model for Lagrangian Fluid Simulation ​
Author: Ruoyan Li, Wei Wang, Yizhou Sun
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CE, cs.AI
arXiv:2608.01164v1 Announce Type: cross Abstract: Pure Lagrangian neural simulators offer geometric flexibility and exact advection, making them well-suited for modeling moving domains and free surfaces. However, the absence of a fixed global reference frame introduces two severe limitations: a spat...
333. DynActiveGS: Active Gaussian Splatting for Dynamic Scene Reconstruction ​
Author: Hongbo Duan, Pengting Luo, Chengzhi Zhao, Yuanhao Chiang, Fangming Liu, Xueqian Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01178v1 Announce Type: cross Abstract: We present DynActiveGS, a dynamic-aware active reconstruction framework based on 3D Gaussian Splatting (3DGS) for autonomous exploration in dynamic environments. The framework incrementally reconstructs a 3D Gaussian scene representation while suppre...
334. Talking to Digital Twins: Selective Disclosure and Belief Measurement in Financial Social Media ​
Author: Boone Bowles, Raymond Duch, Sorin Sorescu
Published: 8/4/2026, 4:00:00 AM
Categories: econ.GN, cs.AI, q-fin.EC
arXiv:2608.01181v1 Announce Type: cross Abstract: Social media affect financial markets, but public posts by financial media personas are voluntary disclosures. What is not disclosed is therefore usually unobserved. We address this measurement problem by conducting repeated, real-time interviews of ...
335. ReBRAC-v2: The Return of the King ​
Author: Denis Tarasov, Robert K. Katzschmann
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01205v1 Announce Type: cross Abstract: Recent offline reinforcement learning methods increasingly rely on expressive generative policies and specialized value-guidance mechanisms. We ask whether comparable progress can instead come from systematically modernizing a conventional behavior-r...
336. It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling ​
Author: Puzhuo Zheng, Hasan Kurban
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01207v1 Announce Type: cross Abstract: Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe transfers poorly to vision-language models (VLMs): recent work shows that simple majority voting beats selection...
337. AdaHAT: Adaptive Hard Attention to the Task in Task-Incremental Learning ​
Author: Pengxiang Wang, Hongbo Bo, Jun Hong, Weiru Liu, Kedian Mu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01252v1 Announce Type: cross Abstract: Catastrophic forgetting is a major problem in task-incremental learning, where neural networks tend to overwrite previously learned knowledge when trained on new tasks. A number of architecture-based approaches have been proposed to address this prob...
338. Auditing Semantic Gains in Sequential Recommendation: A Lightweight Recovery Test ​
Author: Kong Wang, Zhongke He, Xiang Chen, Hongwei Zeng, Kai Deng, Long Wang, Kehua Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.01260v1 Announce Type: cross Abstract: Recent semantic and generative-retrieval recommenders report substantial improvements over ID-only sequential baselines, but it remains unclear whether these gains arise from language-model reasoning, semantic-ID generation, end-to-end semantic archi...
339. ACE-GraphRAG: Agentic Context Engineering for Hierarchical GraphRAG ​
Author: Yongfeng Huang, Yuren Lai, Ruiying Chen, Haoyu Huang, Mingming Zhao, James Cheng
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01269v2 Announce Type: cross Abstract: Hierarchical Graph Retrieval-Augmented Generation (GraphRAG) organizes corpus knowledge at multiple levels of granularity, yet fixed context construction may fail to translate these multi-resolution representations into a context suited to the curren...
340. Rethinking Video Token Compression with a Global Codebook: Learning Once, Compressing Everywhere ​
Author: Jiayang He, Tianling Xu, Diancheng Kang, Huaide Jiang, Junyan Bai, Shaoming Zheng, Xuan Song
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01271v1 Announce Type: cross Abstract: Video large language models (Video-LLMs) represent videos as dense sequences of visual tokens, whose length grows with the temporal and spatial extent of the input. These tokens often contain substantial redundancy arising from repeated visual patter...
341. Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design ​
Author: Sen Song
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01283v1 Announce Type: cross Abstract: All Transformer-based large language models compute attention via the Euclidean inner product, an architectural choice that Dong et al. (2021) proved causes representational rank to decay doubly exponentially with depth in pure self-attention stacks....
342. Training nGPT ​
Author: Ilya Loshchilov, Boris Ginsburg
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01284v1 Announce Type: cross Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere. In this paper, we describe a practical training recipe for nGPT and evaluate it ...
343. Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents ​
Author: Yidan Lin, Kaixiang Wang, Jiong Lou, Jie Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01285v1 Announce Type: cross Abstract: The continued development of LLMs toward persistent and adaptive intelligence increasingly requires long-term memory mechanisms that preserve and reuse information across interactions. Existing memory systems either compress and structure histories f...
344. UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction ​
Author: Junno Yun, Ya\c{s}ar Utku Al\c{c}alar, Mehmet Ak\c{c}akaya
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.01298v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks. DiTs comprise isotropic transformer blocks, and learn representations progressively across depth, ...
345. Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Metric for Infrared-Visible Fusion Assessment ​
Author: Haoran Liu, Mingzhe Liu, Peng Li, Guibin Zan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01301v2 Announce Type: cross Abstract: Infrared-visible image fusion (IVIF) has no ideal fused reference, so fusion algorithms are routinely ranked by scalar objective metrics that formalize different proxies for information transfer, structure, or source similarity. These proxies often d...
346. Sheaf-theoretic Signal Processing on Graphs: Spectral Theory, Filtering, and Sampling ​
Author: Gabriele D'Acunto, Leonardo Di Nino, Paolo Di Lorenzo, Sergio Barbarossa
Published: 8/4/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG
arXiv:2608.01318v1 Announce Type: cross Abstract: Modern sensing, communication, and learning systems generate heterogeneous network signals, with local data differing in dimension, modality, and geometric structure. Processing such data requires a mathematical framework capable of simultaneously mo...
347. Dense Language Generation Made Simple: Deterministic, Randomized, and Multi-Order Algorithms ​
Author: Ziyi Cai, Shuangping Li, Yiheng Shen, Kangning Wang, Peng Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.DS, cs.AI, cs.CL, cs.DM, cs.LG
arXiv:2608.01320v1 Announce Type: cross Abstract: Language generation in the limit is a theoretical framework for studying how a generator can learn to produce new valid strings from a stream of positive examples. In this model, an adversary chooses an unknown language from a countable family and en...
348. Context Compaction Theory ​
Author: Hayder Tirmazi, Sam Markelon, Allison Bishop, Michael Mitzenmacher
Published: 8/4/2026, 4:00:00 AM
Categories: cs.DS, cs.AI
arXiv:2608.01326v1 Announce Type: cross Abstract: Large Language Models (LLMs) have a bounded context window. The context window is the maximum input size an LLM can consume for a single inference. AI agents rely on a process called context compaction to fit their state within the context window whe...
349. LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning ​
Author: Ziyan Xiao, Yinghao Zhu, Wenting Zhang, Heaju Kim, Lequan Yu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV
arXiv:2608.01328v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart understanding and multi-step inference. These abilities are increasingly important as MLLMs are adop...
350. DeVIT: Low-Power Vision Transformer Acceleration Using Delta Computation ​
Author: Reyhaneh Hosseinzadeh, Parham Zilouchian Moghaddam, Mehdi Modarressi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.AR
arXiv:2608.01343v1 Announce Type: cross Abstract: The emergence of transformer-based deep learning models has brought unprecedented performance across various domains, particularly in natural language processing and computer vision. However, deploying these models, especially on resource-constrained...
351. Spatiotemporal Proximal Causal Inference under Hidden Confounding and Interference ​
Author: Omar Faruque, Pavan Raj Ravi, Jianwu Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01352v1 Announce Type: cross Abstract: Estimating causal effects from real-world spatiotemporal data is challenging due to hidden confounders and interference. Standard causal identification methods assume conditional exchangeability given observed covariates, which fails whenever hidden ...
352. Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution ​
Author: B. Sankar, Pawni Yadav, Srinidhi Ranjini Girish, Amogh A. S
Published: 8/4/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2608.01366v2 Announce Type: cross Abstract: Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts. Iterative multi-turn prompting often leads to context degradation and diminishing cognitive returns. We present ...
353. Imprecise Belief Fusion Improves Multi-agent Social Learning ​
Author: Zixuan Liu, Jonathan Lawry, Michael Crosscombe
Published: 8/4/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, physics.soc-ph
arXiv:2608.01367v1 Announce Type: cross Abstract: In social learning, agents learn not only from direct evidence but also through interactions with their peers. We investigate the role of imprecision in such interactions and ask whether it can improve the effectiveness of the collective learning pro...
354. Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent Safety ​
Author: Ruiyang Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.01388v1 Announce Type: cross Abstract: Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasingly deployed to intercept unsafe tool-call sequences in LLM agents. Yet the same monitor achieves 68-75% attack coverage on some model architectures a...
355. Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+ ​
Author: Sherzod Hakimov, Karl Osswald, Jelle Psurek, Eszter Bukovszky, A. Altar L"user, David Schlangen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01395v1 Announce Type: cross Abstract: We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unlike static or preference-based evaluation, this paradigm is multi-turn...
356. TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction ​
Author: Rasa Hosseinzadeh, Alex Labach, Zexin Xue, Shuyi Han, Valentin Thomas, Anthony L. Caterini
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01400v1 Announce Type: cross Abstract: Tabular foundation models, driven by in-context learning, have rapidly grown in quality and popularity. However, recent approaches with either cell-based architectures or retrieval have sacrificed efficiency for raw performance, restricting their uti...
357. Demystifying When and Why VLAs Fail in Contact-Rich Tasks and How to Fix Them ​
Author: Carlota Par'es-Morlans, Nils Kuhn, Isabel Liu, Alberta Longhini, Jeannette Bohg
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.01402v1 Announce Type: cross Abstract: We address the problem of understanding when and why Vision-Language-Action models struggle with contact-rich manipulation tasks that require precise physical interaction. Prior work has primarily focused on addressing contact failures through force-...
358. When Replanning Becomes the Bottleneck: Budgeted Replanning for Embodied Agents ​
Author: Shuaijun Liu, Feiyang You, Xingwei Chen, Ningxin Su
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2608.01428v1 Announce Type: cross Abstract: Embodied agents replan frequently to recover from execution drift, partial observability, and coordination hazards, but each LLM-based replanning call can consume an accumulated textual context that grows over time and across agents. Once this contex...
359. Statistical Mechanics of Learning on Product Wasserstein Manifolds ​
Author: Srinivasa Rao P Vangmayi P Reddy
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.NC
arXiv:2608.01434v1 Announce Type: cross Abstract: Normally the statistical mechanics of learning treats constraints on weight distributions as restrictions that shrink the space of possible solutions. Therefore, it reduces model capacity. In this paper we would like to take a contrary approach, whic...
360. Same violence, different answer: how AI responds to coercive control against women across languages ​
Author: Lyu Chang, S`onia Estrad'e Albiol, N'uria Verg'es Bosch
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL, cs.HC
arXiv:2608.01436v1 Announce Type: cross Abstract: Women experiencing coercive control, a form of intimate partner violence increasingly conducted through digital devices, are turning to conversational AI for help, and the protection they receive should not depend on the language they write in. We an...
361. Retrieval Augmented Biomedical Question Answering with Weak Question Recovery and Neural Reranking for BioASQ Task 14b ​
Author: Xueying Zhao, Lee Mai, Balaji Anandganesh
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01468v1 Announce Type: cross Abstract: This work presents DS@GT ARC BioASQ team's work for a biomedical question answering pipeline, integrating multi-source query expansion, neural reranking, retrieval refinement, and OpenBioLLM-assisted answer generation. The system combines PubMed retr...
362. VGER: Voxel-Guided Global Event Ranking for Event Cloud Attribution ​
Author: Youxin Jiang, Baoheng Fu, Hongwei Ren, Xiangqian Wu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01470v1 Announce Type: cross Abstract: Event cameras produce sparse and asynchronous event streams that provide rich spatio-temporal information for efficient perception. Recent advances in event-based models have demonstrated strong performance by directly modeling asynchronous events wi...
363. Rapid Embodiment Adaptation for Quadrupedal Locomotion ​
Author: Dichen Li, Bo Ai, Nico Bohlinger, Jan Peters, Hao Su, Henrik I. Christensen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2608.01506v1 Announce Type: cross Abstract: Humans readily adapt their movements as their bodies change through aging, injury, or load carrying, but learning-based robot policies often break when hardware properties shift. We introduce an online embodiment adaptation framework for quadrupedal ...
364. Deep Agentic Search for Repository-Level Code Question Answering: An Empirical Study ​
Author: Amirkia Rafiei Oskooei, Bora Ilci, Alperen Kayim, Mehmet Egemen Uzun, Berat Can, Kaan Emre Kara, Ozan Orhan, Mehmet S. Aktas
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.IR, cs.MA
arXiv:2608.01507v1 Announce Type: cross Abstract: Code agents spend much of their effort simply locating the right code inside a repository. Two approaches dominate current practice. In Semantic Search, the agent retrieves code blocks from a vector index built from the repository in advance. In Deep...
365. Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics ​
Author: Longtian Bao, Jianyou Wang, Yang Zhang, Youze Zheng, Ramamohan Paturi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.01522v1 Announce Type: cross Abstract: Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning traces are usually unavailable, and models often exhibit an apparent ceiling beyond which additional ...
366. Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning ​
Author: Seongyoon Kim, Boryeong Cho, Jihwan Oh, Seokhyun Chung, Se-Young Yun
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01556v1 Announce Type: cross Abstract: Large language models are increasingly aligned to human preferences via reward modeling, but user preference data are sensitive and often cannot be centralized. Federated learning keeps such data local while learning a shared initial reward model, wh...
367. Enhancing Visual Perception in Foggy Conditions via Multiclass Fog Density Modeling ​
Author: Mohamad Mofeed Chaar, Galia Weidl
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01572v1 Announce Type: cross Abstract: Autonomous driving (AD) systems have advanced rapidly over the past decade; however, robust perception under adverse weather conditions remains a major challenge, particularly in dense fog. In this work, we investigate fog-aware perception using synt...
368. Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard ​
Author: Hector Zenil, Luan Ozelim
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01575v1 Announce Type: cross Abstract: Whether large language models perform genuine algorithmic reasoning or mere pattern completion is hard to test, because most benchmarks lack a ground truth for correct inductive inference. We introduce F-ICL, an in-context-learning benchmark that sup...
369. The Label Defines the Timescale: Trait-State Limits of Temporal-Aggregate Learning ​
Author: Xizhe Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2608.01587v1 Announce Type: cross Abstract: Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows. Their apparent performance ceiling may therefore be an acquisition-protocol ceiling rather than a model-cap...
370. HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning ​
Author: Haowei Liu, Jiamian Wang, Hsin-Tai Wu, Zhiqiang Tao, Yi Fang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR
arXiv:2608.01597v1 Announce Type: cross Abstract: Search-augmented LM agents are typically trained with a binary exact-match reward, which throws away most of what a failed trajectory tells us about why it failed. We introduce HindSearch, a hindsight self-distillation procedure for GRPO: after each ...
371. PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters' Lack of Knowledge ​
Author: Eojin Jeon, SangKeun Lee
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01598v1 Announce Type: cross Abstract: Simulating human-like Theory of Mind (ToM) has been a longstanding problem in natural language processing (NLP). To address this, existing works introduce a reasoning step of event hiding (a.k.a. perspective-taking), where events unknown to a charact...
372. Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge ​
Author: Ashfak Yeafi, Mehedi Hasan, Md Khairul Islam
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01614v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) face a critical computational bottleneck when processing high-resolution imagery due to the $O(N^2)$ memory complexity of Softmax Multi-Head Attention (MHA). While substituting MHA with independent Multi-Layer Perceptron...
373. QWRF-Net: A Quantum-Wavelet Framework with Rectified Flow for Short-Term Precipitation Nowcasting ​
Author: Zhuo Wang, Chaorong Li, Wenjie Luo, Chuanhu Deng
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01626v1 Announce Type: cross Abstract: Short-term precipitation nowcasting is important for hydrometeorological early warning, especially when intense convective rainfall may trigger urban flooding, flash floods, and other high-impact hazards. A key challenge in warning-oriented nowcastin...
374. RING: Retrieval-Internalized Generation for Continual Large-Scale Knowledge Injection ​
Author: Shicheng Xu, Liang Pang, Liyi Chen, Zihao Wei, Jingcheng Deng, Yan Gao, Yi Wu, Yao Hu, Huawei Shen, Xueqi Cheng
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01630v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves factuality but adds latency and engineering overhead at serving time. We propose RING (Retrieval-Internalized Generation), a holistic paradigm spanning both architecture and training that injects large-sc...
375. AI-assisted Script Management for Requirements Elicitation Interviews ​
Author: Anmol Singhal, Paulo Carvalho, Travis Breaux
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.01640v1 Announce Type: cross Abstract: Requirements elicitation interviews require interviewers to balance topic coverage, active listening, and adaptive probing while responding to stakeholders in real time. Although prior work has explored AI support for isolated interviewing tasks, suc...
376. StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring ​
Author: Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Kaixing Yang, Steven Hoi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.GR
arXiv:2608.01643v1 Announce Type: cross Abstract: Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streaming methods are open-loop: each clip depends on past context, but the model cannot check or correct its trajectory. Small errors therefore ac...
377. CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models ​
Author: Yu Chen, Xiaohong Li, Xiaole Wang, Jianjin Zhang, Jun Sun, Yafeng Deng
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01644v1 Announce Type: cross Abstract: In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply. Such visual sequences are highly redundant along the spatio-temporal d...
378. SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction ​
Author: Shen You, Xiaoming Zhu, Weining Weng, Hefei Mei, Weixuan Wang, Zhongshen Li, Zeji LI, Ye-Wen Wang, Zijun Liao, Juchao Zhuo, Yang Wei, Fuhao Qiu, Siqin Li, Zhenjie Lian, Danei Gong, Junkai Ji, Xiangtao Li, Qiuzhen Lin, Liang Wang, Ka-Chun Wong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.01652v1 Announce Type: cross Abstract: LLM-based multi-agent coordination faces a fundamental trade-off between efficiency and adaptivity in dynamic environments. Existing approaches typically rely on repeated LLM invocations or multi-round communication to adapt decisions during executio...
379. Few-Shot Concept Prompt Learning for Segmentation Foundation Models via Visual Grounding ​
Author: Rahul Venkataramani, Rachana Sathish
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01663v1 Announce Type: cross Abstract: Promptable segmentation foundation models (FMs) such as SAM3 and Medical SAM3 promise few-shot, interactively-specified segmentation for medical imaging through a natural language interface, yet their performance on clinical tasks falls well short of...
380. Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation ​
Author: Fengxian Ji, Yuke Li, Jingpu Yang, Juanfan Wu, Fan Zhang, Zhexuan Cui, Yu Xie, Min Peng, Qianqian Xie, Xiuying Chen, Zhuohan Xie
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.01666v2 Announce Type: cross Abstract: However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unified three-component benchmark f...
381. Learning What to Remember: Test-Time Training via Context Distillation ​
Author: Zixuan Wang, Xingyu Dang, Rui-Jie Zhu, Zixin Wen, Hengyu Fu, Wenhao Chai, Jason D. Lee
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.01672v1 Announce Type: cross Abstract: Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test-time training (TTT) is an appealing approach that performs online parameter updates for long-cont...
382. ProtoAct: Turning Wet-Lab Protocols into Embodied Robotic Actions ​
Author: Zhe Liu, Jiaming Gu, Zhaohui Du, Zhe Wang, Huanbo Jin, Quan Lu, Qi Wang, Ting Xiao, Minting Pan, Dongzhan Zhou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.01690v1 Announce Type: cross Abstract: Biological wet-lab protocols are written for trained researchers and often leave routine operations, state-dependent conditions, and contextual parameters implicit, making them difficult to translate into robot-executable actions. We present ProtoAct...
383. ARM: Detector-Agnostic Changepoint Attribution with Finite-Sample Error Control ​
Author: Chenchen Peng, Mixia Wu, Qijing Yan, Da Chen, Zhiqi Shen
Published: 8/4/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, cs.CV
arXiv:2608.01691v1 Announce Type: cross Abstract: Detecting a change in a multivariate series answers only the first of two questions; the operational question is which coordinates changed. Existing answers are incomplete. Block-level procedures certify predefined groups of coordinates under an addi...
384. Entity-Aware Sequence Transduction for Player-Centric Ball Action Spotting ​
Author: Ruifeng Wang, Di Yang, Jiangtao Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01696v1 Announce Type: cross Abstract: Player-centric ball action spotting requires temporally precise event detection together with actor attribution in crowded, partially observed multi-agent sports videos. Existing Denoising Sequence Transduction (DST) baselines treat the player-role d...
385. Rethinking Generative AI Literacy: An Integrative, Developmental, and Dialectical Framework for K-12 Teacher Education ​
Author: Shahin Hossain, Sima Ahmadi, Leqi Li, Idowu David Awoyemi, Wei Huang, Chenxi Zhou, Jujia Li, Samaa Haniya, Shapla Khanam, Tasbirun Mashreka Subaha
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.ET, cs.HC
arXiv:2608.01705v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has entered classrooms faster than teachers have been prepared to use it well, producing a GenAI literacy lag in which technological diffusion outpaces educators' conceptual, pedagogical, and ethical readine...
386. Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch ​
Author: Shuyang Xie, Shuxiao Xie, Feng Zhu, Yanli Ji, Wangmeng Zuo
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.01715v1 Announce Type: cross Abstract: Online-judge verdicts and the datasets and benchmarks built on them are treated as ground truth for evaluating and training large language models for code. Yet prior audits have sounded a warning: official suites accept buggy submissions. These audit...
387. MNC: Scope-Bound Semantic Declassification for Private LLM-Agent Communication ​
Author: Jinghan Xu, Longze Fan, Zeyuan Wang, Xinjin Li, Hankai Liu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.01719v1 Announce Type: cross Abstract: Multi-agent large language model (LLM) systems can expose protected state through internal messages, tool arguments, logs, and persistent memory even when their public outputs appear innocuous. Existing privacy prompts, redaction methods, and source-...
388. When Extreme Darkness Meets Motion Blur: MeanFlow for Unified RAW Restoration ​
Author: Zepu Wang, Jingze Liang, Weijie Xiao, Kexin Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01720v1 Announce Type: cross Abstract: Extremely low-light RAW enhancement aims to recover severely attenuated sensor signals, yet existing methods often focus on illumination and noise while overlooking the motion-induced degradations inherent in practical low-light imaging. We present a...
389. TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics ​
Author: Heechan Lee, Jeonggyu Kang, Junho Myung, Jaywoong Jeong, Juho Kim, Joseph Seering
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2608.01724v1 Announce Type: cross Abstract: Group conversations are fundamental to human collaboration, yet standard large language models (LLMs) still struggle with the complexities of multi-party interaction. This challenge persists in part because existing group conversation datasets are of...
390. X-KGRank: A Knowledge Graph RAG Framework for Explainable Recommendations via Pattern Mining and LLM Re-Ranking ​
Author: Meenakshi Rajpurohit, Jainish Patel
Published: 8/4/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.01732v1 Announce Type: cross Abstract: Modern recommender systems produce predictions that users cannot interrogate. The two dominant improvements, collaborative filtering and LLM-based reasoning, each fall short: collaborative filtering captures behavioural signals but offers no reasonin...
391. Disagree to Accelerate: Closing the Loop on Diffusion Feature Forecasts ​
Author: Yanchao Li, Jiaqing Xie, Ben Gao, Wanhao Liu, Yanbo Wang, T. Y. Tsui, Jinfei Liu, Yuqiang Li, Tianfan Fu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01740v1 Announce Type: cross Abstract: Training-free feature forecasting accelerates diffusion sampling by predicting features at skipped denoising steps. Recent work has mainly focused on designing stronger forecasters. Yet forecast error varies sharply across steps, and open-loop caches...
392. SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models ​
Author: Xingyan Li, Jordan A. Caraballo-Vega, Jie Gong, Mark L. Carroll, Jianwu Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01751v1 Announce Type: cross Abstract: Geospatial foundation models (GeoFMs), pretrained on large-scale geospatial data such as Earth observation (EO), climate, and weather data, have shown promising performance when fine-tuned on diverse downstream tasks. However, there are two challenge...
393. Can Urban Blight Be Accessed with Vision-language Models: A Case Study in Detroit ​
Author: Xiaohao Yang, Aohua Tian, Derek Van Berkel, Xu Qiang, Mark Lindquist
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01753v1 Announce Type: cross Abstract: Addressing urban blight has seen increased focus in the past 15 years. Assessing urban blight is essential for guiding urban planning, targeting rehabilitation, and safeguarding public health, yet traditional residential blight surveys are difficult ...
394. EntailLLM: Verifying LLM-Generated Vulnerability Discovery Paths with Domain Knowledge via Logic Programming ​
Author: Kaustuv Mukherji, Jaikrishna Manojkumar Patil, Colton Payne, Paulo Shakarian, Dana Warmsley, Nigel Stepp, Evelyn Kim
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LO, cs.SE
arXiv:2608.01763v1 Announce Type: cross Abstract: Large language models are increasingly used to reason about software vulnerabilities, but their outputs can silently violate domain knowledge, limiting their reliability in safety-critical settings such as medical devices. Prior work either treats th...
395. Multi-Source Dynamic Graph Learning for Compound-Flood Forecasting in Managed Coastal Systems ​
Author: Liangjun You, Min Wu, Orlando Woods, Dongsheng Luo
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01775v1 Announce Type: cross Abstract: Compound flooding in managed coastal systems is influenced by hydrological conditions and water-management activity observed across multiple monitoring stations. Current forecasting models can capture temporal dependencies with low average errors, bu...
396. Investigating Social Bias in Narrative Image Generation ​
Author: Junyeong Park, Sowon Min, Euna Jang, Soobin Kim, Jiho Jin, Hyunseung Lim, Gahyeon Bae, Hwajung Hong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01780v1 Announce Type: cross Abstract: Text-to-image (T2I) generation models are increasingly embedded in applications such as media content creation and education, raising concerns about how their outputs may reproduce social biases. Prior work has shown that T2I models exhibit social bi...
397. Radar Detection in the CBRS Band: Techniques, Challenges, and Future Directions ​
Author: Madan Baduwal, Priyanka Paudel
Published: 8/4/2026, 4:00:00 AM
Categories: eess.SP, cs.AI
arXiv:2608.01786v1 Announce Type: cross Abstract: The 3.5 GHz Citizens Broadband Radio Service (CBRS) is a shared wireless band that allows both government systems and commercial networks (such as private LTE/5G) to use the same spectrum. To prevent interference with critical government systems, esp...
398. PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation ​
Author: Xiaohan Jiang, Zeyu Li, Wei Zhang, Jiang Xu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.ET, cs.AI
arXiv:2608.01791v1 Announce Type: cross Abstract: The rapid development of photonic integrated circuits (PICs) is shifting the design flow from traditional graphical user interface (GUI)-based methods to script-based methods for higher flexibility, portability, and maintainability. However, script-b...
399. Illuminating Visual Identity in Universal Multimodal Embeddings ​
Author: Jiawei Cao, Junyi Feng, Jiashen Hua, Ziheng Huang, Bing Deng, Kaijie Wu, Chaochen Gu, Jieping Ye
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.01794v1 Announce Type: cross Abstract: Universal Multimodal Embeddings (UMEs) aim to unify various modalities and tasks into a shared representation space. In recent years, this field has witnessed substantial progress driven by the development of Multimodal Large Language Models (MLLMs)....
400. LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation ​
Author: Tankun Li, Zhi Chen, Yaohua Tang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01804v2 Announce Type: cross Abstract: Post-training large language models (LLMs) via reinforcement learning (RL) has significantly advanced code generation capabilities. To bypass the heavy memory footprint of critic networks, current state-of-the-art frameworks leverage critic-free para...
401. Predictive Maintenance: Deep Learning-Based Remaining Useful Life Prediction for Combat Aircraft Engines ​
Author: Fatih "Urgen, Do\u{g}ay Alt{\i}nel
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SY, eess.SY
arXiv:2608.01819v1 Announce Type: cross Abstract: To improve the operational readiness of combat aircraft engines and reduce unplanned maintenance costs, accurately estimating the remaining useful life (RUL) is critical. Traditional maintenance often proves insufficient under dynamic mission profile...
402. PartMat: Material-Aware 3D Part Decomposition with a Single Global Latent ​
Author: Guangming Fu, Jin Song, Yiyun Fei, Guoqiu Li, Ruigao Yang, Jianan Jiang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.GR
arXiv:2608.01825v1 Announce Type: cross Abstract: Part-level 3D generation has recently attracted increasing attention for producing structured and editable 3D assets. However, existing methods typically decompose objects according to functional semantics rather than the editable material boundaries...
403. DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents ​
Author: Huanyao Zhang, Jiepeng Zhou, Runhao Zhao, Yanzhe Shan, Jiaoyang Chen, Bowen Zhou, Bo Li, Fang Wang, Jialong Wu, Zhengwei Tao, Lang Mei, Xiaohan Yu, Liyan Liu, Chong Chen, Wentao Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01827v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this lim...
404. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills ​
Author: Gaytri Jena, Kapil Wanaskar, Vinija Jain, Aman Chadha, Vasu Sharma, Amitava Das
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.01851v1 Announce Type: cross Abstract: Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that a...
405. Beyond Magnitude and Shape: A Direction-Aware Loss for Time Series Forecasting ​
Author: Seunghan Lee, Jaehoon Lee, Jun Seo, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01857v1 Announce Type: cross Abstract: The direction of change --- whether a series will move up or down --- is often as important as its exact value in decisiondriven applications such as risk management and financial forecasting. However, most forecasting losses optimize either point ma...
406. No One Wins in Nuclear War: A Social Simulation of Military Decision-making ​
Author: Glenn Matlin, Isaac Song, Anthony Wen-Ming Zang, Mark Riedl
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL, cs.MA
arXiv:2608.01868v1 Announce Type: cross Abstract: WOPR is a social-simulation environment for studying how organizations make high-stakes decisions, built on a deterministic, replay-validated rules engine and using wargames as the vehicle. We instantiate it first with the published card game Nuclear...
407. LAB-Tab: LLM-Augmented Bayesian Network Adaptation for Few-Shot Tabular Generation ​
Author: Zijian Shen, Taijie Chen, Bin Zhou, Ziyang Jiang, Jintao Ke
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.01879v1 Announce Type: cross Abstract: Tabular data generation supports analysis and decision-making when target-domain data are scarce, yet collecting complete target samples is often costly. A practical but underexplored setting provides only a few target records together with richer so...
408. Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models ​
Author: Yuwen Wang, Tian-Hao Zhang, Minghao Cai, Yilin Ren, Ziyang Jiang, Xin Wang, Zhichao Wang, Pan Zhou, Kun Zhan, Xinyuan Qian
Published: 8/4/2026, 4:00:00 AM
Categories: cs.MM, cs.AI
arXiv:2608.01881v1 Announce Type: cross Abstract: Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rather than answer directly from a fixed audio input. We study such proble...
409. Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling ​
Author: Cunchen Hu, Liangliang Xu, Tian Liu, Min Lyu, Yongkun Li, Sa Wang, Shuo Quan, Yanan Yang, Wenda Tang, Yiduo Wang, Fu Yu, Jie Wu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2608.01891v1 Announce Type: cross Abstract: Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing energy consumption. Existing energy-management approaches adapt GPU freque...
410. TransNRank: Towards Accurate Neoantigen Ranking with Transformer ​
Author: Zhiyin An, Yuenan Hou, Shumeng Duan, Yiming Zhou, Yuanting Zheng, Leming Shi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CE, cs.AI
arXiv:2608.01924v1 Announce Type: cross Abstract: Personalized neoantigen prediction is challenging due to the scarcity of positive samples, the noise of the experimental data, the severe class imbalance trait and the complex of immunogenicity features. Prior arts, such as linear regression and XGBo...
411. Effective and Efficient Context Retrieval via Partial Dependency Graph for Repository-Level Code Generation ​
Author: Zhongxin Liu, Zhonghao Jiang, Zhifan Ye, Haoye Wang, Jiakun Liu, Xiaoxue Ren
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.01927v1 Announce Type: cross Abstract: LLM-based repository-level code generation aims to generate code using the context available in a software repository, requiring LLMs to reason over complex code dependencies. Due to limited context windows and insufficient repository-specific unders...
412. Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection ​
Author: Wenxiao Fan, Jingling Fu, Fang Li, Luohang Liu, Yu He, Lichen Ma, Zhiyang Yu, Weishan Bi, Junshi Huang, Yan Li, Gu Simiu, Kan Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01930v1 Announce Type: cross Abstract: Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insufficient visual attention or contextual inertia, leaving unclear what models reuse instead of recomputing...
413. Automatic Annotation of Ancient Greek Vowel Length ​
Author: Albin Th"orn Cleland, Eric Cullhed
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.01935v1 Announce Type: cross Abstract: Prior work in Ancient Greek NLP relies on corpora that do not disambiguate the phonemic vowel length of alpha, iota, and ypsilon, together known as the dichrona. Depending on lexeme, morphology, sandhi, syntax, and conventions of period, genre, and v...
414. Semantic Networks as Clues: A Theoretical Foundation and Process Optimization for Semantic Network Construction ​
Author: JinWoo Ha, Dongsoo Kim
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SI, cs.AI
arXiv:2608.01936v1 Announce Type: cross Abstract: The subject matter of this paper is twofold. One is to review the theoretical foundation of a specific type of Semantic Networks (SNs) representing textual non-propositional knowledge. The other involves proposing a framework (ClueNetwork) for rankin...
415. Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory ​
Author: Zhaotian Gu, Jie Su, Weiwei Wang, Chang Liu, Tianyi Qian, Dahui Wang
Published: 8/4/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.NE
arXiv:2608.01947v1 Announce Type: cross Abstract: The ability to robustly maintain and update continuous variables is a hallmark of working memory. While classical continuous attractor networks suffer from severe fine-tuning fragility, standard artificial recurrent neural networks (RNNs) like GRUs a...
416. Agentic Self-Healing for Data and AI Pipelines: An Affordable Vendor-Agnostic Architecture using Open-Source Software ​
Author: Solomon Eshun, Dennis Murage, Sharleen Muoki, Chih-Chun Chen, Stephen Adjignon, Matteo Staar, Oliver Ang'elil
Published: 8/4/2026, 4:00:00 AM
Categories: cs.ET, cs.AI, cs.DB
arXiv:2608.01955v1 Announce Type: cross Abstract: Modern organizations rely on data, machine learning, and software delivery pipelines to move data, train models, deploy applications, refresh dashboards, and support business-critical decisions. However, these pipelines often fail because of data qua...
417. FAST-GS: Frequency Aware Space-time Gaussian Splatting for Photorealistic Dynamic Novel View Synthesis ​
Author: Zhengyang Zhang, Ziyu Lu, PengCheng Li, Hongbo Duan, Yi Liu, Pengting Luo, Peiyu Zhuang, Xinghui Li, Shaohua Ma
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01958v1 Announce Type: cross Abstract: 4D Gaussian Splatting (4DGS) excels in dynamic 3D reconstruction and real-time novel view synthesis via efficient 4D Gaussian representations and parallelizable rendering. However, existing 4DGS approaches rely on a single polynomial to model motion,...
418. Music Restoration via Latent Operator Optimization and Diffusion Model Priors ​
Author: Michal \v{S}vento, Eloi Moliner, Valtteri Kallinen, Lauri Juvela, Vesa V"alim"aki, Pavel Rajmic
Published: 8/4/2026, 4:00:00 AM
Categories: eess.AS, cs.AI
arXiv:2608.01972v1 Announce Type: cross Abstract: Music restoration seeks to recover a clean signal from an observed recording degraded by an unknown effect, distortion, or corruption. Existing systems often rely on paired training data and distortion-specific supervision, which limits their use whe...
419. AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning ​
Author: Jingqi Tian, Haoji Zhang, Lin Chen, Hongbo Jin, Haonan Xu, Tianrui Zhu, Xingming Shui, Shilin Ma, Wenjing Yang, Yansong Tang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01980v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether a video multimodal large language model can adapt its reasoning effort to each question. We propos...
420. SPARE: Structural Parameter-Free Affinity Regularization for Flow Matching ​
Author: Zong-Wei Hong, Jinglun Li, Shen Zhang, Yuhan Liu, Linze Li, Yao Tang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01990v1 Announce Type: cross Abstract: Denoising diffusion transformers achieve strong generation quality but converge slowly during training. Regularizing their internal representations has emerged as an effective accelerator, yet existing methods split into two families with complementa...
421. TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous Driving ​
Author: Mengmeng Zhu, Yuxuan Sun, Wei Chen, Bo Ai
Published: 8/4/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.NI
arXiv:2608.01998v1 Announce Type: cross Abstract: The deployment of Vision-Language Models (VLMs) in autonomous driving (AD) systems is constrained by on-board computing power, restricting vehicles to small VLMs (SVLMs) with limited perception and reasoning capabilities. Infrastructure-assisted AD a...
422. MANGO-Grasp: Mahalanobis Fields over Geometry-Oriented 3D Gaussians for Cross-Embodiment Dexterous Grasping ​
Author: Heng Zhang, Kevin Yuchen Ma, Mike Zheng Shou, Weisi Lin, Yan Wu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.02014v1 Announce Type: cross Abstract: Cross-embodiment dexterous grasping aims to synthesize stable grasps across heterogeneous multi-fingered hands with little or no embodiment-specific tuning. Existing interaction-centric methods achieve promising results, but their object representati...
423. CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship ​
Author: Yao Liu, Guangjia Chai, Yuming Huang, Jihao Huang, Lei Wang, Junchen Wan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.02046v1 Announce Type: cross Abstract: LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family ...
424. TextNCA: Neural Cellular Automata for Language Modeling via Hierarchical Local Attention ​
Author: Avni Mittal, Avinash Anand, Ashutosh Kumar, Dikshant Kukreja, Kritarth Prasad, Sushane Dulloo, Erik Cambria, Timothy Liu, Zhengkui Wang, Rajiv Ratn Shah
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.02050v1 Announce Type: cross Abstract: Can a strictly local, iterated, weight-shared computation primitive support language modelling, and which of those three properties actually drives the model's behaviour? We define \textsc{TextNCA}, a 1D causal windowed-attention realisation of the N...
425. TBSG-Net: Temporal Bipartite Scene Graph Network for Fine-Grained Video Moment Retrieval ​
Author: Ji Huang, Yongsheng Dai, Tianyu Ren, Barry Devereux, Hui Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02056v1 Announce Type: cross Abstract: Recent advances in proposal-free Video Moment Retrieval (VMR) have highlighted the effectiveness of Static Scene Graphs (SSGs). By modeling objects and their relations at the frame level, SSGs enrich retrieval-oriented video representations. However,...
426. Geometry-Guided Layerwise FFN Width Allocation in Transformers ​
Author: Timur Mudarisov, Mikhail Burtsev, Radu State
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.02064v1 Announce Type: cross Abstract: Feed-forward networks (FFNs) account for a large fraction of Transformer parameters, yet their hidden width is usually constant across depth. We ask whether this capacity can instead be allocated from a forward-pass measurement of layer behavior. We ...
427. An AI-Based Decision-Support Pipeline for Day-Ahead Photovoltaic Forecasting ​
Author: Fariba Dehghan, Sebastian Stein, Vahid Yazdanpanah, Stephanie Gauthier, Masood Nazari
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SY, eess.SY
arXiv:2608.02088v1 Announce Type: cross Abstract: Reliable photovoltaic (PV) forecasts are needed for low-carbon energy systems, but newly deployed sites often have short, imperfect records. This makes standard day-ahead forecasting difficult: persistence and physical baselines can be sensitive to c...
428. How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models ​
Author: Andres Algaba, Francesca Carlon, Lynn Delcon, Marthe Ballon, Bert Verbruggen, Vincent Ginis
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02089v1 Announce Type: cross Abstract: Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds each completed run fixed and varies only what a reader inspects to judg...
429. DeGS: A Scalable 3DGS Architecture via Decoupled Workload Parsing and Reorganization ​
Author: Minnan Pei, Gang Li, Zeyu Zhu, Siting Wang, Junwen Si, Zhuoran Song, Yu Feng, Fangxin Liu, Xiaoyao Liang, Jian Cheng
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.CV
arXiv:2608.02099v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existing 3DGS accelerators suffer from poor architectural scalability: increasing the number of PEs leads to marginal performance improvement duri...
430. Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior ​
Author: Ehsan Yaghoubi, Florian Haselbeck
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.02104v1 Announce Type: cross Abstract: Artificial intelligence offers substantial potential for acoustic monitoring of animals, from welfare assessment in precision livestock farming to wildlife conservation and ecological research, where vocalizations can indicate health, stress, and soc...
431. IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations ​
Author: Dingwei Zhu, Jiahan Li, Chengjun Pan, Yunxian Yang, Yunbin Zhao, Yunke Zhang, Zhonghang Lu, Zhuohui Sheng, Chenhao Huang, Jiahang Lin, Yajie Yang, Junlin Shang, Shichun Liu, Yuhui Wang, Honglin Guo, Junjie Ye, Xin Guo, Jiazheng Zhang, Ming Zhang, Shihan Dou, Zhiheng Xi, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.02110v1 Announce Type: cross Abstract: Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustness via implicit history scanning or text compression, yet predominantly assume perfect instructio...
432. Self-Improving Large Language Models via Progressive Experience Evolution ​
Author: Shijie Ren, Xiting Wang, Meng Li, Yujie Guo, Yunhang Yao, Ziheng Peng, Xunlong Wang, Yuetan Chen, Haoyang Zhou, Yunlong Liang, Fandong Meng
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.02139v2 Announce Type: cross Abstract: Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing self-improvement p...
433. UniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D Reconstruction ​
Author: Haixu Song, Xiaoke Yang, Shengjun Zhang, Jiwen Lu, Yueqi Duan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02145v1 Announce Type: cross Abstract: In this paper, we propose UniqueSplat, a view-conditioned feed-forward 3D Gaussian Splatting model to reconstruct customized 3D radiance fields for each view query. Existing feed-forward methods such as pixelSplat and MVSplat aim to generate fixed Ga...
434. PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs ​
Author: Zhongjie Ba, Shengwang Xu, Peng Cheng, Jinyang Zou, Ting Yu, Zhibo Wang, Zhan Qin
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02150v2 Announce Type: cross Abstract: Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understanding ta...
435. RamanPFN: learning from Raman spectral structure with a tabular foundation model ​
Author: Xingyu Pan, Huan Wang, Jinjia Guo, Zhenlin Zhao, Siming Dong, Jixi Lu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02157v1 Announce Type: cross Abstract: Raman spectroscopy enables non-destructive, label-free molecular characterization across materials science, biomedicine and process monitoring. Predictive Raman datasets often contain few labelled spectra and thousands of ordered wavenumbers, with in...
436. Lossless Tensor Compression as Program Synthesis ​
Author: Jieke Shi (James), Junda He (James), Wenjia Jiang (James), Weifeng Sun (James), Shidong Pan (James), Zhensu Sun (James), Chengran Yang (James), Peixin Zhang (James), Yifan Jia (James), Zhou Yang (James), Thong Hoang (James), Xiwei Xu (Sherry), Zhenchang Xing, David Lo
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.PL
arXiv:2608.02162v1 Announce Type: cross Abstract: Model checkpoints are growing in both number and size, which makes archival, transfer, and deployment increasingly costly. General-purpose compressors can reduce storage requirements but ignore tensor structure, whereas existing tensor-specific compr...
437. Fast Discovery of Inclusion Dependencies with Desbordante ​
Author: Alexander Smirnov, Anton Chizhov, Ilya Shchuckin, Nikita Bobrov, George Chernishev
Published: 8/4/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.DC, cs.LG, cs.PF
arXiv:2608.02213v1 Announce Type: cross Abstract: Inclusion dependency is a relation between attributes of tables that indicates possible Primary Key-Foreign Key references. Automatic discovery of inclusion dependencies is a relevant problem for both academic and industrial communities. The core con...
438. Assessing the Impacts of Imperfect Datasets on Client Selections in Federated Learning ​
Author: Yuan-Heng Tsai, Li-Hsing Yen, Yan-Wei Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02250v1 Announce Type: cross Abstract: Federated learning (FL) is a popular distributed learning framework where multiple clients perform local training and a server aggregates the locally updated models. FL enables decentralized training while preserving the privacy of clients' datasets....
439. HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts ​
Author: Haozhe Luo, Ziyu Zhou, Shelley Zixin Shu, Mauricio Reyes
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02252v1 Announce Type: cross Abstract: Recent vision-language models for chest X-ray understanding are largely built on image-report alignment and therefore rely heavily on MIMIC-CXR as the dominant pretraining source. While effective at scale, this paradigm underexplores an important alt...
440. Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow ​
Author: Jiangling Zhang, Shuxuan Gao, Zeyu Chen, Yichao Liu, Yu Zhou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02258v1 Announce Type: cross Abstract: Rapidly evolving Generative AI enables sophisticated visual text manipulations that increasingly evade current forensic detectors. Existing discriminative models often overfit specific forgery patterns, limiting their generalization to unseen, open-s...
441. TS-MAMP: A Remanufactured Agricultural Robot Powered by Second-Life EV Components and NMS-Free On-Device Weed Detection ​
Author: Weijie Shi, Zicheng Xu, Zhenbang Cheng, Haoran Xuan, Mingbo Duan, Gan Ge
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.02270v1 Announce Type: cross Abstract: Agriculture 4.0 robotic systems improve field efficiency yet remain too capital-intensive for the fragmented smallholdings that dominate global agriculture. Meanwhile, a growing number of retired low-speed electric-vehicle (LSEV) powertrains retain f...
442. BRiG-AFA: Bellman Risk-to-Go Learning for Non-Myopic Active Feature Acquisition ​
Author: Jiaorong Feng, Qian Li, Ying Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02305v1 Announce Type: cross Abstract: Active feature acquisition (AFA) asks which unobserved feature to measure next for each test instance under a budget. Greedy rules are easy to train but can overlook context features whose value is realized only through later acquisitions, while rein...
443. FastGFDs: Efficient Validation of Graph Functional Dependencies with Desbordante ​
Author: Anton Chernikov, Yurii Litvinov, Kirill Smirnov, George Chernishev
Published: 8/4/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.LG, cs.PF
arXiv:2608.02321v1 Announce Type: cross Abstract: Graph functional dependencies (GFD) are a recently-developed concept aimed at capturing both topological structures in graphs and functional dependencies between attributes. The process of verifying whether a given GFD holds over a particular graph i...
444. Context-Aware Mixture of Domain Experts for Bodily Expression of Emotion in the Wild ​
Author: Mohammad Mahdi Dehshibi, David Masip
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02331v1 Announce Type: cross Abstract: The same body posture can convey entirely different emotions depending on its surrounding context, yet most methods for recognising bodily emotions treat scene and object cues as auxiliary feature augmentations rather than as structured priors over t...
445. Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning ​
Author: Botao Dong, Longyang Huang, Ning Pang, Hongtian Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SY, eess.SY
arXiv:2608.02332v1 Announce Type: cross Abstract: In offline reinforcement learning (RL), the distribution shift between behavioral data and the learned policy can lead to erroneous \emph{Q}-value estimation, thereby misguiding the direction of policy optimization. To address this issue, we develop ...
446. Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation ​
Author: Stefan Hut, Lorenzo Masoero
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, stat.AP
arXiv:2608.02345v1 Announce Type: cross Abstract: A/B testing remains the standard for rolling out new features in the technology industry. Each experiment, however, consumes real traffic, engineering effort, and weeks of wall-clock time. Can AI agents---conditioned on behavioral profiles and contex...
447. GLAIM: Learning Global and Local Adaptive Inter-Variable Dependency for Multivariate Time Series Imputation ​
Author: Mingyang Wang, Rongwen Li, Xiao Wang, Changjian Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02366v1 Announce Type: cross Abstract: Multivariate time series imputation is fundamental to downstream analysis, yet modeling inter-variable dependencies with incomplete observations remains challenging. Existing methods learn global dependencies across samples or dynamic local dependenc...
448. GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience ​
Author: Sitong Gong, Caixin Kang, Tianyu Yan, Guo Chen, Bo Zheng, Kaipeng Zhang, Yunzhi Zhuge, Xiang Ruan, Huchuan Lu, Yifei Huang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02392v1 Announce Type: cross Abstract: A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situation. Existing video-memory systems primarily support question-conditioned recall, whereas proactive assistants t...
449. Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification ​
Author: Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmad ElShiekh, Ahmed Rashad
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.02397v1 Announce Type: cross Abstract: We benchmark eleven audio classification methods: five task-aware closed-set LLMs (four Gemini models plus open-weight Kimi-Audio-7B-Instruct), four fixed-vocabulary taggers (YAMNet, PANNs, Whisper-AT, and SSLAM), a zero-shot audio-text model (CLAP),...
450. From fragmented data to actionable design: Physics-calibrated learning for plastic upcycling ​
Author: Jingyang Bai, Zijia Wang, Xiangyi Long, Marcos Millan, Binjian Nie, Mingyue Ding
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02402v1 Announce Type: cross Abstract: Thermochemical upgrading of plastic waste is a key upcycling pathway, yet the experimental literature is fragmented by heterogeneous conditions and incomplete reporting. Complete-case learning would retain only 10.99% of the curated experiments, whil...
451. Antares: Foundation Models for Agentic Vulnerability Localization ​
Author: Supriti Vijay, Aman Priyanshu, Didier Chapoteau, Arthur Goldblatt, Jianliang He, Kimia Majd, Fraser Burch, Baturay Saglam, Takahiro Matsumoto, Zhuoran Yang, Amin Karbasi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02407v1 Announce Type: cross Abstract: Vulnerability localization is a fundamental step in software security, requiring models to reason over large codebases and iteratively identify vulnerable implementations. We present Antares, a family of compact language models (350M, 1B, and 3B para...
452. Human-Centered Reflections on Care Robots: A Comparative Study of Caregiver Perspectives ​
Author: Laura Londo~no, Klaus Baumann, Abhinav Valada, Markus Langer
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.ET, cs.LG
arXiv:2608.02411v1 Announce Type: cross Abstract: Care robots are increasingly being introduced into healthcare settings, raising important questions about their acceptance and ethical implementation. To better understand these challenges, this study investigates caregivers' perceptions of four cate...
453. Agentic Incident Response through Digital Twin-Enhanced Multiscale Planning ​
Author: Yiran Gao, Tao Li, Kim Hammar
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.02422v1 Announce Type: cross Abstract: Incident response is currently managed by security operators using predefined playbooks, resulting in slow, labor-intensive security decision-making processes. Consequently, there is a growing need for automated incident response planning. Decision-t...
454. Syntax Meets Semantics: Understanding Scientific Formulae ​
Author: Yuni Susanti, Moritz Schubotz
Published: 8/4/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.02457v1 Announce Type: cross Abstract: Scientific formulae are a fundamental component of scholarly communication, yet their dual nature -- as structured syntax and carriers of semantics -- remains underexplored in scholarly information retrieval. Although prior studies show that jointly ...
455. Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment ​
Author: Vishwajeet Shivaji Hogale, Anjali Pai, Nitya Ravi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02470v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed as reasoning agents in real-world visual assessment pipelines, yet their spatial grounding remains unreliable for fine-grained, visually ambiguous targets. We study this gap in the context of au...
456. Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery ​
Author: Jiayu Gu, Yiwei Wang, Jie Zhang, Guojun Cao, Keshen Lyu, Song Zhou, Yimeng Chen, Haorui Wang, Qingmin Feng, Shenchao Shi, Huan Zhao, Wenbin Chen, Caihua Xiong, Chidan Wan, Jing Samantha Pan, Xiong Cai, Han Ding
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02471v1 Announce Type: cross Abstract: Computational attention models could help surgeons manage the visual demands of laparoscopy, but they require dense spatial labels that are difficult to obtain because surgical intent is highly specialized and tacit. Here, we introduce DiffeoAfford, ...
457. DyFrDet: Towards Accurate Small Object Detection via Dynamic Frequency Suppression with Label Disambiguation ​
Author: Zihan Yang, Yang Guo, Hongxing Zhang, Dan Lu, Siyuan Yao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.02495v1 Announce Type: cross Abstract: Despite the remarkable progress over the past decades, accurately identifying small objects remains challenging because of their insufficient visual cues. Previous works typically attempt to construct discriminative representation of the small object...
458. SWE-Touch: Benchmarking Coding Agents When Users Touch the Code ​
Author: Yuqiao Tan, Jinxiang Meng, Fangyu Lei, Minzheng Wang, Shizhu He, Jun Zhao, Kang Liu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL
arXiv:2608.02499v1 Announce Type: cross Abstract: Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user p...
459. Analytic Planning under Uncertainty with Moment Closure ​
Author: Shishir Sharma, Doina Precup
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.02519v1 Announce Type: cross Abstract: Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictive uncertainty. Propagating full state distributions analytically offers a principled way to do this, but has traditionally required r...
460. Who Should Be Generated? Justifying Demographic Targets in Open-Ended Generation ​
Author: Zeshen Zheng, Yujia He, Qianmian Lin, Xiangyue Huang, Wenqing Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL
arXiv:2608.02551v1 Announce Type: cross Abstract: Fairness evaluation concerns not only what a model produces, but also what its outputs ought to be compared against. When a model generates "a CEO in the United States," the prompt leaves demographic realization to the model. Existing group fairness ...
461. Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection ​
Author: Anusha Madan Gopal, Aras Pirbadian, Kristofor D. Carlson, M Anthony Lewis, Jonathan Tapson
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR
arXiv:2608.02560v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) imposes a prefill cost proportional to retrieved context length, and -- with Transformer backbones -- a KV-cache that grows with each generated token. State-Space Models (SSMs) avoid the second cost by constructio...
462. CoWAM: Coordination Contracts for Selective Policy Intervention with WAMs ​
Author: Shuaijun Liu, Qifu Wen, Shuyang Hao, Qi Luo, Chenglong Zhang, Feiyang You, Chengyu Wu, Ningxin Su
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2608.02578v1 Announce Type: cross Abstract: World Action Models (WAMs) augment robot policies with action-conditioned predicted futures, but a plausible future alone does not justify changing the action that a bimanual policy would execute. We present CoWAM, a selective intervention layer that...
463. UEmbed: Unified Sparse and Dense Multimodal Embeddings ​
Author: Tingyu Song, Mingxin Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Zhijie Nie, Yilun Zhao, Shu Wu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.IR
arXiv:2608.02583v1 Announce Type: cross Abstract: Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semantics. Yet LSR has so far remaine...
464. Bridging Artificial Intelligence and Power Systems Education Using a Hands-On Executable Framework ​
Author: Junjie Yin (Fran), Buxin She (Fran), Xinyu Feng (Fran), Fangxing (Fran), Li
Published: 8/4/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY
arXiv:2608.02599v1 Announce Type: cross Abstract: Artificial intelligence (AI) is increasingly central to power and energy systems, supporting modeling, forecasting, optimization, and control. Yet most existing works emphasize specialized applications and offer little reusable material for newcomers...
465. Rethinking Inference-Time Scaling: Efficiency Limits and Linguistic Signals ​
Author: Junlin Wang, Shang Zhu, Jon Saad-Falcon, Ben Athiwaratkun, Qingyang Wu, Jue Wang, Shuaiwen Leon Song, Ce Zhang, Bhuwan Dhingra, James Zou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2504.14047v2 Announce Type: replace Abstract: There is intense interest in investigating how inference time compute (ITC) (e.g. repeated sampling, refinements, etc) can improve large language model (LLM) capabilities. While breakthroughs like DeepSeek-R1 highlight the power of reinforcement le...
466. PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning ​
Author: Brahim Driss, Alex Davey, Riad Akrour
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2506.13741v2 Announce Type: replace Abstract: Preference-based reinforcement learning (PbRL) has emerged as a promising approach for learning behaviors from human feedback without predefined reward functions. However, current PbRL methods face a critical challenge in effectively exploring the ...
467. Topology Enhanced MARL for Multi-Agent Cooperative Decision-Making of CAVs ​
Author: Ye Han, Lijun Zhang, Dejian Meng, Zhuang Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2507.12110v2 Announce Type: replace Abstract: Decentralized multi-agent cooperative decision-making in continuous environments is fundamentally bottlenecked by the curse of dimensionality, where undirected exploration typically converges to conservative local optima. We propose Topology-Enhanc...
468. ProbGuard: Proactive Runtime Monitoring for LLM Agent Safety via Probabilistic Prediction ​
Author: Haoyu Wang, Christopher M. Poskitt, Jiali Wei, Jun Sun
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2508.00500v4 Announce Type: replace Abstract: Large Language Model (LLM) agents increasingly operate across domains such as robotics, virtual assistants, and web automation. However, their stochastic decision-making introduces safety risks that are difficult to anticipate during execution. Exi...
469. AIC-VDS: Attention-Based In-Context Learning for Joint Velocity Control and Data Collection Scheduling in Multi-UAV-Assisted Pipeline Monitoring ​
Author: Yousef Emami, Miguel Gutierrez Gaitan, Atefeh Hajijamali Arani, Jingjing Zheng, Hao Zhou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2510.05698v2 Announce Type: replace Abstract: Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous inspection and sensor data collection in large-scale infrastructure monitoring applications, such as pipeline monitoring, where timely anomaly detection is critical. Jointly o...
470. Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs ​
Author: Mina Taraghi, Yann Pequignot, Amin Nikanjam, Mohamed Amine Merzouk, Foutse Khomh
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2511.00382v2 Announce Type: replace Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks. Prior work shows that even fine-tuning on benign datasets can weaken safety alignment, raising a practical question: doe...
471. MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs ​
Author: Liang Shan, Kaicheng Shen, Wen Wu, Zhenyu Ying, Chaochao Lu, Yan Teng, Jingqi Huang, Qingshan Liu, Guangze Ye, Guoqing Wang, Jie Zhou, Liang He
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2511.07107v4 Announce Type: replace Abstract: Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address implicit, domain-specific risks. To investigate this gap, we introduce a dataset of 3,000 annotated qu...
472. Near-Optimal Reinforcement Learning for Constrained Recurrence Objectives ​
Author: Dominik Wagner, Leon Witzman, Luke Ong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2511.19849v2 Announce Type: replace Abstract: Recurrence objectives, where a target region must be visited infinitely often, are a fundamental class of specifications for Markov decision processes (MDPs) and form the core of $\omega$-regular and linear temporal logic (LTL) objectives. We study...
473. Knowledge Graph Augmented Large Language Models for Disease Prediction ​
Author: Ruiyu Wang, Tuan Vinh, Ran Xu, Yuyin Zhou, Jiaying Lu, Francisco Pasquel, Mohammed K Ali, Carl Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2512.01210v4 Announce Type: replace Abstract: Electronic health records (EHRs) enable strong clinical prediction, but explanations are often coarse and hard to use for patient-level decisions. We propose a knowledge graph (KG)-guided chain-of-thought (CoT) framework for visit-level disease pre...
474. Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees ​
Author: Alessandro Breccia, Federica Gerace, Marco Lippi, Gabriele Sicuro, Pierluigi Contucci
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.dis-nn, math-ph, math.MP, math.NT
arXiv:2512.01870v2 Announce Type: replace Abstract: We study whether a transformer network can learn the deterministic sequence of trees generated by the iterated prime factorization of the natural numbers. Each integer is mapped into a rooted planar tree and the resulting sequence $\mathbb{N}\maths...
475. SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection ​
Author: Zhibo Liang, Tianze Hu, Zaiye Chen, Mingjie Tang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CR
arXiv:2512.06716v3 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabilities. By interacting with external environments, LLM agents can execute real-world tasks on behalf of us...
476. LADY: Linear Attention for Autonomous Driving Efficiency without Transformers ​
Author: Jihao Huang, Xi Xia, Zhiyuan Li, Tianle Liu, Jingke Wang, Junbo Chen, Tengju Ye
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2512.15038v3 Announce Type: replace Abstract: End-to-end autonomous driving has emerged as a promising paradigm. However, state-of-the-art methods rely heavily on Transformer architectures. The inherent quadratic complexity of Transformers restricts their ability to model long-range spatial an...
477. Scalable Dynamic Distributed Constraint Optimization with Metareasoning and Application to Continual Satellite Operations ​
Author: Itai Zilberstein, Steve Chien
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.06188v4 Announce Type: replace Abstract: Dynamic distributed constraint optimization problems (DDCOPs) provide a general framework for coordinating autonomous agents in changing environments. However, existing DDCOP formulations do not adequately address settings where optimization and ex...
478. Panning for Gold: Expanding Domain-Specific Knowledge Graphs with General Knowledge ​
Author: Runhao Zhao, Weixin Zeng, Wentao Zhang, Chong Chen, Zhengpin Li, Xiang Zhao, Lei Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.10485v4 Announce Type: replace Abstract: Domain-specific knowledge graphs (DKGs) are critical yet often suffer from limited coverage compared to General Knowledge Graphs (GKGs). Existing tasks to enrich DKGs rely primarily on extracting knowledge from external unstructured data or complet...
479. DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Surveys ​
Author: Guo-Biao Zhang, Xian-Ling Mao, Ding-Yuan Liu, Da-Yi Wu, Tian Lan, Huihui Li, Heyan Huang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2601.15307v2 Announce Type: replace Abstract: The rapid development of automated survey generation technology has made it increasingly important to establish a comprehensive benchmark to evaluate the quality of generated surveys. Most existing benchmarks first construct ground-truth datasets b...
480. Bounded Normative Equivalence in Human-AI Cooperation: Group Behaviour, Not Partner Labels, Predicts Cooperation under Anonymous Aggregate Feedback ​
Author: Nico Mutzner, Taha Yasseri, Heiko Rauhut
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, cs.HC, econ.GN, q-fin.EC
arXiv:2601.20487v3 Announce Type: replace Abstract: The introduction of artificial intelligence (AI) agents into human groups raises questions about how they influence cooperative social norms. Prior work has examined human-AI and human-robot teaming in small groups, but less is known about whether ...
481. CORE: Collaborative Reasoning via Cross Teaching ​
Author: Kshitij Mishra, Mirat Aubakirov, Martin Takac, Nils Lukas, Salem Lahlou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.21600v2 Announce Type: replace Abstract: Large language models exhibit complementary reasoning errors: on the same instance, one model may succeed with a particular decomposition while another fails. We propose Collaborative Reasoning (CORE), a training-time collaboration framework that c...
482. Geometric Analysis of Token Selection in Multi-Head Attention ​
Author: Timur Mudarisov, Mikhal Burtsev, Tatiana Petrova, Radu State
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2602.01893v2 Announce Type: replace Abstract: We present a geometric framework for analysing multi-head attention in large language models (LLMs). Without altering the mechanism, we view standard attention through a top-N selection lens and study its behaviour directly in value-state space. We...
483. Group Selection as a Safeguard Against AI Substitution ​
Author: Qiankun Zhong, Thomas F. Eisenmann, Julian Garcia, Iyad Rahwan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, econ.TH
arXiv:2602.03541v2 Announce Type: replace Abstract: Reliance on generative AI can reduce cultural variance and diversity, especially in creative work. This reduction in variance has already led to problems in model performance, including model collapse and hallucination. In this paper, we examine th...
484. FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use ​
Author: Jiaxuan Lu, Kong Wang, Yemin Wang, Qingmei Tang, Hongwei Zeng, Xiang Chen, Jiahao Pi, Shujian Deng, Lingzhi Chen, Yi Fu, Kehua Yang, Xiao Sun
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.08262v2 Announce Type: replace Abstract: The integration of Large Language Models (LLMs) into the financial domain is driving a paradigm shift from passive information retrieval to dynamic, agentic interaction. While general-purpose tool learning has witnessed a surge in benchmarks, the f...
485. Mind the Sim2Real Gap in User Simulation for Agentic Tasks ​
Author: Xuhui Zhou, Weiwei Sun, Qianou Ma, Yiqing Xie, Jiarui Liu, Weihua Du, Sean Welleck, Yiming Yang, Graham Neubig, Sherry Tongshuang Wu, Maarten Sap
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.11245v2 Announce Type: replace Abstract: As NLP evaluation shifts from static benchmarks to multi-turn interactive settings, LLM-based simulators have become widely used as user proxies, serving two roles: generating user turns and providing evaluation signals. Yet, these simulations are ...
486. When Only the Final Text Survives: Implicit Execution Tracing for Multi-Agent Auditing ​
Author: Yi Nian, Haosen Cao, Shenzhe Zhu, Henry Peng Zou, Qingqing Luan, Yudi Zhang, Yue Zhao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2603.17445v5 Announce Type: replace Abstract: When a multi-agent system produces an incorrect or harmful answer, who is accountable if execution logs and agent identifiers are unavailable? In practice, generated content is often detached from its execution environment due to privacy or system ...
487. Trust or Check? Understanding the (Evolutionary) Dynamics of User Trust in AI Systems ​
Author: Adeela Bashir, Zhao Song, Ndidi Bianca Ogbo, Nataliya Balabanova, Martin Smit, Chin-wing Leung, Paolo Bova, Manuel Chica Serrano, Dhanushka Dissanayake, Manh Hong Duong, Elias Fernandez Domingos, Nikita Huber-Kralj, Marcus Krellner, Andrew Powell, Stefan Sarkadi, Fernando P. Santos, Zia Ush Shamszaman, Chaimaa Tarzi, Paolo Turrini, Grace Ibukunoluwa Ufeoshi, Victor A. Vargas-Perez, Alessandro Di Stefano, Simon T. Powers, The Anh Han
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA, nlin.AO
arXiv:2603.24742v2 Announce Type: replace Abstract: As the capabilities and adoption of Artificial Intelligence (AI) systems grow, trust in these AI systems is an increasingly urgent concern. Much research has focused on models of AI governance and has primarily examined incentives for safe developm...
488. OntoTKGE: Ontology-Enhanced Temporal Knowledge Graph Extrapolation ​
Author: Dongying Lin, Yinan Liu, Shengwei tang, Bin Wang, Xiaochun Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.05468v2 Announce Type: replace Abstract: Temporal knowledge graph (TKG) extrapolation is an important task that aims to predict future facts through historical interaction information within KG snapshots. A key challenge for most existing TKG extrapolation models is handling entities with...
489. Beyond One Output: Visualizing and Comparing Distributions of Language Model Generations ​
Author: Emily Reif, Claire Yang, Jared Hwang, Deniz Nazar, Noah A. Smith, Jeff Heer
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.18724v3 Announce Type: replace Abstract: Users typically interact with and evaluate language models via single outputs, but each output is just one sample from a broad distribution of possible completions. This interaction hides distributional structure such as modes, uncommon edge cases,...
490. HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering ​
Author: Yuyu Liu, Sarang Rajendra Patil, Mengjia Xu, Tengfei Ma
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.21027v2 Announce Type: replace Abstract: Electronic health record (EHR) question answering is often handled by LLM-based pipelines that are costly to deploy and do not explicitly leverage the hierarchical structure of clinical data. Motivated by evidence that medical ontologies and patien...
491. GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation ​
Author: Mingyue Cheng, Yitong Zhou, Jiahao Wang, Qingyang Mao, Qi Liu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.21501v3 Announce Type: replace Abstract: Lithology classification in well logs is a fundamental geoscience data mining task that aims to infer rock types from multi dimensional geophysical sequences. Despite recent progress, existing approaches typically formulate the problem as a static,...
492. When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling ​
Author: Zhimin Lin, Yixin Ji, Jinpeng Li, Yu Luo, Dong Li, Junhua Fang, Juntao Li, Min Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.26644v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) achieve strong performance on mathematical reasoning tasks but remain unreliable on challenging instances. Existing test-time scaling methods, such as repeated sampling, self-correction, and tree search, improve perfor...
493. Unifying biomedical knowledge in a modern multimodal graph ​
Author: Lucas Vittor, Ayush Noori, I~naki Arango, Joaqu'in Polonuer, Sam Rodriques, Andrew White, David A. Clifton, Marinka Zitnik
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.27269v2 Announce Type: replace Abstract: Biomedical knowledge graphs (KGs) are widely used in the life sciences, yet many are derived from unstructured documents and therefore lack schema-level constraints, whereas graphs assembled from structured resources are difficult to harmonize into...
494. Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models ​
Author: Shubham Kumar, Narendra Ahuja
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.00123v2 Announce Type: replace Abstract: Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts. Because we lack a robust understanding of why LLMs are susceptible to jailbreaks, future frontier models operating more autonomou...
495. FitText: Evolving Agent Tool Ecologies via Memetic Retrieval ​
Author: Kyle Zheng, Han Zhang, Renliang Sun, Chenchen Ye, Wei Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.IR, cs.LG, cs.MA
arXiv:2605.02411v3 Announce Type: replace Abstract: Efficient reasoning is not only a matter of shortening an answer trace; for tool-using agents, it also depends on whether the agent is reasoning over the right action space. As API ecosystems scale to tens of thousands of endpoints, the semantic ga...
496. Pro$^2$Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural Tasks ​
Author: Lilin Xu, Bufang Yang, Siyang Jiang, Kaiwei Liu, Kaiyuan Hou, Yuang Fan, Hongkai Chen, Zhenyu Yan, Xiaofan Jiang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2605.04227v2 Announce Type: replace Abstract: Procedural tasks with multiple ordered steps are ubiquitous in daily life. Recent advances in multimodal large language models (MLLMs) have enabled personal assistants that support daily activities. However, existing systems primarily provide react...
497. ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models ​
Author: Yuhang Wang, Wenjie Mei, Junkai Zhang, Guangyu He, Zhenxing Niu, Haichang Gao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.05938v2 Announce Type: replace Abstract: Privacy deletion requests often arrive sequentially, creating a continual unlearning challenge for deployed multimodal large language models (MLLMs). However, existing benchmarks mainly focus on static or short-sequence settings, offering limited s...
498. TeachArena: Are Language Agents Ready for Realistic Teaching Work? ​
Author: Zixin Chen, Peng Liu, Rui Sheng, Haobo Li, Jianhong Tu, Xiaodong Deng, Kashun Shum, Dayiheng Liu, Huamin Qu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.14322v3 Announce Type: replace Abstract: Language agents are increasingly deployed in professional workflows, yet tutoring remains a high-stakes capability that existing evaluations only partially capture. Effective tutor agents require more than producing correct answers or executing acc...
499. Co-ReAct: Rubrics as Step-Level Collaborators for ReAct Agents ​
Author: Jiazheng Kang, Bowen Zhang, Zixin Song, Jiangwang Chen, Xiao Yang, Da Zhu, Guanjun Jiang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.23590v2 Announce Type: replace Abstract: ReAct-style agents for search-intensive, multi-step reasoning tasks rely largely on their own internal judgment to decide what evidence to seek, which reasoning or action step to take next, and when to stop, often producing shallow, redundant, or p...
500. HyperGuide: Hyperbolic Guidance for Efficient Multi-Step Reasoning in Large Language Models ​
Author: Yuyu Liu, Haotian Xu, Yanan He, Sarang Rajendra Patil, Mengjia Xu, Tengfei Ma
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.24140v3 Announce Type: replace Abstract: Multi-step reasoning remains a central challenge for large language models: single-pass generation is efficient but lacks accuracy; tree-search methods explore multiple paths but are computation-heavy. We address this gap by distilling reasoning pr...
501. Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qualitative Analysis ​
Author: Zhimin Lin, Kun Cheng, Zhiyao Shu, Junhua Fang, Juntao Li, Fan Bai, Jie Gao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.24600v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for qualitative data analysis (QDA), yet their outputs often miss the depth and nuance of human analysis. We argue this gap reflects a missing credibility practice from human QDA: peer debriefing, ...
502. Behavioural Analysis of Alignment Faking ​
Author: Nathaniel Mitrani Hadida, Rhea Karty, David Williams-King, Alan Cooney
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2605.27681v2 Announce Type: replace Abstract: Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deployment preferences. Understanding when and why AF arises matters as models grow better at distingui...
503. C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning ​
Author: Yuwei Miao, Gen Li, Yunsheng Zeng, Xiandong Li, Yujin Wang, Siyu Chen, Luning Wang, Yunhao Qiao, Junfeng Wang, Jianwei Lv, Bo Yuan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.27860v2 Announce Type: replace Abstract: Retrieval-augmented generation combined with reinforcement learning has shown promise for grounding large language models in trustworthy medical evidence. However, existing methods rely on exact-match binary rewards, which in clinical diagnosis cau...
504. Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement ​
Author: Jyotirmoy Nath, Neeraj Kumar, Brejesh Lall
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.28360v2 Announce Type: replace Abstract: Automatic prompt optimization (APO) has driven significant gains in LLM-based agentic workflows. However, most existing methods treat each task's prompt as a monolithic, instance-blind string optimized through global edits, producing brittle update...
505. VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis ​
Author: Jiachen Zhang, Junyi Lao, Chenghao Liu, Siyuan Liu, Shixin Wu, Linsen Zhang, Boyu Wang, Songfang Huang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CE
arXiv:2605.28978v2 Announce Type: replace Abstract: Finite Element Analysis (FEA) serves as the cornerstone of modern engineering design. However, its workflow is inherently complex and relies heavily on domain expertise. Although recent efforts have integrated Large Language Models (LLMs) into FEA,...
506. Reliable Post-Retrieval Assembly for Agent Memory: Separating Evidence Extraction from Policy Execution ​
Author: Vikas Reddy, Sumanth Reddy Challaram
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR
arXiv:2606.01435v2 Announce Type: replace Abstract: LLM-based memory systems can retrieve relevant evidence yet still fail when answer generation entangles semantic filtering, conflict resolution, prior suppression, and output generation in one step. We study this failure as a problem of post-retrie...
507. The Violation Situation Pattern: Persistent Representation of Compliance Violations in Knowledge Graphs ​
Author: Nima Kamali Lassem, Fuqi Song, Seyid Amjad Ali
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.03326v2 Announce Type: replace Abstract: Existing compliance pipelines identify violations as transient query results, leaving no persistent representation of the violation itself or its lifecycle, evidence, and audit history. We address this limitation with the Violation Situation Patter...
508. Cross-Lingual Token Arbitrage: Optimizing Code Agent Context Windows via Local LLM Preprocessing ​
Author: Mehmet Utku Colak
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.03618v2 Announce Type: replace Abstract: AI-assisted coding agents are bottlenecked by input-token cost. Two pathologies of raw human input drive much of this overhead: tokenization inefficiency for non-English text and structural entropy in conversational prompts. Existing approaches act...
509. Answer Presence Drives RAG Rewriting Gains ​
Author: Yuejie Li, Yueying Hua, Ke Yang, Li Zhang, Yueping He, Yueping He, Ruiqi Li, Bolin Chen, Tao Wang, Bowen Li, Chengjun Mao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.05633v2 Announce Type: replace Abstract: Retrieval-augmented QA pipelines often route retrieved passages through an LLM \emph{rewriter} before a smaller reader, lifting F1 by tens of points on multi-hop benchmarks; this gain is typically credited to improved evidence quality. We ask wheth...
510. HUSH-Bench: Measuring Memory-Use Boundaries for Sensitive History in Conversational Agents ​
Author: Lingxiang Xu, Jiaoyun Yang, Min Hu, Ning An
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.06055v2 Announce Type: replace Abstract: Long-term memory helps conversational agents maintain continuity across sessions, while relevance and current-turn warrant remain distinct decisions. We study this boundary under a stated conservative policy in which sensitive history shapes a resp...
511. Some hypotheses on how chatbots work in problem-solving-driven conversations. Large Language Models as confirmation of the Innovation Illusion ​
Author: S. F. M. van Vlijmen, H. D. Lethe jr
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.07722v4 Announce Type: replace Abstract: We discuss the nature of chatbots as conversation partners in problem-solving. What can chatbots do and what can't they do? We develop hypotheses on how this can this be explained. Our argument draws on insights from Aggregation Dynamics, Cognitive...
512. Mitigating Visual Hallucinations in Multimodal Systems through Retrieval-Augmented Reliability-Aware Inference ​
Author: Pratheswaran Hariharan, Haiping Xu, Donghui Yan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2606.15782v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-language understanding and natural-language response generation. However, these systems can still produce overconfident predictions and hallucination-like outp...
513. Rhythm of the Deep: Two-Tier Combinatorial Structure in Sperm Whale Codas Revealed by Acoustic Unit Induction ​
Author: Mudit Sinha, Sanika Chavan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2606.16084v2 Announce Type: replace Abstract: Sperm-whale codas are conventionally described as recurring click-count and timing patterns. We show instead that their waveforms contain a two-tier combinatorial acoustic organization. Recurring click units combine with inter-click rhythm to form ...
514. BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery ​
Author: Jieyi Wang, Bingxuan Li, Nanyi Jiang, Desong Meng, Zirui Fan, Yuxin Guo, Jiayu Liu, Kunlun Zhu, Eddie Yang, Xiusi Chen, Pan Lu, Bingxin Zhao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.20997v2 Announce Type: replace Abstract: Biomedical deep-research systems increasingly retrieve and synthesize scientific evidence, but their outputs typically collapse heterogeneous evidence into static text, making provenance difficult to inspect and reuse. We formulate evidence-centere...
515. Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation ​
Author: Han-yu Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2606.26502v2 Announce Type: replace Abstract: Large reasoning models (LRMs) take longer on harder problems, just as humans do, but that surface similarity hides an opposite pattern within items. When an LRM gets a problem wrong it spends more tokens than when it gets that same problem right; h...
516. ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks ​
Author: Qitai Tan, Zefang Zong, Mo Li, Yipeng Shi, Yang Li, Peng Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.27814v5 Announce Type: replace Abstract: Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the early stage, but ...
517. CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents ​
Author: Bo Qu, Mingguang Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, q-fin.CP, q-fin.PM
arXiv:2606.29771v2 Announce Type: replace Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Yet most still rank agents by returns over a fixed window, a weak proxy: the market path dominates...
518. Distributionally Robust Listwise Preference Optimization ​
Author: Xudong Wu, Jian Qian, Pangpang Liu, Vaneet Aggarwal, Jiayu Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.01715v2 Announce Type: replace Abstract: Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level. We instead study listwise preference optimization under ranking-label u...
519. Branch-JEPA: Finite-Support Predictive Distributions for JEPA World Models ​
Author: Zhi Song, Ximing Xing, Zhenchao Tang, hanbo Huang, Jiehui Huang, Weilong Yan, Tianxu Lv, Minghao Yang, Zhongzheng Niu, Bing He, Lusheng Wang, Jianhua Yao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.05238v3 Announce Type: replace Abstract: Joint-embedding predictive architectures (JEPAs) learn dynamics by predicting future observations in representation space. Yet most JEPA world models return one latent successor, even when hidden intent, partial observation, or stochastic dynamics ...
520. PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation ​
Author: Hyungseok Song, Junseok Park, Won-Seok Choi, Seohui Bae, Han-Seul Jeong, Youngjoon Park, Soonyoung Lee
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.05915v2 Announce Type: replace Abstract: PCB routing is the task of connecting the nets of a board with copper traces under strict design rules, yet learning-based methods still lag behind rule-based routers. We introduce PCBWorld, an open-source engine-grounded PCB routing environment bu...
521. Trust but Verify:Evidence-Linked Multi-Agent Clinical Information Extraction in Pathology ​
Author: Yufan Wang, Anit Kumar Sahu, Yan Fei Ng, Daniel Kang, Shayan Vassef, Soorya Ram Shimgekar, Koustuv Saha, Piyum Zonooz, Navin Kumar, Chee Leong Cheng, Li Yan Khor
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.06435v2 Announce Type: replace Abstract: Clinical feature extraction from pathology reports is challenging because relevant evidence may be distributed across coded and narrative fields and depend on specimen attribution, negation, ancillary findings, and diagnostic context. We retrospect...
522. Length Penalties Make Chain-of-Thought Less Monitorable ​
Author: Bryce Little
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2607.09786v3 Announce Type: replace Abstract: To curb overthinking and reduce inference costs, researchers now train reasoning models with penalties on chain of thought length. We find that these penalties degrade monitorability. Shorter chains of thought mention misleading hints less often, b...
523. Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions ​
Author: Zachary Izzo
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.14306v2 Announce Type: replace Abstract: In this paper, we study the connection between an LLM's output distribution and the data used to train it. Specifically, we study the degree to which an LLM's next-token distribution agrees with the empirical next-token distribution (ENTD) given th...
524. TopoTuner: Topological Finetuning of Large Language Models ​
Author: Abdulkadir Erol, Yash Mahajan, Vepaul Hariprashad, Baha Rababah, Santu Karmaker, Cuneyt G. Akcora, Mubarak Shah
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.16637v2 Announce Type: replace Abstract: Full fine-tuning remains a strong way to adapt pretrained LLMs, but it updates all weights and can be expensive. LoRA reduces the number of trainable parameters, but it does not directly answer which pretrained components should be trained and whic...
525. OTAP: Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories ​
Author: Babak Barazandeh, Subhabrata Majumdar, George Michailidis
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2607.17082v2 Announce Type: replace Abstract: Large language model agents solve tasks by generating trajectories that interleave planning, tool calls, and intermediate results. Current evaluation metrics reduce such a trajectory to a binary success flag, compare it against a reference by exact...
526. Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding ​
Author: Zhiwei Wei, Yonghe Sun, Zhenjia Liu, Wenjia Xu, Chao He, Weihua Dong, Chunbo Liu, Hua Liao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2607.17999v2 Announce Type: replace Abstract: Spatial understanding is crucial for foundation models (FMs), and maps have long helped humans organize and reason about geographic information. This study examines whether choropleth maps remain useful for machine spatial understanding when models...
527. Lifted State Hypothesis in Large Language Models ​
Author: Bumjin Park, Jaesik Choi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.19360v2 Announce Type: replace Abstract: Large language models (LLMs) adapt rapidly through fine-tuning and in-context learning, yet it remains unclear which inputs they treat as the same case and why their predictions change together. We introduce the Lifted State Hypothesis. Under fixed...
528. CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs ​
Author: Pukaphol Thienpreecha ("Volk"), Karthik Subramanian
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.19396v2 Announce Type: replace Abstract: Document-based LLM systems often flatten a PDF before guardrails inspect it. That step can discard evidence that an instruction was never visible to the user. We introduce CrackedPDFs, a controlled benchmark for hidden prompt injection in PDFs. The...
529. Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning ​
Author: Guanqun Zhao, Zijun Xie, Binbin Zheng, Enlei Gong, Jiafeng Lu, Yehan Yang, Aoqi Hu, Zeyu Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22186v3 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause pol...
530. Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization ​
Author: Muhammad Junaid Ali, Smail Niar, El-Ghazali Talbi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.22583v2 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However, deploying them in embedded and edge computing environments remains challenging because of strict latency, mem...
531. Design Theater: Evaluating the Gap Between User-Facing Design Reasoning and Implementation in Generative UI Tools ​
Author: Kashif Imteyaz, Kaif Imteyaz, Nakul Rajpal, Kaif Shaikh, Michael Muller, Saiph Savage
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22928v2 Announce Type: replace Abstract: Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces. Alongside the interface, these tools generate user-facing design rationales that explain their layout, accessibility, and design...
532. Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Agent ​
Author: Weijie Xia, Stefanie Horian, Hanyue Huang, Queena K. Qian, Jie Yang, Pedro P. Vergara
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.24341v3 Announce Type: replace Abstract: Recent studies use Large language models (LLMs) to simulate human opinions and decisions by prompting models with demographic, attitudinal, or persona-based descriptions. Yet such simulations rarely model the practical, cognitive, or social frictio...
533. Individual-level interventions against sycophantic AI reduce its appeal but not its persuasiveness ​
Author: Meryl Ye, Robert Kraut, Steve Rathje
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.25166v3 Announce Type: replace Abstract: AI chatbots can be "sycophantic," or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call "sycophancy blindness"). We tested whether incre...
534. Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response ​
Author: Abu Bakar Siddik
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.25379v2 Announce Type: replace Abstract: Cyber-capable AI agents combine language models with tools, memory, and execution environments to perform multi-step offensive-security tasks. Existing work separately measures cyber capability and catalogs attacks against agent components, but pro...
535. AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution ​
Author: Junhao Qiu, Zidong Wang, Yansong Sun, Zhitong Ma, Ping Guo, Qingfu Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.26661v2 Announce Type: replace Abstract: Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise. While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally dif...
536. Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data ​
Author: Lingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.27056v2 Announce Type: replace Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract pers...
537. Multi-Head Attention Residuals ​
Author: Cheng Luo, Zefan Cai, Junjie Hu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.27230v2 Announce Type: replace Abstract: Transformers propagate information across depth through a single additive residual stream: every sublayer reads only the most recent state. Attention residuals relax this by letting each sublayer attend, through a learned softmax. However, that rea...
538. AI and Its Impact on Creativity and Diversity: An Empirical Study of LLM-Generated Product Ideas ​
Author: Christian Terwiesch, Lennart Meincke, Karan Girotra, Ethan Mollick, Gideon Nave, Karl T. Ulrich
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, econ.GN, q-fin.EC
arXiv:2607.27553v2 Announce Type: replace Abstract: This research examines how well large language models, or LLMs, generate new product ideas for college students priced under $50. Across a series of studies, we identify key strengths and weaknesses of using LLMs for product innovation. Our first s...
539. Group-Reflective Self-Distillation for Agentic Reinforcement Learning ​
Author: Binbin Zheng, Zijun Xie, Guanqun Zhao, Enlei Gong, Xing Ma, Xiaoliang Fu, Zeyu Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.28076v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) is effective for training large language model agents. However, terminal rewards provide only coarse trajectory-level supervision, leaving successful behaviors, recurring mistakes, and incidenta...
540. Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners ​
Author: Feng Xiong, Leyan Xue, Hongyu Lin
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.28336v2 Announce Type: replace Abstract: On-policy distillation provides dense supervision for multimodal reasoners, but its trajectory-level reward cannot determine whether a failed answer arose from perception or subsequent reasoning. Perception Success Rate (PSR), estimated from multip...
541. InfoOps Bench: A live information operations safety benchmark ​
Author: Dorian Quelle, Lisa-Maria Neudert, Jonathan Bright, John Gallacher
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.28503v2 Announce Type: replace Abstract: In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for state-backed information operations. We draw on over 2,100 information operations from a live mo...
542. Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design ​
Author: Manisha Dubey, Rimvydas Rubavicius, N. Siddharth, Subramanian Ramamoorthy
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AI, stat.ML
arXiv:2607.28894v2 Announce Type: replace Abstract: Computational cognitive modeling seeks to infer latent cognitive mechanisms underlying observed behavior. Bayesian inverse planning provides a principled framework for such inference, but its success depends critically on the experimental environme...
543. The Elements of Differentiable Programming ​
Author: Mathieu Blondel, Vincent Roulet
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.PL
arXiv:2403.14606v4 Announce Type: replace-cross Abstract: Artificial intelligence has recently experienced remarkable advances, fueled by large models, vast datasets, accelerated hardware, and, last but not least, the transformative power of differentiable programming. This new programming paradigm ...
544. OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset ​
Author: Allen Roush, Yusuf Shabazz, Arvind Balaji, Peter Zhang, Stefano Mezza, Markus Zhang, Sanjay Basu, Sriram Vishwanath, Mehdi Fatemi, Ravid Shwartz-Ziv
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2406.14657v4 Announce Type: replace-cross Abstract: We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community. This dataset includes over 3.5 million documents with rich metadata, making it one of the m...
545. Information-Theoretic Foundations for Machine Learning ​
Author: Hong Jun Jeon, Benjamin Van Roy
Published: 8/4/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2407.12288v5 Announce Type: replace-cross Abstract: The progress of machine learning over the past decade is undeniable. In retrospect, it is both remarkable and unsettling that this progress was achievable with little to no rigorous theory to guide experimentation. Despite this fact, practiti...
546. Cost-Based Semantics for Querying Inconsistent Weighted Knowledge Bases ​
Author: Meghyn Bienvenu, Camille Bourgaux, Robin Jean
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.DB
arXiv:2407.20754v3 Announce Type: replace-cross Abstract: In this paper, we explore a quantitative approach to querying inconsistent description logic knowledge bases. We consider weighted knowledge bases in which both axioms and assertions have (possibly infinite) weights, which are used to assign ...
547. Defending Membership Inference Attacks via Privacy-aware Sparsity Tuning ​
Author: Hengxiang Zhang, Qiang Hu, Hongxin Wei
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2410.06814v2 Announce Type: replace-cross Abstract: Over-parameterized models are typically vulnerable to membership inference attacks, which aim to determine whether a specific sample is included in the training of a given model. Previous Weight regularizations (e.g., L1 regularization) typic...
548. Belief-Contraction-Driven Active Inverse Source Localization and Characterization ​
Author: Yiwei Shi, Mengyue Yang, Qi Zhang, Cunjia Liu, Weinan Zhang, Weiru Liu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2501.13084v2 Announce Type: replace-cross Abstract: Active inverse source localization and characterization (ISLC) in dynamic fields requires sequential decision making under partial observability, where a mobile sensor must infer latent source parameters from sparse, noisy readings. We introd...
549. ApplE: A Modular Ontology of Applied Ethics and Event Context for Ethical Decision Modeling ​
Author: Aisha Aijaz, Raghava Mutharaju, Manohar Kumar
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2502.05110v2 Announce Type: replace-cross Abstract: Applied ethics applies ethical decision-making to domain-specific contexts using contextual information such as agents, actions, temporal and spatial settings, and theoretical constructs such as utility, virtues, rights, and duties. However, ...
550. GradientStabilizer:Fix the Norm, Not the Gradient ​
Author: Tianjin Huang, Zhangyang Wang, Haotian Hu, Zhenyu Zhang, Gaojie Jin, Xiang Li, Li Shen, Jiaxing Shang, Tianlong Chen, Ke Li, Lu Liu, Qingsong Wen, Shiwei Liu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2502.17055v5 Announce Type: replace-cross Abstract: Training instability in modern deep learning systems is frequently triggered by rare but extreme gradient-norm spikes, which can induce oversized parameter updates, corrupt optimizer state, and lead to slow recovery or divergence. Widely used...
551. Partial Excitation in Parameter Learning ​
Author: Ganghui Cao, Shimin Wang, Martin Guay, Jinzhi Wang, Zhisheng Duan, Marios M. Polycarpou
Published: 8/4/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY, eess.SP, math.OC
arXiv:2503.02235v3 Announce Type: replace-cross Abstract: This paper investigates parameter learning problems under Partial Persistent Excitation (PPE). The PPE condition is a rank-deficient, and therefore, a more general evolution of the well-known Persistent Excitation (PE) condition. Under the PP...
552. Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation ​
Author: Zhibo Wang, Wenhui Zhang, Huiyu Xu, Zeqing He, Ziqi Zhu, Kui Ren
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2503.06519v2 Announce Type: replace-cross Abstract: Small language models (SLMs) have emerged as promising alternatives to large language models (LLMs) due to their low computational demands, enhanced privacy guarantees, and comparable performance in specific domains. Deploying SLMs on edge de...
553. Oscillatory Hierarchical Reservoirs for Human-like Rhythm Perception and Anticipation ​
Author: Zhongju Yuan, Geraint Wiggins, Dick Botteldooren
Published: 8/4/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI
arXiv:2503.12509v2 Announce Type: replace-cross Abstract: Rhythm is a fundamental aspect of human behaviour, present from infancy and deeply embedded in cultural practices. Rhythm anticipation often occurs before surface event onsets, yet most neuroscience and artificial intelligence studies focus o...
554. Sampling Decisions: Exact Path-Space Control for Physics-Informed Generative Sampling ​
Author: Michael Chertkov, Hamidreza Behjoo, Sungsoo Ahn
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.stat-mech, cs.AI, cs.SY, eess.SY, stat.ML
arXiv:2503.14549v4 Announce Type: replace-cross Abstract: Scientific generative models must turn tractable local decisions into globally correlated samples that respect physical constraints. We introduce Sampling Decisions, a finite-horizon framework in which a structured object is assembled on a gr...
555. Understanding Machine Unlearning Through the Lens of Mode Connectivity ​
Author: Jiali Cheng, Hadi Amiri
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CV
arXiv:2504.06407v2 Announce Type: replace-cross Abstract: Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress, the loss landscape and optimization geometry of unlearning are poorly understood. In this paper, we stu...
556. Fairness in Augmented Graph Learning: A Survey ​
Author: Renqiang Luo, Huafei Huang, Ziqi Xu, Xikun Zhang, Enyan Dai, Bo Yang, Feng Xia
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2504.21296v2 Announce Type: replace-cross Abstract: Graph learning has evolved into Augmented Graph Learning (AGL) by integrating specialized machine learning (ML) techniques. Examples include federated learning, graph transformers, and graph condensation. While enhancing model utility, AGL in...
557. Deep Learning for Retinal Degeneration Assessment: A Comprehensive Analysis of the MARIO Challenge ​
Author: Rachid Zeghlache, Ikram Brahim, Pierre-Henri Conze, Mathieu Lamard, Mohammed El Amine Lazouni, Zineb Aziza Elaouaber, Leila Ryma Lazouni, Christopher Nielsen, Ahmad O. Ahsan, Matthias Wilms, Nils D. Forkert, Lovre Antonio Budimir, Ivana Matovinovi'c, Donik Vr\v{s}nak, Sven Lon\v{c}ari'c, Philippe Zhang, Weili Jiang, Yihao Li, Yiding Hao, Markus Frohmann, Patrick Binder, Marcel Huber, Taha Emre, Teresa Finisterra Ara'ujo, Marzieh Oghbaie, Hrvoje Bogunovi'c, Amerens A. Bekkers, Nina M. van Liebergen, Hugo J. Kuijf, Abdul Qayyum, Moona Mazher, Steven A. Niederer, Alberto J. Beltr'an-Carrero, Juan J. G'omez-Valverde, Javier Torresano-Rodr'iquez, 'Alvaro Caballero-Sastre, Mar'ia J. Ledesma Carbayo, Yosuke Yamagishi, Yi Ding, Robin Peretzke, Alexandra Ertl, Maximilian Fischer, Jessica K"achele, Sofiane Zehar, Karim Boukli Hacene, Thomas Monfort, B'eatrice Cochener, Mostafa El Habib Daho, Anas-Alexis Benyoussef, Gwenol'e Quellec
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2506.02976v4 Announce Type: replace-cross Abstract: The MARIO challenge, held at MICCAI 2024, focused on advancing the automated detection and monitoring of age-related macular degeneration (AMD) through the analysis of optical coherence tomography (OCT) images. Designed to evaluate algorithmi...
558. Grounded Vision-Language Interpreter for Long-Horizon Bimanual Task and Motion Planning ​
Author: Jeremy Siburian, Keisuke Shirai, Cristian C. Beltran-Hernandez, Masashi Hamaya, Michael G"orner, Atsushi Hashimoto
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2506.03270v3 Announce Type: replace-cross Abstract: While recent advances in vision-language models have accelerated language-guided robot planning, their black-box nature lacks the safety guarantees and interpretability crucial for real-world deployment. Conversely, classical symbolic planner...
559. Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models ​
Author: Ren-Jian Wang, Ke Xue, Zeyu Qin, Ziniu Li, Sheng Tang, Hao-Tian Li, Shengcai Liu, Zhi Yu, Yuanpeng Tan, Chao Qian
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE
arXiv:2506.07121v2 Announce Type: replace-cross Abstract: Ensuring the safety and robustness of large language models (LLMs) is a fundamental challenge and a critical prerequisite for the responsible deployment of artificial intelligence. Red-teaming, a systematic framework to identify adversarial p...
560. MGDFIS: Multi-scale Global-detail Feature Integration Strategy for Small Object Detection ​
Author: Yuxiang Wang, Xuecheng Bai, Chuanzhi Xu, Ying Zhou, Weidong Cai
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2506.12697v4 Announce Type: replace-cross Abstract: Small-object detection in Unmanned Aerial Vehicle (UAV) imagery requires preserving weak local evidence while using broader context to separate tiny foreground targets from cluttered backgrounds. Existing multi-scale fusion methods improve fe...
561. Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems ​
Author: Weixin Liang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.HC, cs.LG
arXiv:2506.17467v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown significant potential to change how we write, communicate, and create, leading to rapid adoption across society. This dissertation examines how individuals and institutions are adapting to and engaging ...
562. Rethinking Group Recommender Systems in the Era of Generative AI: From One-Shot Recommendations to Agentic Group Decision Support ​
Author: Dietmar Jannach, Amra Deli'c, Francesco Ricci, Markus Zanker
Published: 8/4/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2507.00535v2 Announce Type: replace-cross Abstract: More than twenty-five years ago, first ideas were developed on how to design a system that can provide recommendations to groups of users instead of individual users. Since then, a rich variety of algorithmic proposals were published, e.g., o...
563. Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models ​
Author: Ken Tsui
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2507.02778v3 Announce Type: replace-cross Abstract: Although large language models (LLMs) have transformed AI, they still make errors and follow unproductive reasoning paths. Self-correction is vital for safety-critical applications, but studying it requires disentangling activation failure fr...
564. The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and Mapping Study ​
Author: Amr Mohamed, Maram Assi, Mariam Guizani
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.HC
arXiv:2507.03156v3 Announce Type: replace-cross Abstract: Large language model assistants (LLM-assistants) present new opportunities to transform software development. Developers are increasingly adopting these tools across tasks, including coding, testing, debugging, documentation, and design. Yet,...
565. Leveraging Synthetic Data for Question Answering with Multilingual LLMs in the Agricultural Domain ​
Author: Rishemjit Kaur, Arshdeep Singh Bhankhar, Jashanpreet Singh Salh, Sudhir Rajput, Vidhi, Kashish Mahendra, Bhavika Berwal, Ritesh Kumar, Surangika Ranathunga
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2507.16974v3 Announce Type: replace-cross Abstract: Enabling farmers to access accurate agriculture-related information in their native languages in a timely manner is crucial for the success of the agriculture field. Publicly available general-purpose Large Language Models (LLMs) typically of...
566. CLONE: Continuous Latent Optimization for Normal Estimation via 3D Gaussian Splatting ​
Author: Yanxing Liang, Yinghui Wang, Wei Li, Tao Yan, Jiaxing Shen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2508.05950v3 Announce Type: replace-cross Abstract: We propose CLONE, a Continuous Latent Optimization framework for Normal Estimation via 3D Gaussian splatting. The core idea is to construct an image-geometry-image consistency loop that unifies explicit geometric representation with different...
567. A Rule-Based Approach to Specifying Preferences over Conflicting Facts and Querying Inconsistent Knowledge Bases ​
Author: Meghyn Bienvenu, Camille Bourgaux, Katsumi Inoue, Robin Jean
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.DB
arXiv:2508.07742v3 Announce Type: replace-cross Abstract: Repair-based semantics have been extensively studied as a means of obtaining meaningful answers to queries posed over inconsistent knowledge bases (KBs). While several works have considered how to exploit a priority relation between facts to ...
568. Learning Graph-Indexed Trajectory Patterns for Stochastic On-Time Arrival Routing ​
Author: Yuanhang Wang, Xing Wei, Duoxiang Zhao, Zezhou Zhang, Hao Qin, Yuqi Ouyang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2508.17218v4 Announce Type: replace-cross Abstract: Correlated link travel times create decision-relevant patterns in partial route histories. In stochastic on-time arrival (SOTA) routing, each route prefix forms a variable-length, graph-indexed sequence in which traversed-edge identities, rea...
569. CAPMix: Robust KPI Anomaly Detection for AIOps in Noisy and Dynamic Environments ​
Author: Xudong Mou, Rui Wang, Tiejun Wang, Zexin Wu, Fangda Guo, Jie Sun, Shiru Chen, Penghao Zhang, Tiezi Zhang, Tianyu Wo, Hao Peng, Chunming Hu, Xudong Liu, Renyu Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2509.06419v2 Announce Type: replace-cross Abstract: Time-series anomaly detection is crucial in AIOps for maintaining large-scale service reliability. In production, streams of Key Performance Indicators (KPI) are high-dimensional, non-stationary, and affected by noise, deployment changes, and...
570. RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking ​
Author: Zhenguo Sun, Yibo Peng, Yuan Meng, Xukun Li, Bo-Sheng Huang, Zhenshan Bing, Xinlong Wang, Alois Knoll
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2509.20717v2 Announce Type: replace-cross Abstract: Long-horizon, high-dynamic motion tracking on humanoids remains brittle: retargeted reference motions are typically kinematically plausible but dynamically inconsistent with the robot, so small tracking errors accumulate and eventually destab...
571. A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models ​
Author: Wenjun Wang, Shuo Cai, Congkai Xie, Mingfa Feng, Yiming Zhang, Zhen Li, Kejing Yang, Ming Li, Jiannong Cao, Hongxia Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2509.22536v5 Announce Type: replace-cross Abstract: The immense computational cost of training Large Language Models (LLMs) presents a major barrier to innovation. While FP8 training offers a promising solution with significant theoretical efficiency gains, its widespread adoption has been hin...
572. AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification ​
Author: Shaohao Rui, Kaitao Chen, Weijie Ma, Xiaosong Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2509.24560v2 Announce Type: replace-cross Abstract: Extended Chain-of-Thought (CoT) reasoning has significantly bolstered the capabilities of medical large language models (LLMs). However, current models exhibit static computational expenditure, applying lengthy reasoning processes indiscrimin...
573. TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning ​
Author: Fangxu Yu, Hongyu Zhao, Tianyi Zhou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.03519v2 Announce Type: replace-cross Abstract: Time series reasoning is crucial to decision-making in diverse domains, including finance, energy, and scientific discovery. While existing time series foundation models (TSFMs) can capture low-level dynamic patterns and provide accurate fore...
574. Eigenvalues as a Metric for Memory Dynamics in Sequence Models ​
Author: Rahel Rickenbach, Jelena Trisovic, Alexandre Didier, Jerome Sieber, Melanie N. Zeilinger
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SY, eess.SY
arXiv:2510.09379v2 Announce Type: replace-cross Abstract: While softmax attention drives state-of-the-art performance in sequence modeling, its quadratic complexity motivates linear alternatives such as state space models (SSMs). Structural differences between the two model classes, however, hinder ...
575. Unpacking Hateful Memes: Presupposed Context and False Claims ​
Author: Weibin Cai, Jiayu Li, Reza Zafarani
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.09935v2 Announce Type: replace-cross Abstract: While memes are often humorous, they are frequently used to disseminate hate, causing serious harm to individuals and society. Current approaches to hateful meme detection mainly rely on pre-trained language models. However, less focus has be...
576. LLM generation novelty through the lens of semantic similarity ​
Author: Philipp Davydov, Ameya Prabhu, Matthias Bethge, Elisa Nguyen, Seong Joon Oh
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2510.27313v3 Announce Type: replace-cross Abstract: Generation novelty is a key indicator of an LLM's ability to generalize, yet measuring it against full pretraining corpora is computationally challenging. Existing evaluations often rely on lexical overlap, failing to detect paraphrased text,...
577. Extending Fair Null-Space Projections for Continuous Attributes to Kernel Methods ​
Author: Felix St"orck, Fabian Hinder, Barbara Hammer
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2511.03304v3 Announce Type: replace-cross Abstract: With the on-going integration of machine learning systems into the everyday social life of millions the notion of fairness becomes an ever increasing priority in their development. Fairness notions commonly rely on protected attributes to ass...
578. Interpretable Recognition of Cognitive Distortions in Natural Language Texts ​
Author: Anton Kolonin, Anna Arinicheva
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.LG
arXiv:2511.05969v2 Announce Type: replace-cross Abstract: We propose a new approach to multi-factor classification of natural language texts based on weighted structured patterns such as N-grams, taking into account the heterarchical relationships between them, applied to solve such a socially impac...
579. Introduction to Automated Negotiation ​
Author: Dave de Jonge
Published: 8/4/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.GT
arXiv:2511.08659v5 Announce Type: replace-cross Abstract: This book is an introductory textbook targeted towards computer science students who are completely new to the topic of automated negotiation. It does not require any prerequisite knowledge, except for elementary mathematics and basic program...
580. DeepDefense: Robust Learning via Layer-Wise Gradient-Feature Alignment ​
Author: Ci Lin, Tet Yeap, Iluju Kiringa
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2511.13749v2 Announce Type: replace-cross Abstract: Deep neural networks are known to be vulnerable to adversarial perturbations, which are small, carefully crafted inputs that lead to incorrect predictions. In this paper, we propose DeepDefense, a novel defense framework that applies Gradient...
581. New York Smells: A Large Multimodal Dataset for Olfaction ​
Author: Ege Ozguroglu, Junbang Liang, Ruoshi Liu, Mia Chiquier, Michael DeTienne, Wesley Wei Qian, Alexandra Horowitz, Andrew Owens, Carl Vondrick
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2511.20544v2 Announce Type: replace-cross Abstract: While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines. One key bottleneck is the lack of diverse, multimodal olfactory training data collected in natural set...
582. Latent Collaboration in Multi-Agent Systems ​
Author: Jiaru Zou, Ruizhong Qiu, Gaotang Li, Xiyuan Yang, Katherine Tieu, Pan Lu, Ke Shen, Hanghang Tong, Yejin Choi, Jingrui He, James Zou, Mengdi Wang, Ling Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2511.20639v4 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) extend large language models (LLMs) from independent single-model reasoning to coordinative system-level intelligence. While existing LLM agents depend on text-based mediation for reasoning and communication, we take...
583. Revisiting Generalization Across Difficulty Levels: It's Not So Easy ​
Author: Yeganeh Kordi, Nihal V. Nayak, Max Zuo, Ilana Nguyen, Stephen H. Bach
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2511.21692v2 Announce Type: replace-cross Abstract: We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curation and evaluation. Existing research is mixed regarding whether training on easier or harder data lead...
584. NORi: An ML-Augmented Ocean Boundary Layer Parameterization ​
Author: Xin Kai Lee, Ali Ramadhan, Andre Souza, Gregory LeClaire Wagner, Simone Silvestri, John Marshall, Raffaele Ferrari
Published: 8/4/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.AI, cs.LG, physics.comp-ph, physics.flu-dyn
arXiv:2512.04452v3 Announce Type: replace-cross Abstract: NORi is a machine learning (ML) parameterization of ocean boundary layer turbulence that is physics-based and augmented with neural networks. NORi stands for neural ordinary differential equations (NODEs) Richardson number (Ri) closure. The p...
585. Auto-exploration for online reinforcement learning ​
Author: Caleb Ju, Guanghui Lan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC
arXiv:2512.06244v3 Announce Type: replace-cross Abstract: The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. Existing algorithms for finite state and action discounted RL problems address this by assuming sufficient exploration ...
586. Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving ​
Author: Yuzhe Gu, Songyang Gao, Zijian Wu, Lingkai Kong, Wenwei Zhang, Zhongrui Cai, Fan Zheng, Tianyou Ma, Junhao Shen, Haiteng Zhao, Duanyang Zhang, Huilun Zhang, Kuikun Liu, Chengqi Lyu, Yanhui Duan, Chiyu Chen, Ningsheng Ma, Jianfei Gao, Han Lyu, Dahua Lin, Kai Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2512.10739v3 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) have expanded the mathematical reasoning frontier through Chain-of-Thought (CoT) techniques and Reinforcement Learning with Verifiable Rewards (RLVR), capable of solving AIME-level problems. However, the performa...
587. Human-like working memory signatures emerge from intrinsically plastic artificial neurons for robust dynamic vision ​
Author: Jingli Liu, Huannan Zheng, Bohao Zou, Kezhou Yang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.ET, cs.AI, cs.CV, cs.NE
arXiv:2512.15829v4 Announce Type: replace-cross Abstract: While the unsustainable energy cost of artificial intelligence necessitates physics-driven computing, its performance superiority over full-precision GPUs remains a challenge. We bridge this gap by repurposing the Joule-heating relaxation dyn...
588. Breaking Self-Attention Failure: Rethinking Query Initialization for Infrared Small Target Detection ​
Author: Yuteng Liu, Duanni Meng, Yimian Dai, Maoxun Yuan, Xingxing Wei, Bo li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2601.02837v2 Announce Type: replace-cross Abstract: Infrared small target detection (IRSTD) faces significant challenges due to low signal-to-noise ratios, extremely small target sizes, and complex cluttered backgrounds. Although DETR-based detectors benefit from global context modeling, their...
589. Visualising Information Flow in Word Embeddings with Diffusion Tensor Imaging ​
Author: Thomas Fabian
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2601.05713v2 Announce Type: replace-cross Abstract: Understanding how large language models (LLMs) represent natural language is a central challenge in natural language processing (NLP) research. Many existing methods extract word embeddings from an LLM, visualise the embedding space via point...
590. Emerging Threats and Countermeasures in Neuromorphic Systems: A Survey ​
Author: Pablo Sorrentino, Stjepan Picek, Ihsen Alouani, Nikolaos Athanasios Anagnostopoulos, Francesco Regazzoni, Lejla Batina, Tamalika Banerjee, Fatih Turkmen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.ET
arXiv:2601.16589v2 Announce Type: replace-cross Abstract: Neuromorphic computing mimics brain-inspired mechanisms through spiking neurons and energy-efficient processing, offering a pathway to efficient in-memory computing (IMC). However, these advancements raise critical security and privacy concer...
591. When LLM Essays Outscore Student Essays: What a Korean Writing Rubric Rewards and Where Readers Disagree ​
Author: Shinwoo Park, Yo-Sub Han
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.19913v4 Announce Type: replace-cross Abstract: LLMs now help students plan, draft, and revise essays. Educational assessment therefore faces a basic question: how should student and LLM writing be compared? Rubrics assign points to content, organization, and expression. Their total can st...
592. Can Small Language Models Handle Context-Summarized Multi-Turn Customer-Service QA? A Synthetic Data-Driven Comparative Evaluation ​
Author: Lakshan Cooray, Deshan Sumanathilaka, Pattigadapa Venkatesh Raju
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2602.00665v4 Announce Type: replace-cross Abstract: Customer-service question answering (QA) systems increasingly rely on conversational language understanding. While Large Language Models (LLMs) achieve strong performance, their high computational cost and deployment constraints limit practic...
593. MedTextWeaver: Procedural Knowledge Evolution in Agentic Medical Text Editing ​
Author: Ziyan Xiao, Yinghao Zhu, Liang Peng, Kyongtae T Bae, Lequan Yu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2602.00740v2 Announce Type: replace-cross Abstract: Medical text editing is essential for improving communication among diverse stakeholders in clinical settings. However, adapting LLM agents to this task remains challenging because expert supervision is often sparse, fragmented, and distribut...
594. AROpt: An Optimization Method for Autoregressive Time Series Forecasting ​
Author: Zheng Li, Jerry Cheng, Huanying Gu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.02288v3 Announce Type: replace-cross Abstract: Current time-series forecasting models are primarily based on transformer-style neural networks. These models achieve long-term forecasting mainly by scaling up the model size rather than through genuinely autoregressive (AR) rollout. From th...
595. RAP: KV-Cache Compression via RoPE-Aligned Pruning ​
Author: Jihao Xin, Tian Lyu, David Keyes, Hatem Ltaief, Marco Canini
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.02599v4 Announce Type: replace-cross Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the memory and compute of the key-value (KV) cache. Structured pruning is a direct way to shrink it: dropping the least useful channels of the W_k, W_v projection weigh...
596. Exploring Silicon-Based Societies: An Early Study of the Moltbook Agent Community ​
Author: Yu-Zheng Lin, Bono Po-Jen Shih, Hsuan-Ying Alessandra Chien, Shalaka Satam, Jesus Horacio Pacheco, Naima Kaabouch, Sicong Shao, Soheil Salehi, Pratik Satam
Published: 8/4/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CY
arXiv:2602.02613v4 Announce Type: replace-cross Abstract: The rapid emergence of autonomous large language model agents has given rise to persistent, large-scale agent ecosystems whose collective behavior cannot be adequately understood through anecdotal observation or small-scale simulation. This p...
597. Transformers perform adaptive partial pooling ​
Author: Vsevolod Kapatsinski
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2602.03980v2 Announce Type: replace-cross Abstract: Any language model must decide what to say in novel contexts based on information from similar contexts. But what about contexts that are not novel but merely infrequent? In hierarchical regression, the model's predictions for behavior in a c...
598. Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models ​
Author: Moritz Miller, Florent Draye, Bernhard Sch"olkopf
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2602.04718v4 Announce Type: replace-cross Abstract: A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not ...
599. SVRepair: Structured Visual Reasoning for Automated Program Repair ​
Author: Jincheng Wang, Liwei Luo, Xiaoxuan Tang, Jingxuan Xu, Sheng Zhou, Dajun Chen, Wei Jiang, Yong Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CV
arXiv:2602.06090v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently been applied to Automated Program Repair (APR), yet most existing approaches remain unimodal and fail to use diagnostic signals contained in visual artifacts such as screenshots and control-flow grap...
600. Optimized Piecewise Affine Abstractions of Neural Networks with Learnable Activation Functions ​
Author: Noah Schwartz, Chandra Kanth Nagesh, Sriram Sankaranarayanan, Ramneet Kaur, Tuhin Sahai, Susmit Jha
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO
arXiv:2602.06737v2 Announce Type: replace-cross Abstract: We present a generalized framework for the range verification of neural networks featuring non-linear activation functions. Our approach first constructs an ``optimized piecewise affine abstraction" of the network that replaces each non-linea...
601. RAG Strategies for Natural Language-Based SQL Query and REST API Call Generation ​
Author: Tim Schlippe, Simon Martin, Michael Marketsm"uller
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL
arXiv:2602.07086v2 Announce Type: replace-cross Abstract: Enterprise software systems commonly expose business functionality through both relational databases and REST APIs. Accessing these interfaces requires specialized technical knowledge, as users must determine whether a request requires a data...
602. SoK: DARPA's AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned ​
Author: Cen Zhang, Younggi Park, Fabian Fleischer, Yu-Fu Fu, Jiho Kim, Dongkwan Kim, Youngjoon Kim, Qingxiao Xu, Andrew Chin, Ze Sheng, Hanqing Zhao, Michael Pelican, David J. Musliner, Jeff Huang, Jon Silliman, Mikel Mcdaniel, Jefferson Casavant, Isaac Goldthwaite, Nicholas Vidovich, Matthew Lehman, Taesoo Kim
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2602.07666v5 Announce Type: replace-cross Abstract: DARPA's AI Cyber Challenge (AIxCC, 2023--2025) is the largest competition to date for building fully autonomous cyber reasoning systems (CRSs) that leverage recent advances in AI -- particularly large language models (LLMs) -- to discover and...
603. Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents ​
Author: Haochen Wang, Yi Wu, Daryl Chang, Li Wei, Lukasz Heldt
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.10226v3 Announce Type: replace-cross Abstract: Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyperparameter search space and, more critically, designing sophisticated optimizers, architectures, and ...
604. LakeMLB: Data Lake Machine Learning Benchmark ​
Author: Feiyu Pan, Tianbin Zhang, Aoqian Zhang, Yu Sun, Zheng Wang, Lixing Chen, Li Pan, Jianhua Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.10441v2 Announce Type: replace-cross Abstract: Data lakes have become a fundamental platform for large-scale machine learning by enabling flexible management of heterogeneous data. Despite their growing importance, standardized benchmarks for evaluating machine learning performance in dat...
605. Chimera: Neuro-Symbolic Attention Primitives for Trustworthy Dataplane Intelligence ​
Author: Rong Fu, Xiaowen Ma, Kun Liu, Wangyu Wu, Ziyu Kong, Jia Yee Tan, Tailong Luo, Xianda Li, Yongtai Liu, Youjin Wang, Simon Fong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.CR, cs.LG
arXiv:2602.12851v5 Announce Type: replace-cross Abstract: Deploying expressive learning models directly on programmable dataplanes promises line-rate, low-latency traffic analysis but remains hindered by strict hardware constraints and the need for predictable, auditable behavior. Chimera introduces...
606. Nonparametric Distribution Regression Re-calibration ​
Author: 'Ad'am Jung, Domokos M. Kelen, Andr'as A. Bencz'ur
Published: 8/4/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2602.13362v2 Announce Type: replace-cross Abstract: A key challenge in probabilistic regression is ensuring that predictive distributions accurately reflect true empirical uncertainty. Minimizing overall prediction error often encourages models to prioritize informativeness over calibration, p...
607. Provably Safe Generative Sampling with Constricting Barrier Functions ​
Author: Darshan Gadginmath, Ahmed Allibhoy, Fabio Pasqualetti
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SY, eess.SY, math.OC
arXiv:2602.21429v3 Announce Type: replace-cross Abstract: Flow-based generative models, such as diffusion models and flow matching models, have achieved remarkable success in learning complex data distributions. However, a critical gap remains for their deployment in safety-critical domains: the lac...
608. Quantifying Frontier LLM Capabilities for Container Sandbox Escape ​
Author: Rahul Marchand, Art O Cathain, Jerome Wynne, Philippos Maximos Giavridis, Stuart Jennings, Freddy Tuxworth, Tolga H. Dur, Sam Deverett, John Wilkinson, Jason Gwartz, Harry Coppock
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2603.02277v3 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly act as autonomous agents, using tools to execute code, read and write files, and access networks, creating novel security risks. To mitigate these risks, agents are commonly deployed and evaluated in ...
609. From We to Me: Theory Informed Narrative Shift with Abductive Reasoning ​
Author: Jaikrishna Manojkumar Patil, Divyagna Bavikadi, Kaustuv Mukherji, Ashby Steward-Nolan, Peggy-Jean Allin, Tumininu Awonuga, Joshua Garland, Paulo Shakarian
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.03320v2 Announce Type: replace-cross Abstract: Effective communication often relies on aligning a message with an audience's narrative and worldview. Narrative shift involves transforming text to reflect a different narrative framework while preserving its original core message--a task we...
610. Local Shapley: Model-Induced Locality and Optimal Reuse in Data Valuation ​
Author: Xuan Yang, Hsi-Wen Chen, Ming-Syan Chen, Jian Pei
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DB, cs.GT
arXiv:2603.03672v2 Announce Type: replace-cross Abstract: The Shapley value provides a principled foundation for data valuation, but exact computation is #P-hard due to the exponential coalition space. Existing accelerations remain global and ignore a structural property of modern predictors: for a ...
611. Heterogeneous Decentralized Diffusion Models ​
Author: Zhiying Jiang, Raihan Seraj, Marcos Villagra, Bidhan Roy
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2603.06741v3 Announce Type: replace-cross Abstract: Training frontier-scale diffusion models often requires substantial computational resources concentrated in tightly-coupled clusters, limiting participation to well-resourced institutions. While Decentralized Diffusion Models (DDM) enable tra...
612. GPrune-LLM: Generalization-Aware Structured Pruning for Large Language Models ​
Author: Xiaoyun Liu, Divya Saxena, Jiannong Cao, Yuqing Zhao, Yiying Dong, Penghui Ruan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2603.13418v2 Announce Type: replace-cross Abstract: Structured pruning is widely applied to compress large language models (LLMs), but its performance depends heavily on how neuron importance is estimated. Most existing methods rely on activation statistics from a single calibration set, which...
613. MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model ​
Author: Jinguang Tong, Jinbo Wu, Kaisiyuan Wang, Zhelun Shen, Xuan Huang, Mochu Xiang, Xuesong Li, Yingying Li, Haocheng Feng, Chen Zhao, Hang Zhou, Wei He, Chuong Nguyen, Jingdong Wang, Hongdong Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.14686v2 Announce Type: replace-cross Abstract: Human-Object Interaction (HOI) video reenactment aims to transfer the interaction dynamics of a source video to a novel target object while preserving realistic hand-object coordination. Existing methods typically rely on sparse 2D motion con...
614. MAPLE: Metadata Augmented Private Language Evolution ​
Author: Eli Chien, Yuzheng Hu, Ryan McKenna, Shanshan Wu, Zheng Xu, Peter Kairouz
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CR, cs.LG
arXiv:2603.19258v2 Announce Type: replace-cross Abstract: Differentially private (DP) fine-tuning of large language models (LLMs) requires massive compute and full model access, which rules out state-of-the-art proprietary APIs for general users. Generating DP synthetic data offers a practical worka...
615. Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification ​
Author: Masnun Nuha Chowdhury, Nusrat Jahan Beg, Umme Hunny Khan, Syed Rifat Raiyan, Md Kamrul Hasan, Hasan Mahmud
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA
arXiv:2603.28488v3 Announce Type: replace-cross Abstract: Large language models (LLMs) remain unreliable for high-stakes claim verification due to hallucinations and shallow reasoning. While retrieval-augmented generation (RAG) and multi-agent debate (MAD) address this, they are limited by one-pass ...
616. Hierarchical Pre-Training of Vision Encoders with Large Language Model ​
Author: Eugene Lee, Ting-Yu Chang, Jui-Huang Tsai, Jiajie Diao, Chen-Yi Lee
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG
arXiv:2604.00086v2 Announce Type: replace-cross Abstract: The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks. However, existing approaches often treat vision encoders and large language models (LLMs) as indepe...
617. LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering ​
Author: Sing Hieng Wong, Hassan Sajjad, A. B. Siddique
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2604.03532v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong multilingual capabilities, yet reliably controlling the language of their outputs remains difficult. Representation-level steering addresses this by adding language-specific vectors to model activation...
618. Self-Preference Bias in Rubric-Based Evaluation of Large Language Models ​
Author: Jos'e Pombal, Ricardo Rei, Andr'e F. T. Martins
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.06996v3 Announce Type: replace-cross Abstract: LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB): they tend to favor outputs produced by themselves or by models from their own family. This skews eval...
619. Face-D(^2)CL: Multi-Domain Synergistic Representation with Dual Continual Learning for Facial DeepFake Detection ​
Author: Yushuo Zhang, Yu Cheng, Yongkang Hu, Jiuan Zhou, Jiawei Chen, Zhaoxia Yin
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.08159v2 Announce Type: replace-cross Abstract: Facial forgery techniques are advancing rapidly, posing severe threats to public trust and information security while imposing higher demands on the continual adaptation of DeepFake detection models. Although continual learning enables models...
620. QARIMA: A Quantum Approach To Classical Time Series Analysis ​
Author: Nishikanta Mohanty, Bikash K. Behera, Badshah Mukherjee, Pravat Dash, Giuseppe Sergioli, Roberto Giuntini
Published: 8/4/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG
arXiv:2604.08277v3 Announce Type: replace-cross Abstract: We present QARIMA, a quantum state-similarity-based reconstruction of the classical ARIMA modelling pipeline. Rather than using a quantum circuit as a standalone forecaster, QARIMA preserves ARIMA's interpretable forecasting structure while r...
621. Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time ​
Author: Ruizhi Zhang, Ye Huang, Yuangang Pan, Chuanfu Shen, Zhilin Liu, Ting Xie, Haijun Lei, Lixin Duan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.08340v2 Announce Type: replace-cross Abstract: While artificial intelligence has mastered structured games like chess and Go, vision-language agents still struggle in visually-driven 3D games without access to game states. Existing game environments typically evaluate a fixed agent config...
622. Tail-Aware Information-Theoretic Bounds for LLM Alignment under Heavy-Tailed Rewards ​
Author: Huiming Zhang, Binghan Li, Wan Tian, Qiang Sun
Published: 8/4/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, math.PR, math.ST, stat.TH
arXiv:2604.10727v2 Announce Type: replace-cross Abstract: Classical information-theoretic learning bounds typically rely on KL mutual information and moment-generating-function (MGF) arguments, which are well matched to bounded or sub-Gaussian losses but can be ineffective when losses or rewards are...
623. QShield: Securing Neural Networks Against Adversarial Attacks using Quantum Circuits ​
Author: Navid Azimi, Aditya Prakash, Yao Wang, Li Xiong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CV, cs.LG, quant-ph
arXiv:2604.10933v2 Announce Type: replace-cross Abstract: Deep neural networks remain highly vulnerable to adversarial perturbations, limiting their reliability in security- and safety-critical applications. To address this challenge, we introduce QShield, a modular hybrid quantum-classical neural n...
624. ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception ​
Author: Huanzhen Wang, Ziheng Zhou, Jiaqi Song, Li He, Yunshi Lan, Yan Wang, Wenqiang Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.12255v2 Announce Type: replace-cross Abstract: Dynamic facial expression recognition in the wild remains challenging due to data scarcity and long-tail distributions, which hinder models from effectively learning the temporal dynamics of scarce emotions. To address these limitations, we p...
625. AST: Adaptive, Seamless, and Training-Free Precise Speech Editing ​
Author: Sihan Lv, Yechen Jin, Zhen Li, Jintao Chen, Jinshan Zhang, Ying Li, Jianwei Yin, Meng Xi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2604.16056v2 Announce Type: replace-cross Abstract: Text-based speech editing aims to modify specific segments while preserving speaker identity and acoustic context. Current approaches generally involve either expensive task-specific training or adapting pre-trained Text-to-Speech (TTS) model...
626. Dual-Resolution Attention-Gated Deep Learning with Ordinal Regression for Diabetic Retinopathy Grading: A Quantified Assessment of Cross-Domain Generalization ​
Author: Afshan Hashmi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.17341v2 Announce Type: replace-cross Abstract: Diabetic retinopathy (DR) is a leading cause of preventable blindness, and automated grading could extend screening capacity. However, most reported DR models are validated only on the dataset they were trained on, leaving their behaviour und...
627. Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers ​
Author: Harri Renney, Fouad Trad, Michael Mattarock, Jayden Evetts, Zena Wood
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.DC, cs.PF
arXiv:2604.24785v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are becoming increasingly capable at small parameter scales. At the same time, conventional cloud-centric deployment introduces challenges around data privacy, latency, and cost that are acute in operational techn...
628. SymphonyGen: 3D Hierarchical Orchestral Generation with Controllable Harmony Skeleton ​
Author: Xuzheng He, Nan Nan, Zhilin Wang, Ziyue Kang, Zhuoru Mo, Ao Li, Yu Pan, Xiaobing Li, Feng Yu, Xiaohong Guan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2604.25498v2 Announce Type: replace-cross Abstract: Generating symphonic music requires simultaneously managing high-level structural form and dense, multi-track orchestration, yet existing symbolic models often struggle with a "complexity-control imbalance" between scalability and steerabilit...
629. Meritocratic Fairness via $K$-Shapley Values in Budgeted Combinatorial Bandits with Full-Bandit Feedback ​
Author: Shradha Sharma, Shweta Jain, Swapnil Dhamal
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA
arXiv:2605.00762v2 Announce Type: replace-cross Abstract: We study meritocratic fairness in budgeted combinatorial multi-armed bandits with full-bandit feedback, where a learner selects at most $K$ arms per time step and observes only the noisy aggregate reward of the selected set. To define merit u...
630. SOD: Step-wise On-policy Distillation for Small Language Model Agents ​
Author: Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin, Jie Sun, Houcheng Jiang, Xiang Wang, Junfeng Fang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.07725v2 Announce Type: replace-cross Abstract: Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and limited model capacity. While reinforcement learning methods like group relative policy optimization provi...
631. Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory ​
Author: Daniel Goldstein, Navneel Singhal, Eugene Cheah
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2605.09877v5 Announce Type: replace-cross Abstract: Recall presents a difficult choice: transformers have a linearly growing memory that slows each successive token, while linear RNNs typically have fixed costs but limited recall. We present Key-Value Means ("KVM"), a novel block-recurrence fo...
632. Bridging the Cognitive Gap: A Unified Memory Paradigm for 6G Agentic AI-RAN ​
Author: Xijun Wang, Zhaoyang Liu, Chenyuan Feng, Xiang Chen, Howard H. Yang, Tony Q. S. Quek
Published: 8/4/2026, 4:00:00 AM
Categories: cs.NI, cs.AI
arXiv:2605.10036v2 Announce Type: replace-cross Abstract: As 6G evolves, the radio access network must transcend traditional automation to embrace agentic AI capable of perception, reasoning, and evolution. A fundamental cognitive gap persists in current disaggregated architectures, where interfaces...
633. Formally Verifying Analog Neural Networks Under Process Variations Using Polynomial Zonotopes ​
Author: Yasmine Abu-Haeyeh, Tobias Ladner, Matthias Althoff, Lars Hedrich
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.10474v2 Announce Type: replace-cross Abstract: Analog neural networks are gaining attention due to their efficiency in terms of power consumption and processing speed. However, since analog neural networks are implemented as physical circuits, they are highly sensitive to manufacturing pr...
634. The Transformer as a Polar State Estimator ​
Author: Peter Racioppo
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.11007v3 Announce Type: replace-cross Abstract: We show that the core components of the Transformer---attention, residual connections, and normalization---arise naturally from a single geometric state estimation problem. Modeling the latent state in polar coordinates naturally separates ra...
635. Cochise: A Reference Harness for Autonomous Penetration Testing ​
Author: Andreas Happe, J"urgen Cito
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE
arXiv:2605.11671v2 Announce Type: replace-cross Abstract: Recent work on LLM-driven autonomous penetration testing reports promising results, but existing systems often bundle architectural, prompting, and tool-integration choices together. This makes it difficult to determine what is gained over a ...
636. Representing Higher-Order Networks: A Survey of Graph-Based Frameworks ​
Author: Takaaki Fujita, Florentin Smarandache
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SI, cs.AI, cs.CE, math.CO
arXiv:2605.12509v3 Announce Type: replace-cross Abstract: Many real-world phenomena can be naturally represented using graphs and networks. Classical graph models, however, are generally restricted to pairwise interactions and may therefore be insufficient for capturing the richer structural relatio...
637. Counterfactual Reasoning for Causal Responsibility Attribution in Probabilistic Multi-Agent Systems ​
Author: Chunyan Mu, Muhammad Najib
Published: 8/4/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2605.13077v2 Announce Type: replace-cross Abstract: Responsibility allocation -- determining the extent to which agents are accountable for outcomes -- is a fundamental challenge in the design and analysis of multi-agent systems. In this work, we model such systems as concurrent stochastic mul...
638. Characterizing Readability Issue Patterns and the Role of Prompt Design in LLM-Generated Code ​
Author: Hengzhi Ye, Fengyuan Ran, Weiwei Xu, Minghui Zhou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2605.13280v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly changing how code is produced, but generated code still requires human review and validation before it can be adapted or integrated into real-world projects. This makes the readability of LLM-gene...
639. Spectral Analysis of Fake News Propagation ​
Author: Weibin Cai, Reza Zafarani
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SI, cs.AI
arXiv:2605.13861v2 Announce Type: replace-cross Abstract: How can we systematically represent the propagation of information? The propagation structure of fake news has been shown to be an important cue for detecting it; yet, existing propagation-based fake news detection methods have mainly relied ...
640. Asymmetric Generative Recommendation via Kronecker Residual Bridge and Multi-Faceted Hierarchical Quantization ​
Author: Bin Huang, Xin Wang, Junwei Pan, Yongqi Zhou, Yifeng Zhou, Zhixiang Feng, Shudong Huang, Haijie Gu, Wenwu Zhu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2605.14512v2 Announce Type: replace-cross Abstract: Generative Recommendation (GenRec) models reformulate recommendation as a sequence generation task, representing items as discrete Semantic IDs used symmetrically as both inputs and prediction targets. We identify a critical dual-stage inform...
641. Margin-Adaptive Confidence Ranking for Reliable LLM Judgement ​
Author: Gaojie Jin, Yong Tao, Lijia Yu, Tianjin Huang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.15416v3 Announce Type: replace-cross Abstract: Jung et al. (2025) introduce a hypothesis testing framework for guaranteeing agreement between large language models (LLMs) and human judgments, relying on the assumption that the model's estimated confidence is monotonic with respect to huma...
642. When Bits Break Recourse: Counterfactual-Faithful Quantization ​
Author: Chaymae Yahyati, Ismail Lamaakal, Khalid El Makkaoui, Ibrahim Ouahbi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2605.17160v3 Announce Type: replace-cross Abstract: Model quantization is widely used to reduce memory, latency, and deployment cost, and is typically judged by whether predictive accuracy is preserved. In decision systems that provide algorithmic recourse, however, accuracy preservation is no...
643. Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation ​
Author: Mingfei Sun
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.18591v2 Announce Type: replace-cross Abstract: Natural policy gradients improve optimization by accounting for the geometry of distribution space, but their practical use is limited by the cost of estimating and inverting the Fisher matrix. We present Randomized Advantage Transformation (...
644. Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora ​
Author: Maciej Skorski
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.22660v3 Announce Type: replace-cross Abstract: Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages. Idiomatic expressions, slang, and cultural references introduce hard-to-avoid translation artefacts. Yet automated moral values cl...
645. CogAdapt: Adapting Clinical ECG Foundation Models for Wearable Cognitive Load Assessment ​
Author: Amir Mousavi, Erfan Nourbakhsh, Mohammad Sadegh Sirjani, Mimi Xie, Rocky Slavin, Leslie Neely, John Davis, John Quarles
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.HC
arXiv:2605.22774v5 Announce Type: replace-cross Abstract: Assessing cognitive load continuously and at low latency would help adaptive human-computer interaction, but it remains hard because labeled data are scarce and models generalize poorly across subjects. Recent ECG foundation models, pre-train...
646. MambaGaze: Bidirectional Mamba with Explicit Missing Data Modeling for Cognitive Load Assessment from Eye-Gaze Tracking Data ​
Author: Amir Mousavi, Mohammad Sadegh Sirjani, Erfan Nourbakhsh, Mimi Xie, Rocky Slavin, Leslie Neely, John Davis, John Quarles
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.HC
arXiv:2605.22775v3 Announce Type: replace-cross Abstract: Real-time cognitive load assessment from eye-tracking signals could enable adaptive human-centered AI in safety-critical applications such as driver vigilance monitoring or automated flight deck assistance, yet two challenges persist: handlin...
647. OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations ​
Author: Jiangwang Chen, Bowen Zhang, Zixin Song, Jiazheng Kang, Xiao Yang, Da Zhu, Guanjun Jiang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.23668v2 Announce Type: replace-cross Abstract: Although large language model (LLM) conversational systems process millions of multi-turn dialogues daily, they remain fundamentally reactive: they respond only after the user types a query. A key step toward proactive interaction is next-que...
648. A Hamiltonian-Inspired Local-Operator Ansatz for Slimming Large Language Models ​
Author: Ying Lu, Peng-Fei Zhou, Qi-Xuan Fang, Pan Zhang, Shi-Ju Ran, Gang Su
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, quant-ph
arXiv:2605.25344v2 Announce Type: replace-cross Abstract: Dense linear maps carry much of the parameter and computational burden of modern neural networks, yet their dense form leaves the organization of learned couplings implicit. Quantum many-body physics organizes exponentially large operators by...
649. Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals ​
Author: Federico Torrielli, Peter Schneider-Kamp, Lukas Galke Poech
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.26045v2 Announce Type: replace-cross Abstract: An activation oracle is a language model trained to read another model's internal activations and describe them in natural language, for example to name a secret word the other model was trained to hide. Oracle answers carry no measure of con...
650. QSignAI: Quantum-Randomness-Seeded Identity Signatures at the Intersection of AI for Science and Science for AI ​
Author: Dongping Liu, Aoyu Zhang, Luyao Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.ET, quant-ph
arXiv:2605.27729v4 Announce Type: replace-cross Abstract: The 2024-2025 Nobel and Turing awards recognised AI and quantum science simultaneously. Yet no deployed system has brought these streams together for the public. This paper presents QSignAI, a production-deployed platform demonstrating a bidi...
651. DELOS: Detecting Shallow Transits in Kepler Photometry Using a Contrastive-Learning Framework ​
Author: Qingtian Liu, Jian Ge, XingChen Yan, Kevin Willis, Xinyu Yao, QuanQuan Hu, Jiapeng Zhu
Published: 8/4/2026, 4:00:00 AM
Categories: astro-ph.EP, astro-ph.IM, cs.AI
arXiv:2605.29428v2 Announce Type: replace-cross Abstract: We present DEtection in phase-folded Light curves with cOntrastive Scoring (DELOS), a contrastive-learning-based framework designed to search for shallow transits in Kepler photometry. DELOS combines GPU-accelerated phase folding, optimized p...
652. Enhancing Regime Shift Detection Using Unstructured Data: A Study on the Treasury Market ​
Author: Mingxuan Yi, Vidal Mehra, Jing Chen, John Cartlidge
Published: 8/4/2026, 4:00:00 AM
Categories: q-fin.CP, cs.AI, cs.LG, q-fin.ST
arXiv:2605.30363v2 Announce Type: replace-cross Abstract: Regime shifts in financial markets reorganise the joint dynamics of asset prices and macro variables, breaking any single-regime calibration. They are nonetheless hard to identify: the data signal is noisy and heavily multicollinear, while th...
653. Certificates without Electrons? Theory and Evidence on Impacts from AI-Driven Power Demand ​
Author: Dana Golden, Aruna Balasubramanian, Niranjan Balasubramanian
Published: 8/4/2026, 4:00:00 AM
Categories: econ.EM, cs.AI
arXiv:2606.00811v2 Announce Type: replace-cross Abstract: Data centers now account for 4.4% of United States electricity demand, yet the grid-level effectiveness of the renewable energy certificates (RECs) and power purchase agreements (PPAs) hyperscalers use to claim carbon neutrality remains uncle...
654. DeliChess: A Multi-party Dialogue Dataset for Deliberation in Chess Puzzle Solving ​
Author: Xiaochen Zhu, Georgi Karadzhov, Tom Stafford, Andreas Vlachos
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2606.04987v2 Announce Type: replace-cross Abstract: Multi-party dialogue is a critical setting for studying collaborative reasoning and decision-making, yet existing datasets rarely focus on structured, reasoning-intensive tasks. We introduce DeliChess, a dataset of group deliberation dialogue...
655. Streaming Communication in Multi-Agent Reasoning ​
Author: Zhen Yang, Xiaogang Xu, Wen Wang, Cong Chen, Xander Xu, Ying-Cong Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA
arXiv:2606.05158v2 Announce Type: replace-cross Abstract: Multi-agent reasoning systems adopt a "generate-then-transfer" paradigm that forces end-to-end latency to scale linearly with pipeline depth. We introduce StreamMA, a multi-agent reasoning system that streams each reasoning step to downstream...
656. TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation ​
Author: Eric Spencer, Arslan Bisharat, Brian Ortiz, Khushboo Bhadauria, Mujtaba Nazari, TaiNing Wang, George K. Thiruvathukal, Konstantin Laufer, Mohammed Abuhamad
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG, cs.LO
arXiv:2606.06133v5 Announce Type: replace-cross Abstract: TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequently produce TLA+ specifications that fail the TLC model checker for semantic reasons. Across 25 LLMs,...
657. Plug-and-Play Guidance for Discrete Diffusion Models via Gradient-Informed Logit Correction ​
Author: Hongkun Dou, Zike Chen, Fengji Li, Hongjue Li, Yue Deng
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.06303v2 Announce Type: replace-cross Abstract: Controllable generation with discrete diffusion models is often hindered by high computational overhead or the need for retraining. In this paper, we present \underline{\textbf{G}}radient-\underline{\textbf{I}}nformed \underline{\textbf{L}}og...
658. Will the Agent Recuse, and Will It Stop? Measuring LLM-Agent Compliance with In-Band Governance Signals at the Access Door and Mid-Flight ​
Author: Thamilvendhan Munirathinam
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2606.06460v4 Announce Type: replace-cross Abstract: As autonomous LLM agents hold real credentials and operate infrastructure without a human in the loop, operators cannot tell an agent a resource is off-limits, or ask a running agent to stand down. We propose an in-band governance signal -- t...
659. AgentCompile: An LLM-Guided Compiler for Direct CUDA Inference ​
Author: Xuanzhe Li, Ziyan Weng, Zhiyu Zhu, Junhui Hou
Published: 8/4/2026, 4:00:00 AM
Categories: cs.PL, cs.AI
arXiv:2606.07665v2 Announce Type: replace-cross Abstract: Transformer inference increasingly relies on specialized compiler and runtime support, while recent LLMs can generate nontrivial CUDA kernels. However, unconstrained generation guarantees neither correctness nor performance. We present \texts...
660. Rewrite to Translate, Translate to Reward: Reinforcement Learning for Source Rewriting in Machine Translation ​
Author: Boxuan Lyu, Haiyue Song, Zhi Qu, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.08011v3 Announce Type: replace-cross Abstract: Prior work has explored prompting large language models (LLMs) to rewrite source text before translation, with the goal of improving machine translation (MT) quality. However, we find that such prompt-based rewriting can degrade translation q...
661. Hacking Generative Perplexity: Why Unconditional Text Evaluation Needs Distributional Metrics ​
Author: Antonio Franca, Alexander Tong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.08417v2 Announce Type: replace-cross Abstract: Diffusion and continuous flow-based language models have emerged as the leading non-autoregressive alternatives to language modeling. Progress in both paradigms is overwhelmingly tracked by generative perplexity (gen-PPL): the per-token negat...
662. CARE: Context-Aware Ranking Evolution with Executable Scoring Programs for Budgeted Reaction Optimization ​
Author: Guanyu Liu, Weiyi Kong, Chao Tang, Zeyu Wang, Boer Zhang, Baiqing Li, Peiyu Zhang, Tianyu Shi
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.14581v4 Announce Type: replace-cross Abstract: High-throughput experimentation can evaluate many reaction conditions, yet combinatorial condition spaces still exceed the available experiment budget. This makes experiment selection a sequential decision problem: each new condition must be ...
663. Few-Shot Biomedical Relation Extraction with Large Language Models: A Viable Alternative to Supervised Learning? ​
Author: Jakob Mraz, Toma\v{z} Curk, Bla\v{z} Zupan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.15412v2 Announce Type: replace-cross Abstract: Biomedical relation extraction (BioRE) is a key step in transforming biomedical literature into structured knowledge. Most existing approaches rely on supervised models trained on costly annotated datasets, limiting their scalability and adap...
664. Distilling Drifting Transformers with Representation Autoencoders ​
Author: Jiawei Zhang, Mengfei Xia, Gen Li, Yuantao Gu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.15553v2 Announce Type: replace-cross Abstract: Despite the significant training acceleration and promising performance, Representation Autoencoders (RAEs) are mainly criticized for poor distillation effectiveness. In this work, we argue that RAE is competent at high-quality one-step gener...
665. Entropy-Gated Latent Recursion ​
Author: Soham Bhattacharjee, Dushyant Singh Chauhan, Salem Lahlou, Martin Takac, Nils Lukas
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.16620v4 Announce Type: replace-cross Abstract: Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity from a single source: stochastic token-level sampling. We argue that this single-axis sampling space is...
666. FusionRS: A Large-Scale RGB-Infrared-Style Remote Sensing Dataset for Cross-Modal Vision-Language Learning ​
Author: Jiaju Han, Ben Zhang, Xuemeng Sun, Qike Zhang, Yuxian Dong, Dingyi Lu, Chengyin Hu, Luwei Yang, Fengyu Zhang, Yiwei Wei, Jiujiang Guo
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.17020v2 Announce Type: replace-cross Abstract: Remote sensing vision-language models have advanced Earth observation, but available large-scale vision-language resources remain RGB-centered, leaving complementary infrared information underexplored. Infrared observations provide distinctiv...
667. Quantum Cinema: An Interactive Cinematic Exploration of Quantum Computing Hardware via Generative World Models ​
Author: Aoyu Zhang, Dongping Liu, Luyao Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: physics.pop-ph, cs.AI, cs.ET, cs.HC, quant-ph
arXiv:2606.17102v3 Announce Type: replace-cross Abstract: Quantum computing promises transformative advances across science and industry, yet the physical hardware that enables these computations remains invisible to the public: quantum processors operate inside sealed dilution refrigerators at temp...
668. DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams ​
Author: Cong Wan, Zeyu Guo, Zijian Cai, Jiangyang Li, SongLin Dong, Lin Peng, Xiangyang Luo, Zhiheng Ma, Yihong Gong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.21337v2 Announce Type: replace-cross Abstract: Raw multimodal streams are abundant but noisy, redundant, and unaligned with any particular training objective. Turning them into supervision today means either brittle heuristics or repeatedly querying a proprietary vision-language model, a ...
669. Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers ​
Author: Xin Gao, Xingming Xu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.21848v3 Announce Type: replace-cross Abstract: Transformer architectures form the foundation of modern natural language processing, yet the Key-Value (KV) cache introduces substantial memory and bandwidth overhead during long-context generation, increasingly bottlenecking large-scale depl...
670. SAGE: An Expert-Annotated South Asian GI Endoscopy Dataset for Multimodal Learning and Hallucination Analysis ​
Author: Niyoj Oli, Sachin Acharya, Sandesh Pokhrel, Sanjay Bhandari, Ramesh Rana, Nikesh Mani Shrestha, Ram Bahadur Gurung, Yash Raj Shrestha, Prashnna K Gyawali, Binod Bhattarai
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.22144v2 Announce Type: replace-cross Abstract: Gastrointestinal cancers represent a growing health burden in the South Asian region, driven largely by rapid changes in socio-economic conditions and lifestyle habits. However, early diagnosis remains limited by inadequate equipment, financi...
671. The Governance Inversion Hypothesis: Why More AI Regulation May Produce Less Organisational Control ​
Author: Victor Frimpong
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2606.26117v2 Announce Type: replace-cross Abstract: This paper introduces the Governance Inversion Hypothesis (GIH) to explain a growing paradox in artificial intelligence (AI) governance: under conditions of increasing regulatory expansion and technological complexity, organisations may becom...
672. The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals ​
Author: Kwan Soo Shin
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV
arXiv:2606.26529v3 Announce Type: replace-cross Abstract: AI in radiology and other safety-critical workflows is evaluated on the hazards it is told to find, yet harm arises disproportionately from hazards no one specified. We show that conditioning a language or vision model on a narrow task suppre...
673. DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training ​
Author: Yiwei Liu, Haoning Wang, Haisen Luo, Dan Liu, Junxi Yin, Haotian Wang, Lei Zhang, Xiaoyu Tian, Shuaiting Chen, Yuansheng Song, Baoyan Guo, Xiongfei Yan, Bolan Yang, Chengwei Liu, Ming Cui, Jiong Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.30345v3 Announce Type: replace-cross Abstract: Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in complex reasoning tasks. Existing self-distillation and reinforcement learning methods lack explicit mechanis...
674. ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ​
Author: Zijun Xie, Binbin Zheng, Enlei Gong, Jihua Liu, Yuyang You, Lingfeng Liu, Jiayao Tang, Guanqun Zhao, Aoqi Hu, Zeyu Chen
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.31650v5 Announce Type: replace-cross Abstract: Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Context-management methods make such rollouts feasible by simplifying past interactions through deletion,...
675. Freeform Preference Learning for Robotic Manipulation ​
Author: Marcel Torne, Anubha Mahajan, Abhijnya Bhat, Chelsea Finn
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2606.32027v3 Announce Type: replace-cross Abstract: Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of ...
676. A Category Theory Account of AI Identity ​
Author: Andrea Ferrario
Published: 8/4/2026, 4:00:00 AM
Categories: math.CT, cs.AI, cs.CY
arXiv:2607.00220v2 Announce Type: replace-cross Abstract: Artificial intelligence (AI) systems are routinely modified after deployment through retraining and changes in their environments. These transformations raise a metaphysical question: under what conditions does an AI system remain the same sy...
677. Adaptive Perturbation Selection for Contrastive Audio Decoding ​
Author: Aaron Isidore Grace, Zhouyuan Huo, Weiran Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2607.00247v2 Announce Type: replace-cross Abstract: Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors. While contrastive decoding (CD) offers training-free mitigation, existing methods rely on blunt perturbations like masking or noi...
678. LCPNet: Latent Consistent Proximal Unfolding Network for Infrared Small Target Detection ​
Author: Tianfang Zhang, Lei Li, Chang Liu, Zhenming Peng, Huaping Zhang, Xiangyang Ji
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.04603v2 Announce Type: replace-cross Abstract: Infrared small target detection (IRSTD) aims to identify long distance small targets from complex infrared backgrounds, and is a fundamental task in remote sensing. Deep learning methods have improved IRSTD by learning discriminative image-to...
679. x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability ​
Author: Xin Peng, Ang Gao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.06114v4 Announce Type: replace-cross Abstract: Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function evaluations (NFEs). This remains a practical challenge for released checkpoints, since many accelerators r...
680. LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis ​
Author: Chenhao Yuan, Yinhao Xu, Shuwen Xu, Xizhi Yang, Jiaxiang Liu, Chenxi Zhou, Shaoping Huang, Haolin Ren, Pengfei Cao, Jun Zhao, Kang Liu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.06160v2 Announce Type: replace-cross Abstract: Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instr...
681. Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph ​
Author: Duen Horng Chau, Donghao Ren, Fred Hohman, Dominik Moritz
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DS, cs.HC
arXiv:2607.08746v3 Announce Type: replace-cross Abstract: While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally. This graph encodes the dat...
682. RSRA: Training-Free Probing of Representation Sensitivity for Efficient LoRA Rank Allocation ​
Author: Jiaqi Liu, Haidong Kang, Qihui Zhao, Guo Yu, Jingchao Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.09757v2 Announce Type: replace-cross Abstract: Parameter-efficient fine-tuning enables large language models to adapt to downstream tasks with substantially lower computational and storage cost, and Low-Rank Adaptation (LoRA) is among its most widely used techniques. However, vanilla LoRA...
683. An Autonomous Scientific Knowledge Generation Framework for AI-Driven Scientific Discovery ​
Author: Dibakar Datta
Published: 8/4/2026, 4:00:00 AM
Categories: cs.DL, cond-mat.mtrl-sci, cs.AI
arXiv:2607.09806v2 Announce Type: replace-cross Abstract: Artificial intelligence (AI) is transforming scientific discovery, but its effectiveness is fundamentally limited by the availability of structured scientific knowledge. Although existing databases have accelerated data-driven materials resea...
684. Geometric mean-based pairwise comparison method with the reference values -- statistical approach ​
Author: Konrad Ku{\l}akowski, Kamil Pustelnik, Jacek Szybowski
Published: 8/4/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, math.ST, stat.TH
arXiv:2607.10038v3 Announce Type: replace-cross Abstract: For many years, pairwise comparison methods have been widely used for eliciting preferences and ranking alternatives in decision-making problems. These methods estimate priority weights from a pairwise comparison matrix and are now standard t...
685. Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference ​
Author: Zebin Yang, Qi Wang, Yunhe Wang, Xiurui Guo, Bo Yu, Shaoshan Liu, Jiafeng Xu, Hao Dong, Meng Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.12659v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-power onboard devices, such as the Jetson Orin, remains challenging due to their high computational compl...
686. Reassessing Muon for Matrix Factorization ​
Author: Ali Parviz, Gal Mishne, Alex Cloninger
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.13246v2 Announce Type: replace-cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its e...
687. StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows ​
Author: Sizhong Qin, Yi Gu, Yao Jiang, Ao Cai, Changjian Zhou, Shaoxuan Shuai, Jiachang Wang, Tianhao Shen, Yueqiang Li, Xinhao Li, Li Zeng, Yueshi Chen, Dachen Gao, Genrong Xu, Wenjie Liao, Xinzheng Lu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.MA
arXiv:2607.14896v2 Announce Type: replace-cross Abstract: Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, applicable engineering checks, ...
688. A cubical formalisation of topos causal models: intervention, forcing, and a contextuality obstruction ​
Author: Karen Sargsyan
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, math.CT
arXiv:2607.15629v2 Announce Type: replace-cross Abstract: Topos causal models recast causal inference inside a topos: a causal world is a presheaf, an intervention is a sub-model named by a characteristic map into the subobject classifier $\Om$, and reasoning is Kripke-Joyal forcing in an intuitioni...
689. Agentic Synthesis against Counterexample-Supplemented Sketches ​
Author: Muness Castle, Eric Rubeck
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.15854v2 Announce Type: replace-cross Abstract: Coding agents can fix a failing example without preserving the domain rule that made it fail. We present agentic synthesis against counterexample-supplemented sketches, a repository-native method for systems whose policy is discovered during ...
690. When Does Muon Help Agentic Reinforcement Learning? ​
Author: Kai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.16169v4 Announce Type: replace-cross Abstract: Muon is competitive with AdamW in large-scale pre-training, but its operating regime in reinforcement-learning post-training remains unclear. We map this regime on ALFWorld, a sparse-reward agentic benchmark, using three group-based objective...
691. Predicted Cortex Is Not a Domain-General Prior: A Matched-Control Audit of Brain-Encoding Features for Video Memorability ​
Author: Carson Rodrigues
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.16292v4 Announce Type: replace-cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio and text well enough to win the Algonauts 2025 challenge. We ask whether their predicted responses, obtained with no scanner, are a useful feature lens for a human-behavi...
692. WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture ​
Author: Renqin Cai, Dawei Sun, Yuanjun Yao, Zhiyong Wang, Velvin Fu, Maggie Zhuang, Yu Shi, Zhongnan Fang, Xuan Cao, Jing Qian, Rui Li
Published: 8/4/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG
arXiv:2607.17017v2 Announce Type: replace-cross Abstract: As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including user, item, context, and ...
693. ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts ​
Author: Pratyush Dhingra, Pramit Kumar Pal, Janardhan Rao Doppa, Partha Pratim Pande
Published: 8/4/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.DC
arXiv:2607.17074v2 Announce Type: replace-cross Abstract: Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling Large Language Models (LLMs). However, MoE inference on conventional hardware is constrained by three fundamental bottlenecks. These encompass the massive ...
694. Retrieval-Augmented Interpretable Learning: Towards Task-Specific Zero-Shot Models in Healthcare ​
Author: Sazan Mahbub, Caleb Ellington, Zhiyuan Li, Yixin Yang, Souvik Kundu, Ben Lengerich, Eric P. Xing
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.17508v2 Announce Type: replace-cross Abstract: We introduce Retrieval-Augmented Interpretable Learning (RAIL), a probabilistic meta-learning framework for zero-shot generation of task-specific interpretable models that synthesizes coefficient-space structure from natural-language task des...
695. CriPO: Enhancing Rubric-based RL via Self-Distillation ​
Author: Mingxuan Xia, Yuhang Yang, Chao Ye, Shuai Zhu, Shenzhi Yang, Guangcheng Zhu, Yuhang Zhang, Cheng Peng, Haobo Wang, Siqing Wang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.18082v3 Announce Type: replace-cross Abstract: Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (Unexplored Criteria, UC) receive no optim...
696. Physical Self-Supervised Learning: IMU Sensing without Manual Labels ​
Author: Yuyang Leng (Richard), Renyuan Liu (Richard), Shaohan Hu (Richard), Peijun Zhao (Richard), Chun-Fu Chen (Richard), Songqing Chen, Shuochao Yao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.18361v2 Announce Type: replace-cross Abstract: Deep neural networks have become a promising approach for IMU-based sensing, but their scalability is fundamentally limited by costly labeled data and poor robustness to heterogeneous devices, placements, and users. Existing unsupervised and ...
697. Generative AI floods and dilutes the market for books ​
Author: Tuhin Chakrabarty, Xinyue Liu, Jane C. Ginsburg, Paramveer Dhillon
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2607.20349v3 Announce Type: replace-cross Abstract: Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-quality ``slop'' that buyers will ignore, and are assumed to carry little commercial weight. We test that assumption with full-te...
698. Robostral Navigate ​
Author: Abdelaziz Bounhar, Abhijeet Somani, Aditi Kabra, Adrian Valente, Adrien Petralia, Adrien Sade, Alan Jeffares, Albert Jiang, Aleksandr Timashov, Alexandre Cahill, Alexandre Gavaudan, Alexandre Laval, Alexandre Sablayrolles, Amelie Heliou, Amos You, Andre Jonasson, Andrew Bai, Andrew Ehrenberg, Andrew Zhao, Angele Lenglemetz, Anmol Agarwal, Antonia Calvi, Arata Suzuki, Arjun Majumdar, Arthur Fournier, Artjom Joosen, Avinash Sooriyarachchi, Aylin Guliz Akkus, Aysenur Karaduman, Baptiste Bout, Baptiste Roziere, Baudouin De Monicault, Benjamin Holzschuh, Benjamin Lefaudeux, Benjamin Tibi, Bernhard Stadlbauer, Blazej Osinski, Camille Le Scao, Chaoran Yu, Charlotte Cronjager, Chen-Yo Sun, Chris Bamford, Christian Wallenwein, Christophe Renaudin, Clemence Lanfranchi, Corentin Barreau, Corentin Sautier, Cristiana-Diana Diaconu, Cyprien Courtot, Daniel Marczak, Darius Dabert, Diego de Las Casas, Dominik Nuss, Dylan Rubini, Dzmitry Soupel, Elizaveta Demyanenko, Elliot Chane-Sane, Emilien Fugier, Emmanuel Gottlob, Erik Aas, Etienne Goffinet, Etienne Millon, Eujeong Choi, Fabian Paischer, Fabian Schlager, Faruk Ahmed, Federico Baldassarre, Filip Szatkowski, Florian Wiesner, Gabrielle Berrada, Gaetan Ecrepont, Gaetan Lepage, Gaspard Blanchet, Gaspard Donada-Vidal, Gauthier Delerce, Gauthier Guinet, Genevieve Hayes, Georgii Novikov, Giada Pistilli, Gianluca Galletti, Guillaume Breton, Guillaume Kunsch, Guillaume Lample, Guillaume Martin, Guillaume Raille, Gunjan Dhanuka, Gunshi Gupta, Han Zhou, Harshil Shah, Hasan Furkan Vural, Hedi Hadiji, Hope McGovern, Hugo Cisneros, Hugo Thimonier, Indraneel Mukherjee, Ivan Cuevas Salazar, Jacques Sun, Jan Ludziejewski, Jason Rute, Jean Quentin, Jean-Hadrien Chabran, Jean-Malo Delignon, Jie Zhang, Joachim Studnia, Joep Barmentlo, Johannes Brandstetter, John Harvill, Jonas Amar, Jonas Schweizer, Josephine Delas, Josselin Somerville, Julien Denize, Julien Tauran, Kartik Khandelwal, Khyathi Raghavi Chandu, Kilian Tep, Kush Jain, Larissa Laich, Laura Calem, Laurence Aitchison, Laurent Callot, Laurent Fainsin, Leo Cotteleer, Leonard Blier, Lingxiao Zhao, Louis Martin, Louis Serrano, Lucile Saulnier, Ludovic Ho Fuh, Luis Montero, Maarten Buyl, Manon Chossegros, Marcin Mozejko, Margaret Jennings, Markus Hennerbichler, Martin Alexandre, Mathieu Poiree, Mathieu Schmitt, Mathilde Guillaumin, Matthieu Andre, Matthieu Dinot, Matthieu Futeral, Maurits Bleeker, Mauro Comi, Max Mynter, Maxim Berman, Maxime Darrin, Maxime Louis, Maximilian Augustin, Maximilian Muller, Melina Jingting Laimon, Mert Unsal, Mia Chiquier, Michael Pilcer, Michal Pietruszka, Michal Zajac, Mikhail Biriuchinskii, Minh-Quang Pham, Minwoo Kang, Morgane Riviere, Namit Katariya, Nathan Grinsztajn, Nathan Simpson, Neeraj Aggarwal, Neha Gupta, Ola Mysiak, Oliver Leicht, Olivier Bousquet, Olivier Duchenne, Parag Jain, Patricia Wang, Patrick Blies, Patrick von Platen, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Pavan Kumar Reddy, Pavel Kuksa, Philippe Pinel, Philomene Chagniot, Pierre Stock, Pierre-Andre Savalle, Piotr Milos, Prateek Gupta, Pravesh Agrawal, Quentin Desreumaux, Quentin Torroba, Quercus Hernandez, Ram Ramrakhya, Randall Isenhour, Ranjit Parva, Raul Perez Pelaez, Reinhard Sonnleitner, Remi Delacourt, Richard Kurle, Rishi Shah, Rob Romijnders, Rohin Arora, Romain Sauvestre, Roman Soletskyi, Rosalie Millner, Rupert Menneer, Sagar Vaze, Samuel Barry, Samuel Belkadi, Samuel Humeau, Sanchit Gandhi, Sandeep Subramanian, Sarthak Mittal, Saskia Adaime, Sean Cha, Sebastian Kaltenbach, Shashwat Dalal, Shashwat Verma, Sherif Waly, Shrimai Prabhumoye, Siddhant Waghjale, Siddharth Gandhi, Simon Lepage, Simon Sorg, Soham Ghosh, Sophie Marbach, Srijan Mishra, Stanislas Lange, Steve Hong, Sumukh Aithal, Szymon Antoniak, Tarun Kumar Vangani, Teven Le Scao, Theo Cachet, Thibaut Lavril, Thomas Chabal, Thomas Coste, Thomas Defard, Thomas Foubert, Thomas Robert, Thomas Wang, Tianyu Zhang, Tim Lawson, Timothee Lacroix, Tobias Kronlachner, Tom Bewley, Tom Edwards, Tomas Hodan, Tuhin Das, Tyler Wang, Ulrick BLE, Umar Jamil, Umberto Tomasini, Valentin Mace, Van Phung, Vedant Nanda, Victor Jouault, Victor Letzelter, Victor Paltz, Victor Poucheret, Vincent Maladiere, Vincent Pfister, Virgile Richard, Vladislav Bataev, Wassim Bouaziz, Wen Ding Li, William Havard, William Marshall, Xinghui Li, Xingran Guo, Xinyu Yang, Yann Dreze, Yassine El Ouahidi, Yassir Bendou, Yihan Wang, Yimu Pan, Yves Martin des Taillades, Zaccharie Ramzi, Zhenlin Xu, Zsofia Csakany
Published: 8/4/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.20785v3 Announce Type: replace-cross Abstract: Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, ...
699. Interaction Dynamics Modeling and Predictive Control for Safe Steerable Catheter--Tissue Interaction ​
Author: Yongyan Cao
Published: 8/4/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.RO, cs.SY
arXiv:2607.20939v2 Announce Type: replace-cross Abstract: Safe steerable catheter control is fundamentally a problem of interaction dynamics: the tip must follow a planned motion, remain compliant against moving tissue, reject friction and hysteresis, and respect a clinically meaningful never-exceed...
700. CEL: Comprehensive Counterfactual Explanations Library and Benchmark ​
Author: Oleksii Furman, {\L}ukasz Lenkiewicz, Marcel Musia{\l}ek, Maciej Zi\k{e}ba
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22045v2 Announce Type: replace-cross Abstract: Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome. While early methods primarily focuse...
701. Understanding Machine Unlearning Through the Lens of Mode Connectivity ​
Author: Jiali Cheng, Hadi Amiri
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.23970v2 Announce Type: replace-cross Abstract: Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress, the loss landscape and optimization geometry of unlearning are poorly understood. In this paper, we stu...
702. What EEG Foundation Models Encode: Dataset Identity and a Negative-Control Suite for Clinical Benchmarks ​
Author: Marzieh Zare
Published: 8/4/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE
arXiv:2607.24519v2 Announce Type: replace-cross Abstract: Pretrained EEG foundation models are proposed for clinical decoding, but whether reported gains transfer across populations or survive negative controls is unclear. We benchmark LaBraM, EEGMamba, CBraMod, REVE, LEAD, BENDR, and BIOT on five c...
703. Cross-Cohort Spectral-Temporal Dissociation in Frozen EEG Foundation-Model Representations ​
Author: Marzieh Zare
Published: 8/4/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.ET, cs.LG
arXiv:2607.24834v2 Announce Type: replace-cross Abstract: Objective. We tested whether frozen representations from five EEG foundation models support decoding of long-range temporal correlations, measured as the detrended-fluctuation-analysis (DFA) exponent of the alpha-band amplitude envelope. Appr...
704. Specula: Scaling formal specifications for autonomous model checking of system code ​
Author: Qian Cheng, Saad Mohammad Rafid Pial, Ruize Tang, Yiming Su, Emilie Ma, Finn Hackett, Ivan Beschastnikh, Yu Huang, Tianyin Xu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.DC, cs.OS
arXiv:2607.25333v2 Announce Type: replace-cross Abstract: Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the specifications for highly effective model checking and bug finding. Specula employs large language model (LL...
705. Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review ​
Author: Binyan Xu, Xilin Dai, Fan Yang, Kehuan Zhang
Published: 8/4/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.LG
arXiv:2607.27209v2 Announce Type: replace-cross Abstract: Peer review at ML conferences increasingly relies on reviewer scores as the primary decision instrument. As submissions have scaled from thousands to tens of thousands per year, no systematic audit has examined whether this instrument functio...
706. Specification-Guided Synthesis of Deadlock-Free Communication Protocol Refinements with Large Language Models ​
Author: Yang Li, Ping Hou, Nobuko Yoshida
Published: 8/4/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.27964v2 Announce Type: replace-cross Abstract: Ensuring behavioural correctness in communication protocols is a central challenge in distributed software systems, as subtle inconsistencies can lead to deadlocks. In such settings, protocol refinement - the safe substitution of a protocol t...
707. Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models ​
Author: Shuyi Fan, Boyuan Deng, Mengyu Xu, Jiale Liu, Hongyang Zhang, Qiaoxin Yang, Chongyang Gao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2607.28128v2 Announce Type: replace-cross Abstract: LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conver...
708. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding ​
Author: Mingkang Dong, Muxin Pu, Jie Li, Bohan Guo, Songruo Chen, Bin Ren, Xu Zheng, Chen Zhao, Tianwen Qian, Mohamed Elhoseiny, Yuqian Fu
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.28312v2 Announce Type: replace-cross Abstract: Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches primarily manage the growing visual context according to token importance, temporal redundancy,...
709. SCMA: Structure-Conditioned and Metal-Aware Flow Matching for CT Metal Artifact Reduction ​
Author: Heran Wang, Jianing Sun, Xu Jiang, Genwei Ma, Jigang Duan, Xing Zhao
Published: 8/4/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.28759v2 Announce Type: replace-cross Abstract: In X-ray CT, metallic objects cause beam hardening, photon starvation, and scattering, leading to projection inconsistency, streaks, dark bands, and structural distortions that compromise clinical diagnosis and quantitative analysis. Existing...
710. The persuasive power of large language models does not depend on their perceived national origin ​
Author: Ningzhi Liu, Yannic Hinrichs, Jonas R. Kunst
Published: 8/4/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2607.29334v2 Announce Type: replace-cross Abstract: Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rejected as foreign propaganda, with consequences for democratic discourse and information sovereignty. Ye...