arXiv cs.AI - 2026-08-10 ​
295 items collected.
1. Towards Multi-Label Graph Foundation Models: from Single-Vector Representation Learning to Multi-Semantic Basis Learning ​
Author: Dongxiao He, Jiayu Zhang, Jitao Zhao, Yi Wang, Di Jin
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06394v1 Announce Type: new Abstract: Multi-label node classification is an important yet challenging task in graph learning, where nodes exhibit multiple semantics simultaneously. Existing methods for multi-label node classification can effectively model multiple labels, while only consid...
2. EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs ​
Author: Bo Liu, Muxuab Yu, Yu Zhang, Pengfei Gao, Yongping Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06398v1 Announce Type: new Abstract: Recent byte-level large language models (LLMs) have made tokenizer-free modeling increasingly competitive by grouping bytes into dynamically sized patches. However, existing byte-patch architectures still apply the same dense feed-forward computation t...
3. Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast ​
Author: Yifan Wang, Jinyi Mu, Mayank Jobanputra, Yu Wang, Soyoung Oh, Isabel Valera, Vera Demberg
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06400v1 Announce Type: new Abstract: Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging. Recent sparse Mixture-of-Experts (MoE) reward models seek to improve interpretability by routing prompts to specialized exp...
4. Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes ​
Author: Aoting Zeng, Kai Wang, Jianwei Wang, Yuxiang Sun, Yizhang He, Wenjie Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.06402v1 Announce Type: new Abstract: Community detection is a fundamental task in graph analytics that aims to identify cohesive groups of entities with similar behaviors or interests. Classic objective-driven methods struggle with complex graph structures, while deep-learning approaches ...
5. ADIAS: Automated Design of Interactive Agentic Systems ​
Author: Lekang Jiang, Bohan Tang, Stephan Goetz, Yiwen Guo
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA
arXiv:2608.06410v1 Announce Type: new Abstract: Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents, which leaves the repair pro...
6. Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ​
Author: Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang, Hao Geng, Minjun Yu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.06411v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual tokens. Visual token pruning can reduce this cost, but requires accurate...
7. WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader ​
Author: Boshui Chen, Huiping Liu, Shaolei Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approach to closing their remaining functional gap. This training regime is bottlenecked by reward design....
8. Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding ​
Author: Ming Wang, Yuqing Zhang, Tingna Xie, Xiangju Li, Xiaocui Yang, Daling Wang, Shi Feng, Yifei Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MM
arXiv:2608.06501v1 Announce Type: new Abstract: Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. Cross-concept underst...
9. KNOWPLAN: Knowledge-Driven AI Agents for Smart Degree Pathway Planning ​
Author: Shuheng Cao, Weijia Zhang, Jiaqi Wu, Xiyun Hu, Yat Yang, Juqy Chen, Zhaoxiang Feng
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06530v1 Announce Type: new Abstract: Planning a degree from official university sources requires solving two problems in order. The institution's curriculum must first be reconstructed from catalogs, departmental pages, JSON endpoints, and PDFs that share no schema, and only then can a st...
10. TaskSense: Focusing on What Matters in World Models ​
Author: SM Mazharul Islam, Manfred Huber
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2608.06544v1 Announce Type: new Abstract: World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve information across the entire visual input. However, task-relevant content often occupies only a sm...
11. Divergent Response Modes in Frontier Language Models Under Steering Pressure ​
Author: Ali Jalal-Kamali
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06578v1 Announce Type: new Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study evaluates behavioral steerab...
12. Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques ​
Author: Hotaka Maeda, Yikai Lu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06609v1 Announce Type: new Abstract: Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation. We aimed to build a near-comprehensive AIE model by predicting item...
13. NxN E-valuation: Hypothesis Certification via a Conformal CRT Null ​
Author: Bin Wang, Yan Zhong
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06621v1 Announce Type: new Abstract: We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a hypothesis be verified without building any case-specific certification procedure---such as constructing a dedicated null hypothesis---as long as a large ...
14. Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation ​
Author: Ziyun Xu, Bosen Ding, Yue Zhang, Ji Qi, Qingyuan Song, Jizhou Huang, Liwei Wang, Jefferey Santelli, Yue Weng, Qichao Que, Zhenheng Yang, Junfeng Pan, Linhong Zhu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06632v1 Announce Type: new Abstract: Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.g., clicks, dwell time) rather than explicit, natural language inputs. As a result, users experience a per...
15. TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure ​
Author: Joshua Zuniga, Srinivasan Subramanian, Ramya Madhuri Narapureddy, Md Abdullah Al Hafiz Khan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2608.06657v1 Announce Type: new Abstract: Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustworthiness depends on the whole loop, not any one model. Yet no standard benchmark captures time-alig...
16. CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models ​
Author: Haiping Liu, Qian Zhao, Lijing Lin, Jingyuan Sun, Hongpeng Zhou
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06659v1 Announce Type: new Abstract: This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression value...
17. Vehicle routing problem using deep reinforcement learning - A case study about truck planning in the industry ​
Author: Siliang Lu, Dan Hu, Lili Wu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06668v1 Announce Type: new Abstract: As an important component of the supply chain industry, transportation has experienced rapid development in the past decade with the assistance of digital platforms and intelligent algorithms. Within the field of transportation research, Vehicle Routin...
18. A Multi-Agent Framework for Automated Coarse-Grained Molecular Dynamics of Polymers ​
Author: Joohee Choi, Junhyeong Lee, Seunghwa Ryu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.06694v1 Announce Type: new Abstract: Coarse-grained (CG) molecular dynamics extends polymer simulation beyond the scales accessible to all-atom (AA) methods, but bottom-up CG modeling is laborious. The CG resolution is a design choice, so a transferable parameter set is generally not avai...
19. AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models ​
Author: Zibo Shao, Baochen Xiong, Chengdong Xu, Linhui Xiao, Kaichen Li, Haoran Gong, Yan Li, Yaguang Song, Xiaoshan Yang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.06699v1 Announce Type: new Abstract: Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are specialized for particular tools or environments, complicating consolid...
20. WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance ​
Author: Zhi Li, Tao Zhou, Yeqing Li, Eugene Ie, Demetri Terzopoulos
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2608.06704v1 Announce Type: new Abstract: Delegating a web task involves more than asking a question; it requires transferring a policy: what to verify, how to handle uncertainty, which preferences matter, and when to stop. Yet, current live-web agents are evaluated solely on the final answer,...
21. MolBioKG: Grounding Out-of-Graph Molecules in Biomedical Knowledge Graphs via Multi-Resolution Structural Anchoring ​
Author: Yiming Zhang, Hikaru Shindo, Shuan Chen, Kaushalya Madhawa, Jun Jin Choong, Yuna Oikawa, Takashi Fujiwara, Keisuke Ozawa
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.06713v1 Announce Type: new Abstract: Biomedical knowledge graphs (KGs) accelerate drug discovery, but standard pipelines assume query molecules already exist as graph entities, leaving unregistered molecules disconnected. We address this cold-start challenge, termed the out-of-graph molec...
22. The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows ​
Author: Junbo Li, Boyi Liu, Canwen Xu, Yite Wang, Yuxiong He, Zhangyang Wang, Qiang Liu, Zhewei Yao
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06714v1 Announce Type: new Abstract: Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search po...
23. bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning ​
Author: Koushik Howlader, Tirtho Roy, Md Tauhidul Islam, Wei Le
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.06727v1 Announce Type: new Abstract: Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. Mixture-of-Recursions (MoR) improves efficiency through adaptive token-choice or expert-choice routing. We ...
24. From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos ​
Author: Yifeng Luo, Yupeng Li, Liang Lan, Tian Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.MM
arXiv:2608.06732v1 Announce Type: new Abstract: Recent text-to-video (T2V) generation models enable fake news videos to be synthesized from scratch, shifting the threat beyond cheap fakes assembled from existing footage. Such news videos can closely match fabricated narratives, creating a modality a...
25. IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents ​
Author: Senhao Wang, Chenghao Cai, Haitao Hu, Mingxing Huang, Xingguang Wang, Wenhao Li, Zecheng Lin
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.06735v1 Announce Type: new Abstract: Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as mathematical reasoning and code execution. In these settings, the environment follows fixed rules an...
26. MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents ​
Author: Zhisheng Chen, Bingfan Zeng, Bangde Cao, Zhengwei Xie, Yuxuan Li, Jinhan Li, Zheng Lu, Xiangchen Guan, Zikai Xiao, Rui Qian, Jingwei Song
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06745v1 Announce Type: new Abstract: Long-horizon agents rely on memory to reuse experiences, yet existing memory systems often assume that evidence can be directly consumed through a fixed representation. This leads to representation mismatch, where relevant information is available but ...
27. Mind the Gap: A Dual Knowledge Graph Framework for Unified Multi-task User Intent Inference ​
Author: Tzu-Cheng Peng (National Taiwan University), Chien Chin Chen (National Taiwan University), Chih-Hao Ku (University of North Texas), Yung-Chun Chang (Taipei Medical University)
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.06752v1 Announce Type: new Abstract: This paper proposes DKG-MTI, a dual knowledge graph framework for unified multi-task user intent inference from online travel reviews. Existing approaches often rely on hierarchical pipelines that suffer from error propagation or retrieval methods that...
28. Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence ​
Author: Ying Chen, Weizhen Li, Zhe Hu, Zhenjiang Li, Rui Jiang, Zhifeng Gu, Lihuang Fang, Jiangping Liu, Lei Yi, Jie Chen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06756v1 Announce Type: new Abstract: Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verif...
29. LiFTER: A Grounded Neuro-Symbolic Microscope for Continuous-Time Dynamic Graph Forecasting ​
Author: Minwoo Yu, Young-guk Ha
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06765v1 Announce Type: new Abstract: Continuous-time dynamic graph models predict future links by compressing past interactions into neural states. Although effective for forecasting, this computation obscures which entities are shared across events and how temporal patterns contribute to...
30. Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts ​
Author: Rulin Zhou, Wanhao Liu, Guoheng Ma, Liangjin Shao, Qiujie Song, Yidu Wang, Guankun Wang, Tong Chen, Long Bai, Luping Zhou, Hongliang Ren
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.06770v1 Announce Type: new Abstract: Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions. However, existing methods lack a unified multimodal control paradigm, ...
31. Evolving Parallel Algorithm Portfolios via Potential-Aware Instance Generation with LLMs ​
Author: Shaofeng Zhang, Shengcai Liu, Zhiyuan Wang, Ke Tang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06808v1 Announce Type: new Abstract: The Automatic Construction of Portfolios via Large Language Models (LLM-ACP) suffers from poor generalization in practical few-shot scenarios when solving complex combinatorial optimization problems. Instance and algorithm co-evolution frameworks addre...
32. Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents ​
Author: Hongxi Yan, Ziyue Huang, Shichao Fan, Qingjie Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06861v1 Announce Type: new Abstract: Training large language model agents in long-horizon environments requires assigning credit from sparse terminal outcomes to individual actions. Existing critic-free methods propagate trajectory-level rewards uniformly across steps, while recent approa...
33. CEDAR: Agent-Orchestrated Tree Search for Goal-Directed Optimization of Complex Systems ​
Author: Yingtao Tian
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06871v1 Announce Type: new Abstract: Complex systems, core objects of study in artificial life, model diverse phenomena through nonlinear, feedback-driven interactions that produce emergent behavior, with applications from population dynamics and biology to economic policy and strategic d...
34. SkillEval: Decomposing Agent Skill Quality into Interpretable Signals ​
Author: Jiahui Han, Qinuo Li, Ziheng Peng, Haotian Wu, Haoze Liu, Danfeng Shan, Guanchu Wang, Huiqi Deng, Ninghao Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06891v1 Announce Type: new Abstract: Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill quality becomes increasingly important. Existing evaluations often measure skill quality by testing whether a skill imp...
35. From Points to Edges: Edge-Conditioned Spectral Operators for Physics-Sensitive PDE Learning ​
Author: Zhentao Tan, Ruijie Quan, Yi Yang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.06894v1 Announce Type: new Abstract: Neural operators have become a central tool for solving partial differential equations (PDEs), with spectral operators offering efficient global mixing across spatial locations. However, many PDEs contain physics-sensitive local structures that are cri...
36. Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework ​
Author: Jing Chen, Yang Sun, Li Zhang, Lin Xu, Jie Shi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06909v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provide limited support fo...
37. Fast LapSum: Exact Differentiable Top-k at Million Scale ​
Author: {\L}ukasz Struski, Joanna Wojciechowicz, Jakub Antczak, Marcin Mazur, Kamil Ksi\k{a}.zek, Jacek Tabor
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06912v1 Announce Type: new Abstract: The top-$k$ operation is a fundamental building block of modern sparse computation, enabling token routing, expert activation, memory selection, and attention pruning. Yet standard hard top-$k$ blocks gradients, while existing continuous (soft) relaxat...
38. ReGraph: Learning to Generate Recipe Graphs from Food Images ​
Author: Guoshan Liu, Bin Zhu, Pengkun Jiao, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06917v1 Announce Type: new Abstract: Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.However, cooking is a structured transformation process in which ingredients undergo state changes through ordered actions,while free-form ...
39. Deal Me Maybe: The Role of Emotions in Multi-Agent Negotiation ​
Author: Massimiliano Luca, Apoorva Singh, Bruno Lepri
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06922v1 Announce Type: new Abstract: Negotiation is a demanding social task for LLM agents, requiring strategic reasoning, persuasion, and interpersonal adaptation. Yet existing benchmarks often treat agents as emotionally neutral, overlooking a key driver of human bargaining behavior. We...
40. TRIBE: Predicting Team Performance via Communication Behavior Ensembles ​
Author: Ali Jalal-Kamali, Nikolos Gurney, David V. Pynadath, Fred Morstatter
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06926v1 Announce Type: new Abstract: Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present TRIBE, a domain independent approach that reveals team behavioral dynamics invisible to traditional...
41. Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery ​
Author: Taolin Han, Yuchen Zhang, Jinghang Wang, Yun Wu, Wai Yuet Chiu, Zhaohai Li, Yifei Zhang, Jinxin Wang, Yuhao Zhou, Chen Zhao, Jiajia Li, Jiaxin Li, Qile Jin, Kewei Sun, Shuang Wu, Weiqi Zhai, Renquan Lv, Junchao Li, Ruodan Chen, Qingteng Chen, Zhibo Yang, Hu Wei, Lin Qu, Shuai Bai, Bing Zhao
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06931v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated q...
42. Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps ​
Author: Yang Shu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06940v1 Announce Type: new Abstract: LLM judge panels are a standard evaluation tool, but prior work reports highly correlated panel errors: nine judges provide roughly the effective information of two independent ones, and aggregation closes only a small fraction of the gap. A natural re...
43. LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents ​
Author: Ivan Majic, Zexian Huang, Franziska H"ubl, Krzysztof Janowicz, Meilin Shi, Mina Karimi, Zilong Liu, Alexandra Fortacz-Lazan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06948v1 Announce Type: new Abstract: AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial tasks and workflows. However, most of the research on their spatial capabilities (e.g., spatial reaso...
44. Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints ​
Author: Paul-Peter Arslan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06949v1 Announce Type: new Abstract: Prior benchmarking work has shown that a single large language model (LLM), forced to make life-or-death resource-allocation decisions, exhibits measurable demographic bias. Real deployments, however, rarely use a single agent: they use pipelines, with...
45. Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation ​
Author: Jonghyun Jee, Aaron Shaw
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs systematically reproduce evaluative hierarchies remains unclear. Prior research on cultural bias in LLM...
46. CAi Copilot: Reducing Operational Workload in Molecular Design through Intent-Driven Agentic Workflows ​
Author: Zhu Wang, Jiangyu Chen, Yingjun Shang, Yuhui Yao, Laiao Lu, Tianfan Fu, Na Zou
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.06961v1 Announce Type: new Abstract: Early-stage molecular design is an iterative process, not just a task of generating molecules. Researchers turn broad goals into design strategies, refine candidates, assess many properties, and gather evidence before synthesis and tests. AI methods ca...
47. Learning in Deep Networks under Dale's Constraint ​
Author: Roy Abel, Shimon Ullman
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06963v1 Announce Type: new Abstract: Biologically plausible learning models aim to explain how neural circuits can implement effective learning under the constraints of real neurons. Although significant progress has been made, a major remaining challenge is that existing models often all...
48. Finding Usable Weight Mechanisms with Tiled SVD ​
Author: Ash Manvi, Samreena Tajreen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06969v1 Announce Type: new Abstract: The dominant approach to mechanistic interpretability trains proxy dictionaries such as sparse autoencoders and labels features from max-activating text. The best such atlases identify con- cepts, but that identity lives in the learned dictionary rathe...
49. FedLBW: A Loss-Based Weighting Strategy for Federated Learning on Non-IID Data in Wireless Networks ​
Author: Majid Kundroo, Tinku Singh, Taehong Kim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.ET, cs.LG
arXiv:2608.07007v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative machine learning (ML) across distributed clients while preserving privacy. However, efficient model convergence in FL remains challenging, especially in wireless networks where non-independent and identical...
50. ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization ​
Author: Yongge Ma, Guoan Wang, Feiyu Wang, Yaoming Li, Qian Zhang, Zihan Yan, Yinjun Han, Tong Yang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07019v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models. Existing PTQ methods typically obtain an initial quantized model through heuristic rules or greedy optimization, and once quantization...
51. ZIPBrain: Can EEG Foundation Models Be Faster, Locally Deployable, but Accurate? ​
Author: Lingwei Li, Yirong Kan, Peng Chen, Xu Cao, Zheng Chen, Yasuhiko Nakashima
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07033v1 Announce Type: new Abstract: This work investigates whether Electroencephalograph (EEG) foundation models (EFMs) can be made faster and locally deployable without sacrificing accuracy. EEG foundation models are a major trend, offering strong general-purpose representations. Howeve...
52. Not All Problems Are Best Modeled as MILP: A DSL-Centric Framework for Flexible and Accurate Optimization Modeling ​
Author: Shaofeng Zhang, Hongyuan Su, Qingwen Peng, Zefang Zong, Shengcai Liu, Ke Tang, Yong Li
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07040v1 Announce Type: new Abstract: Solving combinatorial optimization problems (COPs) requires not only efficient algorithms but also carefully crafted formulations. While recent works have leveraged LLMs to automate optimization modeling, current frameworks predominantly rely on a rigi...
53. Unsupervised Adaptation of PDE Foundation Models ​
Author: Ziye Song, Zhao Wei, Xin Yu, Ivor Tsang, Yueming Lyu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07053v1 Announce Type: new Abstract: Pretrained partial differential equation (PDE) foundation models can generalize across different equations, but adapting them to unseen PDE systems typically requires dense solution data, which is often expensive or unavailable. To address this limitat...
54. BONSAI: Evolvability-Guided Tree Search over Skills ​
Author: Yash Priya Shastri, Anand Eswaran, Adnan Qidwai, Pankaj Thorat, Sachin Joshi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07056v1 Announce Type: new Abstract: A skill is a naturallanguage document that steers a frozen agent whose weights cannot be updated so any capability the agent lacks must be supplied in prose Optimising a skill is therefore optimising text against a score and the standard recipe which k...
55. PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks ​
Author: Hui Xie, Tong Shi, Haotong Qin, Aishan Liu, Xiaode Liu, Jinyang Guo
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07066v1 Announce Type: new Abstract: Spiking neural networks (SNNs) enable sparse and event-driven computation, but their low-bit deployment remains incomplete because recurrent membrane states are commonly retained in floating point even after weight quantization. Quantizing these states...
56. DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding ​
Author: Hanshu Yao, Janfeng Zhong, Niu Lian, Jinpeng Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, cs.MM
arXiv:2608.07067v1 Announce Type: new Abstract: Long-document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limited by static retrieval and fragile cross-round memory. Mainstream single-round methods commit to a fixed top-$k$ ...
57. MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents ​
Author: Zhiyuan Liu, Tinghong Ye, Chenghao Liu, Yizhuo Li, Songfang Huang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07068v1 Announce Type: new Abstract: Long-horizon agents accumulate growing contexts during interaction, impairing performance and stability. Compact memory mitigates this problem by compressing and rewriting the history retained between model invocations. Learning what to retain typicall...
58. Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking ​
Author: Devin Pereira, Willem Zuidema
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.07077v1 Announce Type: new Abstract: The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). Current models solve the standard formulation of the puzzle, but still struggle with the flat-to-flat variant (where initial and...
59. MemWM: Memory-Augmented Text-Based World Model ​
Author: Yujun Wang, Tao Zhang, Jinhe Bi, Aniri, Wenxuan Ye, Boliang Liu, Sikuan Yan, Shuning Wang, Xuebing Zhou, S"oren Pirk, Hinrich Sch"utze, Yunpu Ma
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07107v1 Announce Type: new Abstract: World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent next-state predictions can still omit task-critical facts, corrupt product attributes, or apply incorr...
60. How Much, Then Where: Credit-Conserving Action-to-Token Allocation for Multi-Turn Agent Reinforcement Learning ​
Author: Lichao Ma, Yang Sun, Shuaitao Zhao, Yangyi Fang, Cong Qin, Xiaoliang Fu, Yuhang Tian, Yuchen Wei, Junbo Zhu, Yang Wei, Lu Pan, Jiaye Lin
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07118v1 Announce Type: new Abstract: Credit assignment in multi-turn agent reinforcement learning operates at two levels: assigning trajectory-level credit to actions and distributing each action's credit across its tokens. In this paper, we introduce FACTOR, which separates these decisio...
61. DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training ​
Author: Xucong Wang, Zhe Zhao, Liheng Yu, Di Wu, Xiaofeng Cao, Pengkun Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07147v1 Announce Type: new Abstract: Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from compilation and tests provides objective verification. However, unlike agent tasks, coding agents face...
62. A MARL Centered Reference Architecture for Large Language Model Augmentation in Smart Manufacturing ​
Author: Fouad Bahrpeyma, Dirk Reichelt
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07148v1 Announce Type: new Abstract: Modern manufacturing imposes six coupled demands on adaptive control: local decisions with global consequences, partial observability, nonstationarity, reflex speed response with long horizon effects, delayed and diffuse outcomes, and dynamics that res...
63. NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs ​
Author: Aditya Katkar, Om Karkele, Kartik Mandhane, Manisha More, Yash Kashid
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it shouldn't. Prompt injection, hallucinated reasoning, and unsafe tool calls form the primary attack surface...
64. Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory ​
Author: Taeil Kim, Kangsan Kim, Sung Ju Hwang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.07169v1 Announce Type: new Abstract: Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which struggle to generate sufficient successful trajectories on their own. We propose Agent Memory Distillatio...
65. SetEasy: A Multi-Modal Classroom Engagement Assessment and Seating Optimization Framework ​
Author: Zhihao Xie, Hongye Yang, Shien Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07188v1 Announce Type: new Abstract: SetEasy optimizes classroom engagement in fixed seating grids. It fuses multimodal sensing (wristband physiology, 4K video, environmental data) and trains a v-Gage model grounded in a revised ISEQ. Each week, two-week engagement forecasts are mapped to...
66. EMAS: Stabilizing Multi-Agent System Evolution through Evidence-Guided Revision ​
Author: Chao Fei, Qingyi Si, Kaihua Liang, Yanghua Xiao, Panos Kalnis, Hongcheng Guo
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07196v1 Announce Type: new Abstract: Many methods for automated multi-agent system design optimize prompts and topologies during an initial design stage and then deploy the resulting system unchanged on subsequent samples. Experience from these samples is rarely consolidated into reusable...
67. Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs ​
Author: Milan Markovic, Goutham Indukuri, Somayajulu Sripada, Colby J. Vorland, Jack Wilkinson, Clare Robertson, Mark Bolland, Andrew Grey, Miriam Brazzelli, Alison Avenell
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07202v1 Announce Type: new Abstract: Systematic reviews of Randomised Controlled Trials (RCTs) are routinely used as evidence for clinical care guidelines. Such evidence has to meet high research integrity standards to prevent low quality or false research outputs influencing the clinical...
68. Beyond the Black Box: Interpretable Models of Human Randomisation Failures ​
Author: Ngoc Linh Dao
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07220v1 Announce Type: new Abstract: Mixed strategy equilibrium predicts i.i.d play: past actions should not help predict future decisions. Human players, however, systematically depart from this benchmark, and in O'Neill's zero sum card game, these departures can be predicted by black bo...
69. From probability to causality in probabilistic logic programming ​
Author: Zora Wurm, Kilian R"uckschlo{\ss}, Felix Weitk"amper
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07230v1 Announce Type: new Abstract: Probabilistic logic programming is a formalism of statistical relational artificial intelligence that supports causal queries, including interventions from outside the system. When the structure of a probabilistic logic program is learned from data, ho...
70. Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models ​
Author: Rens Anderson, Tessa Verhoef, Amirhossein Zohrehvand
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.NE
arXiv:2608.07243v1 Announce Type: new Abstract: Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM creativity by adapting...
71. WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN ​
Author: Yuehao Huang, Yunzi Wu, Xiaotao Zhang, Xinhai Li, Jiankun Dong, Jiajun Lv, Chi Zhang, Chenjia Bai, Yong Liu, Xuelong Li
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.RO
arXiv:2608.07267v1 Announce Type: new Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions. Although s...
72. Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons ​
Author: Guilin Zhang, Kai Zhao
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, math.ST, stat.TH
arXiv:2608.07303v1 Announce Type: new Abstract: Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and they are easy to get wrong. We report a case study in which a simple AutoML engine, Orcetra, appeared t...
73. An End-to-End Agent Auditing Engine ​
Author: Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains. The fast-evolving harness ecosystem has also made rigorous capability evaluation increasingl...
74. QFCQT: A Chaotically Gated Quantformer Framework for Volatile Time-Series Forecasting ​
Author: Junkai Lin, Siqi Hou, Raymond Lee
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07363v1 Announce Type: new Abstract: Forecasting non-stationary time series remains difficult due to long-range dependencies, local volatility bursts, structural shifts, and nonlinear oscillatory behaviors. Although Transformer-based forecasters are effective for modeling long-term tempor...
75. Curriculum as Code: An AI-Assisted Architecture for Instructional Design in STEM Education ​
Author: Henrique Mohallem Paiva
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.07364v1 Announce Type: new Abstract: Contribution: This paper presents a six-phase AI-assisted instructional design architecture based on the Curriculum as Code paradigm, integrating Generative AI with LaTeX and Python to automate the creation of reproducible, visually consistent, and tec...
76. People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe ​
Author: Maria-Louisa Wightman, Guillaume Bied, Tijl De Bie
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07367v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly used as a primary source of information and advice, understanding their alignment to humans in terms of values becomes a pressing concern. A growing literature has leveraged large scale surveys to invest...
77. FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings ​
Author: Sasan Mansouri, Daniel Saad, Mark Wahrenburg, Manu Weissel, Fabian Woebbeking
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.DB, econ.GN, q-fin.EC, q-fin.GN
arXiv:2608.07400v1 Announce Type: new Abstract: Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence. Similar facts and disclosures recur across sections of a filing, across...
78. GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks ​
Author: Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno, Lynda Tamine
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, cs.LG
arXiv:2608.07411v1 Announce Type: new Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark ...
79. ResidencyRL: Reinforcement Learning in Simulated Clinical Environments ​
Author: Valentin Li'{e}vin, Samuel Schmidgall, Tim Strother, Alex Bijamov, Akshay Goel, Anil Palepu, Chunjong Park, Vahid Balazadeh, Min Woo Sun, Marius Guerard, Justin Chen, Dave Steiner, Vikram Dhillon, Ibrahim Azar, Akhil Mehta, Nicholas Spetsieris, Shilpan Shah, Maen Abdelrahim, Amit Dahiya, Yun Liu, Katherine Chou, Yossi Matias, Avinatan Hassidim, Dale R. Webster, Quoc V. Le, Raia Hadsell, Joelle Barral, Carey Radebaugh, Aleksandra Faust, Shekoofeh Azizi, Mike Schaekermann, Po-Hsuan Cameron Chen, Tao Tu, David Racz, Lin Yang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.07418v1 Announce Type: new Abstract: In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy. Much of clinical reasoning rel...
80. CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing ​
Author: Yan Zhou, Yue Ouyang, Kaiyang Zheng, Suncheng Xiang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07424v1 Announce Type: new Abstract: Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a chain of thought, or applying a stronger evaluator. Under a fixed inference budget, these choices compete. This paper formulates test-t...
81. A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy ​
Author: Bhavika Jalli, Nikhil Korati Prasanna, Jayanta Choudhury
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.PF
arXiv:2608.07427v1 Announce Type: new Abstract: LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom network analytics and numerical time-series data analysis (NTSDA), where raw multivariate KPI windows from 4G/5G...
82. TEPA: Revoking Stale Memories for Conflict-Robust Language Agents ​
Author: Yan Zhou, Yue Ouyang, Kaiyang Zheng, Suncheng Xiang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07429v1 Announce Type: new Abstract: Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also creates a central falsifiability problem: when the world changes, stale memories can remain retrievable and pollute the prompt. We characte...
83. Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers ​
Author: Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau, Anass Belfatmi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.07436v1 Announce Type: new Abstract: Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on $(a+b) \bmod 113$ grok and later lose generalization. Across five seeds ...
84. Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ​
Author: Jiacheng Miao, Jin Mu, Guanhua Chen, James Zou
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses end-to-end. However, ...
85. PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents ​
Author: Mohammad Amanlou, Parham Abed Azad, Farbod Davoodi, Mostafa Masumi, Behnam Bahrak, Abdol-Hossein Vahabie
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC
arXiv:2608.07438v1 Announce Type: new Abstract: Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolved conflict also shape what becomes accessible. We present PsychoAgent, a cognitive architecture for LLM agents that separates factual ...
86. Blast Radius ​
Author: MY Pitsane, Hope Mogale
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.07440v1 Announce Type: new Abstract: Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates an incoming prompt's reach through coupled context and code channels. NECROPHORESIS enables reversi...
87. SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent ​
Author: Mingxuan Zheng, Yujin Zhou, Chuxue Cao, Boqin Yin, Yuyao Zhang, Jiapeng Sun, Shuaishuai Gong, Sirui Han, Yike Guo
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.07449v1 Announce Type: new Abstract: LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifacts that are loaded into the agent's context without weight updates. Recent methods refine skills thro...
88. Interaction Creates Dynamical AI Behavior Absent in Isolation ​
Author: Bella Xinrui Li, Frank Yingjie Huo, Neil F Johnson
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.dis-nn, cond-mat.stat-mech, physics.soc-ph
arXiv:2608.07457v1 Announce Type: new Abstract: What will happen when AI agents interact in daily life, e.g. when one AI starts bossing another around? We find a counterintuitive answer that opens new avenues for out-of-equilibrium Physics. When a boss AI directs a stream of messages at the subordin...
89. Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation ​
Author: Yuma Asato, Kiyoaki Shirai, Natthawut Kertkeidkachorn
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.05726v1 Announce Type: cross Abstract: Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts. However, LLM evaluators tend to generate particular score...
90. Multimodal Drivers' Emotion Recognition and Safety-Oriented Intervention for Intelligent Transportation Systems ​
Author: Chang Liu, Dalai Mengke, Hanbo Zhou, Jia Hu, Peter Mihajlik, Tamas Sziranyi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.06378v1 Announce Type: cross Abstract: Driver emotions can affect risk perception, decision-making, and vehicle control under complex road conditions. Existing studies mainly focus on driver emotion recognition, while limited attention has been given to context-aware intervention that joi...
91. Mobile Interaction for Assessing Fatigue, Sleep, and Activity in Neurodegenerative and Chronic Diseases ​
Author: Julian Fierrez, Alejandro Pe~na, Aythami Morales, Ruben Tolosana, Ruben Vera-Rodriguez, Meenakshi Chatterjee, Ahmaniemi Teemu, Wan-Fai Ng, Walter Maetzler, Nikolay V. Manyakov, Jennifer Kudelka, Ralf Reilmann, C. Janneke van der Woude, Kristen Davies, Victoria Macrae, IDEA-FAST Consortium
Published: 8/10/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.06380v1 Announce Type: cross Abstract: Fatigue, sleep, or disturbances in daily activities are common symptoms among patients with neurodegenerative disorders (NDD) and immune-mediated inflammatory diseases (IMID). The current assessment of such symptoms is usually conducted using patient...
92. Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent Collaboration ​
Author: Mateus Levi Sim~oes Fernandes, Alberto Sardinha
Published: 8/10/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.MA
arXiv:2608.06381v1 Announce Type: cross Abstract: Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting generalizability to state-of-the-art teaming research. We provide the first systematic evaluation of XAI ...
93. TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation ​
Author: Guanzhi Deng, Haibo Wang, Kuan Wu, Xiangru Jian, Shing Yin Wong, Sichun Luo, Zhuoran Wang, Linqi Song
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.06396v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models route each token through a small subset of experts, making routing patterns useful for identifying task-relevant experts during downstream adaptation. Yet current approaches have two limitations: task experts ...
94. Agentic Planning for Symbolic Execution ​
Author: Daniel Koh Ji Yang, Yannic Noller, Corina S. Pasareanu, Youcheng Sun
Published: 8/10/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.SE
arXiv:2608.06397v1 Announce Type: cross Abstract: Symbolic execution seeks to explore feasible program paths, yet a practical run may exhaust its resources while much program behaviour remains unreached. We investigate a complementary way of extending its practical reach by reasoning about how the s...
95. Recovering Explanations from Transformed Rule-Based Ontologies ​
Author: Alex Ivliev, Markus Kr"otzsch, Maximilian Marx
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.DB
arXiv:2608.06399v1 Announce Type: cross Abstract: Datalog rules are often used to define ontologies over Knowledge Graphs. Rule reasoners routinely optimise such ontologies by rewriting their rules into a form that can be evaluated more efficiently. These transformations preserve the entailed facts,...
96. TransSLR: A Lightweight Transformer for Sign Language Recognition ​
Author: Lucia Yen Wanchi, Samuel Johnny, Victor Tolulope Olufemi, Emmanuel Aaron, Moise Busogi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.06407v1 Announce Type: cross Abstract: Automated Sign Language Recognition for under-represented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifies this gap: the only available bench-mark, CASL-W60, has a best reported accuracy of 69.93%, and we...
97. Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models ​
Author: Linkai Peng, Baorian Nuchged
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.06409v1 Announce Type: cross Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a generation-aligned diagnostic ...
98. WorldMark: A Plug-and-Play World Knowledge Interface for Cross-Host Language Model Watermarking ​
Author: Song Xiao, Yuqi Yuan, Yanshuo Zhang, Kejun Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.06416v1 Announce Type: cross Abstract: Watermarking traces the provenance of text produced by large language models by embedding statistically detectable signals during decoding. Existing schemes fall into logits-based, sampling-based, entropy-aware, and adaptive-strength families, yet al...
99. Risk-Aware Decision Policies for Agents Under Noisy Perception ​
Author: David Szczecina
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06420v1 Announce Type: cross Abstract: Perception in biological systems is inherently noisy, requiring organisms to make decisions under uncertainty where misclassification can be costly or fatal. We present an Artificial Life predator-prey model of foraging under noisy perception, and co...
100. ED-CSP: Crystal Structure Prediction from Electron Diffraction ​
Author: Germain Poloudenny, Ya"el Fr'egier, Arnaud Demorti`ere
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06448v1 Announce Type: cross Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. Existing ED-based learning methods mainly predict crystallographic labels, reconstruct structures fr...
101. CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training ​
Author: Amine Lbath, Manan Suri, Aurelien Delaitre, Vadim Okun, Massih-Reza Amini, Ram D. Sriram, Dinesh Manocha
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE
arXiv:2608.06471v1 Announce Type: cross Abstract: Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in real-world software. Generally available agents can already aid attackers, who only need to find one exploitable...
102. StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection ​
Author: Zhuoxin Zhan, Akbar Rafiey, Avery Ma, Leila Pishdad, Layla El Asri
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as web pages. In this paper, we introduce multi-step indirect prompt injection, a new attack class aga...
103. LyEvO: Lyapunov-Guided Evolutionary Optimization for Safe and Robust Sim-to-Real Policy Learning ​
Author: Riccardo Curcio, Hongpeng Cao, Marco Caccamo
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG, cs.NE
arXiv:2608.06481v1 Announce Type: cross Abstract: Training controllers that are safe and robust in simulation, and systematically assessing their readiness for real-world deployment, remain key challenges in sim-to-real transfer. To address this, we propose LyEvO, a physics-grounded framework that c...
104. Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events ​
Author: Ming Wang, Peidong Wang, Xiaocui Yang, Daling Wang, Shi Feng, Fiona Fui-Hoon Nah, Ee-Peng Lim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SI
arXiv:2608.06485v1 Announce Type: cross Abstract: Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coh...
105. Agentic AI: User Empowerment or Enclosure? ​
Author: David Gamba, Daniel M. Romero, Grant Schoenebeck
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2608.06510v1 Announce Type: cross Abstract: Agentic AI promises a more flexible form of digital agency: systems that can act on users' behalf, from filtering content to negotiating prices to selecting services. Whether it will empower users is an open question, and we argue that the answer dep...
106. CertBind from Multimodal Connectivity to Certifiable Retrieval Decisions ​
Author: Shuheng Cao, Zhenhao Zhang, Ruiqi Chen, Renjie Cao, Weijia Zhang, Siyu Zhang, Jiaxin Liu, Xiangyu Zeng, Haotian Geng, Fan Gu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06516v1 Announce Type: cross Abstract: Lightweight connectors make frozen multimodal encoders composable at the representation level. Deployment exposes a second problem at the level of task decisions. A connected route can expand cross-modal reach while changing an established native ret...
107. TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade ​
Author: Debodeep Banerjee, Amitangshu Dasgupta
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.06549v1 Announce Type: cross Abstract: LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks. In realpolitik, negotiations are longitudinal data, where participating parties can...
108. SyncSBC: Decentralized Swarm Behavior Prediction for Synchronized Autonomous Control ​
Author: Varun Raveendra, Connor Mattson, Daniel S. Brown
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.06587v1 Announce Type: cross Abstract: Robot swarms utilize many independent limited-sensing agents to produce complex emergent behaviors without requiring centralized control. However, little research explores how agents can infer swarm-level behavior from purely local perception, a capa...
109. Beyond "AI Language": The case for the idiolectal nature of LLM output ​
Author: Karolina Rudnicka, Thomas Stephan Juzek
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.06589v1 Announce Type: cross Abstract: While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human idiolects. We analyse ...
110. Flowing Through States: Neural ODE Regularization for Reinforcement Learning ​
Author: Mohamed Ghanem, Bernd Finkbeiner
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06595v1 Announce Type: cross Abstract: Neural networks applied to sequential decision-making tasks typically rely on latent representations of environment states. While environment dynamics dictate how semantic states evolve, the corresponding latent transitions are usually left implicit,...
111. SLED: Scalable Location Encoding via Distillation ​
Author: Kevin Lane, Zhongying Wang, Esther Rolf, Morteza Karimzadeh
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.06612v1 Announce Type: cross Abstract: The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer size of the Earth Observations (EO), differing modalities, and different sensor types pose significant ...
112. Do 3D Medical Foundation Models See Through MRI Artifacts? A Controlled Study of Representation Robustness ​
Author: Julia Anna Mielcarz, Daniel Klaaby, Mostafa Mehdipour Ghazi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.06613v1 Announce Type: cross Abstract: Self-supervised 3D medical foundation models are increasingly used as general-purpose feature extractors, yet their sensitivity to MRI artifacts remains poorly understood. We present a controlled evaluation of representation robustness across five pr...
113. Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval ​
Author: Linhai Ma, Ethan F. Wei, Xueqing Peng, Yan Wang, Lingfei Qian, V'ictor Guti'errez-Basulto
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.06614v1 Announce Type: cross Abstract: Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this ...
114. Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation ​
Author: Chunhui Shi, Xinwen Fu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06631v1 Announce Type: cross Abstract: Cryptanalytic extraction has been demonstrated for ReLU networks, for networks using componentwise activations such as GELU or SiLU, and for a Transformer's final projection matrix. These methods do not recover the bias-free Gated Linear Unit (GLU) f...
115. Bypassing Krum: Selection-Aware Backdoor Attacks in Federated Learning ​
Author: Srinivasan Subramanian, Md. Abdullah Al Hafiz Khan, Kazi Aminul Islam
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06637v1 Announce Type: cross Abstract: Robust aggregation methods are widely used in federated learning to mitigate the impact of adversarial client behavior. Distance-based aggregation rules, such as Krum and Multi-Krum, select updates that are closest to the majority under the assumptio...
116. MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering ​
Author: Jakub Po'cwiardowski, Mateusz Modrzejewski
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.06638v1 Announce Type: cross Abstract: Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze two public text-to-MIDI systems of contrasting design: the purpose-built encoder--decoder text2midi and MIDI-LLM...
117. Characterizing the Quality Profile of AI-Generated C++ in Production ​
Author: Michael Tran, Fred Lewis, Kun Yang, Saksham Thakur, Aditya Kini, Aditya Patil, Milad Hashemi, Parthasarathy Ranganathan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.06640v1 Announce Type: cross Abstract: The widespread integration of AI coding assistants offers undeniable boosts to engineering velocity. Yet, recent studies point to a growing trade-off, revealing persistent challenges with code quality and maintainability. Industry leaders, including ...
118. SoRoMoX: Fast, Differentiable, and Parallelizable Soft Robot Models ​
Author: Maximilian St"olzle, Solange Gribonval, Daniel Feliu-Talegon, Vito Daniele Perfetta, Michele Martini, Chuhan Zhang, Kiwan Wong, Mohammed Tarnini, Anup Teejo Mathew, Federico Renda, Daniela Rus, Cosimo Della Santina
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.06650v1 Announce Type: cross Abstract: Reduced-order models based on Cosserat-rod theory are now well established, and modeling theory is no longer the primary bottleneck in soft-robot control. Their implementations, however, do not support the differentiable, GPU-parallel, and control-or...
119. Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models ​
Author: Zhuoheng Huang, Mukesh Singh
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.06690v1 Announce Type: cross Abstract: Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different systems question: can trusted authorization determine which newly trained parameters are reachable by the forwar...
120. Online Monitoring and Corrective Steering of Programming Agents ​
Author: Shuyang Liu, Saman Dehghan, Ji Young Kim, Jatin Ganhotra, Martin Hirzel, Reyhaneh Jabbarvand
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL, cs.LG
arXiv:2608.06701v1 Announce Type: cross Abstract: Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or the issue description lacks the information needed to localize and repair it. As a result, agents traverse long t...
121. Scalable Long-Horizon Planning with Staggered Updates for Lifelong MAPF ​
Author: Vaibhav Sanjay, Jiaoyang Li
Published: 8/10/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.RO
arXiv:2608.06702v1 Announce Type: cross Abstract: Lifelong Multi-Agent Path Finding (LMAPF) requires generating collision-free paths for large agent fleets under strict real-time constraints. Reactive frameworks such as PIBT and Enhanced PIBT (EPIBT) scale effortlessly to thousands of agents through...
122. Dueling World Models: Advantage-Style Action Channels for Common-Mode Distractor Rejection ​
Author: Jiazhuo Li, Yiming Fei, Zhiruo Zhou, Heikichi Hayashi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06706v1 Announce Type: cross Abstract: Latent world models plan by predicting future states from an action, but when a scene contains motion the agent does not control, they quietly go action-blind: predictions for different actions become indistinguishable even as the training loss keeps...
123. Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors ​
Author: Saeid Shokoufa, Mohammad Erfan Sadeghi, Mehdi Kamal, Massoud Pedram
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06723v1 Announce Type: cross Abstract: The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making accurate estimation essential for sustainable artificial intelligence deployment and hardware-aware de...
124. KReF: Training-Free Retrieval for Long-Term Time-Series Forecasting and Predictive Uncertainty ​
Author: Yang Zhang, Rui Su
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06748v1 Announce Type: cross Abstract: Probabilistic long-term time-series forecasting commonly relies on trained models. Training-free conformal methods typically construct intervals around a pre-existing point forecaster and do not natively represent a complete predictive distribution; ...
125. Progressive Content Refinement with Decaying Reward Joint LinUCB ​
Author: Shion Ishikawa, Pablo Loyola, Young-joo Chung, Yun Ching Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.06750v1 Announce Type: cross Abstract: Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-based Self-Refine to traditional bandit approaches often rely on static options or overlook the saturation effect....
126. Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation ​
Author: Kuan Xing, Ye Wang, Changyi Gan, Yuheng Li, Thao Nguyen, Yi Chang, Yilin Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.06751v1 Announce Type: cross Abstract: Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical shortcuts, such as recurring motifs, generic palettes, or overrepresented period signatures, rather...
127. Hidden Gauge Controls Feature Specialization in ReLU Networks ​
Author: Tongxi Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06766v1 Announce Type: cross Abstract: Training changes a network's predictions while allocating task-relevant structure across its internal units. In an overparameterized ReLU network, several neurons can begin with exactly the same functional role, yet one may acquire a teacher feature ...
128. Genotypic Triggers: Exposing Pharmacogenomic Blind Spots via Host-Specific Backdoors in Generative Antimicrobial Peptide Models ​
Author: Doniyorkhon Obidov, Xiaolong Guo, Yonghui Li, Kaichen Yang
Published: 8/10/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI, cs.CL
arXiv:2608.06779v1 Announce Type: cross Abstract: Large Language Models (LLMs) have accelerated drug discovery, particularly in the automated design of antimicrobial peptides (AMPs). However, current validation pipelines for peptide generation models overlook historical precedents showing that certa...
129. HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation ​
Author: Yuebo Luo, Ahmad Sedigh Baroughi, Philip Stachura, Le Chen, Venkatram Vishwanath, Zhenman Fang, Caiwen Ding
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AR, cs.AI
arXiv:2608.06791v1 Announce Type: cross Abstract: Application-specific FPGA accelerators offer substantial performance and energy-efficiency gains across many application domains, but developing them is costly, often requiring months of specialized effort. Even with high-level synthesis (HLS), desig...
130. Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training ​
Author: Oseong Choi, Hoeinn Kim, Jihoon Lee, Byungsoo Kang, Taeyeong Jang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.06792v1 Announce Type: cross Abstract: Foundation model(FM) for recommendation has shown strong ability to model long-horizon sequential user behavior. In practice, a single pretrained foundation model is often adapted to diverse downstream serving surfaces through Supervised Fine-Tuning(...
131. LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes ​
Author: Doniyorkhon Obidov, Honggang Yu, Xiaolong Guo, Kaichen Yang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.06795v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) enables efficient specialization and distribution of large language models through compact adapters. However, untrusted adapters introduce a supply-chain threat: a backdoored adapter can cause a model to generate harmful co...
132. Coupling Planning with Episodic Memory in LLM Agents for Software Issue Resolution ​
Author: Jiahao Zhang, Yifan Zhang, Yu Huang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.06811v1 Announce Type: cross Abstract: Resolving a real software issue with a large language model (LLM) agent is a long repair episode, often tens to hundreds of steps spanning exploration, hypothesis, implementation, and verification. Success depends on both the base model's local reaso...
133. FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding ​
Author: Quanquan Li, Hongbo Zhang, Yihe Chi, Jingyu Li, Xidong Xi, Liuyang Song, Hongzhen Zhang, Yuxiang Huang, Jing Ke, Siyuan Ma, Junyi Lin, Guitao Cao
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.06819v1 Announce Type: cross Abstract: Token-level collaboration allows a large language model (LLM) to assist a small language model (SLM) when their predictions diverge. Existing methods either use LLM-generated intervention tokens or rank candidates with the LLM's next-token probabilit...
134. Control-Anchored Residual Flow Matching Conditioned on Gene Geometry for Virtual Cell Perturbation Modeling ​
Author: Quanquan Li, Yihe Chi, Liuyang Song, Hongbo Zhang, Jingyu Li, Xidong Xi, Conghua Wei, Yijie Sun, Yu Chen, Xin Liu, Qi Hu, Jing Ke, Guitao Cao
Published: 8/10/2026, 4:00:00 AM
Categories: q-bio.MN, cs.AI
arXiv:2608.06824v1 Announce Type: cross Abstract: A central task in virtual cell modeling is predicting single-cell transcriptional responses to unseen genetic perturbations and drug combinations, and biological networks provide valuable priors on gene relationships. Existing graph-based models comm...
135. Investigating Quantum-Embedded Transformers on Classical Datasets for Cross-Modality Classification ​
Author: Hao-Yuan Chen
Published: 8/10/2026, 4:00:00 AM
Categories: quant-ph, cs.AI
arXiv:2608.06846v1 Announce Type: cross Abstract: We test whether a parameterized quantum circuit (PQC) improves a hybrid quantum-classical model's performance on classical datasets, using an interface-matched classical map as the control while holding all other components fixed. Our architecture, Q...
136. Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry ​
Author: Yehan Yang, Junyuan Shang, Yang Li, Guanqun Zhao, Shuohuan Wang, Dianhai Yu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.06849v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods typically decide which tokens or heads to preserve from runtime attention scores, observati...
137. Bridging the Gap Between Hyperdimensional Computing and Kernel Methods via the Nystr\"om Method ​
Author: Quanling Zhao, Anthony Hitchcock Thomas, Ari Brin, Xiaofan Yu, Tajana Rosing
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06860v1 Announce Type: cross Abstract: Hyperdimensional computing (HDC) is an approach from the cognitive science literature for solving information processing tasks using data represented as high-dimensional random vectors. The technique has a rigorous mathematical backing, and is easy t...
138. Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection ​
Author: Xuechao Zou, Shun Zhang, Kai Li, Yi Zhou, Xinyu Sun, Yuhui Chen, Zhe Wu, Congyan Lang, Junliang Xing
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MA
arXiv:2608.06865v1 Announce Type: cross Abstract: The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmarks provide limited coverage of ...
139. FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition ​
Author: Ghani Haider, Majid Kundroo, Boyun Eom, Dong Hwan Park, Chen Chen, Taehong Kim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.DC, cs.LG
arXiv:2608.06876v1 Announce Type: cross Abstract: In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentralized intelligence paradigm for Video Anomaly Recognition (VAR). This task is vital for maintaining high-fidelity ...
140. Georeferencing Non-Gazetteered Place Names using Biological Specimen Records ​
Author: Aneesha Fernando, Surangika Ranathunga, Kristin Stock, Raj Prasanna, Christopher B. Jones
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2608.06884v1 Announce Type: cross Abstract: Biological specimen records collected by natural history institutions constitute a rich source of temporal geographic knowledge, capturing biodiversity information about regional landscapes as they were recorded at different times. Using digitised da...
141. Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests ​
Author: Seitaro Ono, Senna Ross, Jun Saiki
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.LG
arXiv:2608.06908v1 Announce Type: cross Abstract: We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT). WEAT is a bias measurement method widely used in both computational social science and AI fairness research. I...
142. MaskFlow: Precise, Consistent and Seamless Regional Image Editing ​
Author: Rui Xu, Yang Yong, Shunzi Yang, Ruihao Gong, Chengtao Lv
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.06929v1 Announce Type: cross Abstract: Regional image editing has attracted considerable attention for its spatial controllability. Although instruction-based and mask-reference-based editing methods can achieve strong semantic alignment, reliable regional control remains challenging, whe...
143. Ask-E: An Environment for Calibrated Question Generation ​
Author: Sarah Pratt, Jae Sung Park, Scott Geng, Ali Farhadi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.06933v1 Announce Type: cross Abstract: Today, we improve models by training and evaluating them on problems at the frontier of their abilities. Creating such problems is itself a demanding task, requiring the ability to probe model limits and generalize beyond existing question distributi...
144. Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning ​
Author: Chen Ling, Hanqian Li, Dongnan Liu, Keyu Qian, Jungang Li, Xinglong liu, Shiyi Wang, Xin Dong, Pengcheng Zhu, Wei Zhou, Linjian Mo, Nai Ding
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.06938v1 Announce Type: cross Abstract: The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commonsense reasoning, which requires models to reason beyond common assumptions. Recent studies mainly improve visu...
145. Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression ​
Author: Alex Kwon
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.06953v1 Announce Type: cross Abstract: Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory. We ask what governs whether it does. Matched notes carry the identical claim and...
146. Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents ​
Author: Takahiro Ezaki, Naoto Imura, Katsuhiro Nishinari
Published: 8/10/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI, cs.CY
arXiv:2608.06968v1 Announce Type: cross Abstract: Language-model agents act on state encodings of their environment, yet these are treated as interchangeable interfaces. Using pretrained language models, we designed a circular-synchronization experiment applying a state-encoding intervention while h...
147. PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue ​
Author: Bo Tang, Jianan Yang, Junyi Zhu, Yiquan Wu, Rui Zhao, Zhengyu Yang, Yang Zhang, Feiyu Xiong, Zhiyu Li, Jiajun Shen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.06975v1 Announce Type: cross Abstract: Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing work falls short on two fronts: representations are typically static profiles that cannot be updated locally without destabilizing u...
148. HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses ​
Author: Xiao Zhang, Yusheng Wang, Yuhao Fei, Dongyuan Li, Zian Liang, Liuyu Xiang, Hongxun Gu, Zhaofeng He
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.06984v1 Announce Type: cross Abstract: Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However, this capability creates delayed safety risks: attacker-influenced content can cross system boundarie...
149. Density-aware Hierarchical Clustering Based on Element-Categorized Connection Subgraphs ​
Author: Yuning Yu, Jos'e Rodr'iguez-Pi~neiro, Xuefeng Yin, Bin Feng
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06990v1 Announce Type: cross Abstract: Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning. Among various clustering methods, hierarchical clustering, density-based clustering, and graph clustering stand out as representative approaches....
150. GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base ​
Author: Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB
arXiv:2608.06992v1 Announce Type: cross Abstract: We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2.0 contains 38.4M triples over 1.6M canonical entities, together with 207.6K consolidated relations and 66K con...
151. Beyond Foundation Models: Dimension-Aware Neural Architecture Search with Small-Data Representation Models for Cryocooler Lifetime Prediction ​
Author: Gregor Molan (Comtrade 360 d.o.o., Letali\v{s}ka cesta 29b, Ljubljana, 1000, Slovenia), Grafika Jati (Comtrade 360 d.o.o., Letali\v{s}ka cesta 29b, Ljubljana, 1000, Slovenia), Francesco Barchi (Alma Mater Studiorum - Universita di Bologna, Department of Electrical, Electronic, and Information Engineering), Andrea Acquaviva (Alma Mater Studiorum - Universita di Bologna, Department of Electrical, Electronic, and Information Engineering), Alja\v{z} Osterman (LE-Tehnika d.o.o., \v{S}uceva 27, Kranj, 4000, Slovenia), Martin Molan (Comtrade AI GmbH, Grafenauweg 8, Zug, 6300, Switzerland)
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06993v1 Announce Type: cross Abstract: Large-scale pretrained time-series models achieve strong results through large-scale pretraining and task-agnostic representation learning, but they rely on abundant, diverse data that industrial and scientific domains often lack. We therefore propos...
152. Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models ​
Author: Xiangkai Ma, Yue Ma, Junjie Wang, Sheng Xu, Mingyang Li, Han Zhang, Yuzheng Zhuang, Wenzhong Li, Zhihao Yuan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.06994v1 Announce Type: cross Abstract: World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visual branches focus on predicting static visual observation, rather than ...
153. An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation ​
Author: Emma Jouffroy, Warren Jouanneau, Marc Palyart
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.07023v1 Announce Type: cross Abstract: Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, directly impacting downstream tasks like accurate talent matching. To address this, we propose a hybrid knowledg...
154. Accounting Graph Transformer for Short-History Multi-KPI Forecasting in Small Businesses ​
Author: Shrutendra Harsola, Vignesh Subrahmaniam
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.07037v1 Announce Type: cross Abstract: Small businesses often have only 12-24 months of accounting history, yet planning and risk workflows require coordinated forecasts across financial statements. We study joint 12-month forecasting of 13 income-statement, balance-sheet, cash-flow, and ...
155. Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering ​
Author: Xiuwei Shang, Li Hu, Xiao Jiang, Jieke Shi, Junda He, Zhou Yang, Shaoyin Cheng, Guoqiang Chen, Weiming Zhang, David Lo
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CR
arXiv:2608.07038v1 Announce Type: cross Abstract: Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereby reducing the cognitive burden of reverse analysis and improving efficiency. However, reliably evaluating HOBR...
156. Soft Redaction of Image Provenance via Zero-Knowledge Proofs ​
Author: Muhammad Awan, John Collomosse
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.07063v1 Announce Type: cross Abstract: Content provenance standards, such as C2PA, are increasingly used to attach signed records of origin, editing history, and rights to digital images. However, provenance transparency can conflict with privacy -- assertions that strengthen trust in an ...
157. AutoIntervene: Calibrated Intervention for Action-Chunking Imitation Learning Policies ​
Author: Jinhe Tang, Weiming Zhi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.HC, cs.LG
arXiv:2608.07065v1 Announce Type: cross Abstract: Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather than single-step commands. Yet perception errors and execution drift can move the robot outside the demonstrati...
158. Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers ​
Author: Xiangyu Meng, Peng Chen, Mingzhen Li, Jianmin Wang, Sen Wang, Guangming Tan, Weile Jia, Mohamed Wahib, Tao Luo, Xun Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.DC, cs.AI
arXiv:2608.07078v1 Announce Type: cross Abstract: Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale. Among existing approaches, LightDock leverages Glowworm Swarm Optimization (GSO) for accuracy, yet suffers fr...
159. LifelongCrossNav: Persistent 3D Semantic Memory for Cross-Floor Multi-Object Navigation ​
Author: Zehui Li, Zihao Sun, Jiawei Xu, Zheqi He, Xiaoqiang Zhang, Jing-Shu Zheng, Lu Liu, Dahui Gao, Xiuwan Chen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.07079v1 Announce Type: cross Abstract: Object-goal navigation has made substantial progress in semantic perception and exploration, yet persistent memory for multi-object navigation and cross-floor navigation are still commonly addressed separately. We present LifelongCrossNav, a framewor...
160. Beyond Isolation: Unlocking Reinforcement Learning Component Synergy for Sample-Efficient Continuous Control ​
Author: Qi Zhao, Guozheng Ma, Yilun Kong, Lu Li, Haoyu Wang, Zilin Wang, Tiantian Zhang, Yuxing Wang, Jian Sha, Yongzhe Chang, Xueqian Wang, Dacheng Tao
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.07086v1 Announce Type: cross Abstract: Reinforcement learning systems are significantly more complex than other machine learning paradigms due to inherent properties, causing RL system design to jointly account for many tightly coupled factors. Despite advances in individual algorithmic c...
161. RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs ​
Author: Qiyanhui Lu, Han Wu, Rongjian Xu, Tingzhang Luo, Cheng Fan, Xinghao Chen, Minjing Dong, Jufeng Yang, Jianyuan Guo
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.07088v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing training-free pruning methods select tokens by importance, diversity, or spatial coverage, but treat ret...
162. Human-Centered Explainable AI for TinyML Edge Devices: A Pareto-Based Selection Framework with LLM-Guided Design ​
Author: Zeinab Dehghani, Dhavalkumar Thakker, Koorosh Aslansefat, Kuniko Paxton, Bhupesh Kumar Mishra, Baseer Ahmad, Rameez Raja Kureshi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.07091v1 Announce Type: cross Abstract: Edge Artificial Intelligence (Edge AI) enables the deployment of AI models directly on local edge devices, while such deployments are subject to strict resource constraints, particularly in clinical applications requiring local and timely inference. ...
163. International Transfer of Stochastic Cortical Self-Reconstruction ​
Author: Fabian Bongratz, Zhizheng Zhuo, Chao Zhang, Yaou Liu, Dennis M. Hedderich, Christian Wachinger
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, q-bio.NC
arXiv:2608.07092v1 Announce Type: cross Abstract: Stochastic cortical self-reconstruction (SCSR) enables personalized mapping of gray matter atrophy, a hallmark of neurodegenerative disorders such as Alzheimer's disease (AD), onto high-resolution cortical surfaces. Unlike conventional normative mode...
164. Geometry-Aware Camera Localization for Bronchoscopy ​
Author: Lumin Chen, Qingyao Tian, Jinpeng Li, Haoyu Jiang, Huai Liao, Xinyan Huang, Hongbin Liu, Dong Yi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.07116v1 Announce Type: cross Abstract: Camera localization in bronchoscopy remains a challenging problem due to stringent accuracy requirements, real-time constraints, and limited training data. Compared to natural scenes, the confined anatomical structures demand millimeter-level precisi...
165. PHOENIX: Fine-Tuned SLM-Powered Autonomous Satellite Lifetime Extension via Predictive Self-Healing and Multi-Agent AI Recovery ​
Author: Sumaiya Islam, Harsha Kumara Moraliyage
Published: 8/10/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.07126v1 Announce Type: cross Abstract: Most CubeSats, small and low-cost satellites roughly the size of a shoebox, do not survive as long as they were designed to: a study of 178 missions found that only 48-65% remain operational after two years, against a designed lifetime of 2-5 years. ...
166. Autonomous discovery of accelerator commissioning algorithms ​
Author: Thorsten Hellert (Lawrence Berkeley National Laboratory)
Published: 8/10/2026, 4:00:00 AM
Categories: physics.acc-ph, cs.AI
arXiv:2608.07138v1 Announce Type: cross Abstract: Simulated commissioning has become essential for de-risking modern light-source design and commissioning, but the procedures being simulated are still designed entirely by human experts. Their labor-intensive redevelopment after lattice changes makes...
167. Interpretable reinforcement learning with decision-tree pruning ​
Author: Mark Leon Ringer, Michel Tokic
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.07151v1 Announce Type: cross Abstract: Reinforcement learning policies are difficult to inspect, but interpreting them is a prerequisite for trustworthiness. Converting a trained policy into explicit decision-tree rules improves transparency and the resulting artifacts often remain too co...
168. Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation ​
Author: Yang Shen, Chonghao Cheng, Ziyi Zhao, Jialuo Zhu, Zhenyi Yi, Qi Zhao, Jian Yang, Yuhui Shi, Chin-Teng Lin
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.07154v1 Announce Type: cross Abstract: Open-source robotics and foundation models have lowered the barrier to embodied AI, yet language-guided laboratory automation still requires reliable alignment from instructions and observations to safe actions. This field report presents an OpenArm-...
169. Fluid-DiT: Graph-Free Diffusion Transformers for Fluid Flow Simulations Learning ​
Author: Shentong Mo, Guolin Ke
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE
arXiv:2608.07161v1 Announce Type: cross Abstract: Simulating complex fluid flows requires capturing full equilibrium distributions rather than just mean trajectories, yet high-fidelity solvers remain computationally prohibitive. Recent advances, such as Diffusion Graph Networks (DGNs), have combined...
170. Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation ​
Author: Francisco Caetano, Tim J. M. Jaspers, Haiko Middeljans, Martijn R. Jong, Rixta A. H. van Eijck van Heslinga, Floor Slooter, Albert J. de Groof, Jacques J. Bergman, Peter H. N. De With, Fons van der Sommen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.07176v1 Announce Type: cross Abstract: Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers. Although representation alignment has improved efficiency in general ...
171. Momba: Network Modernization Improves Multi-Objective Reinforcement Learning ​
Author: Adam \v{S}tafa, Santeri Heiskanen, Petr Novotn'y, Joni Pajarinen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.07180v1 Announce Type: cross Abstract: Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in sample efficiency and asymptotic performance without altering the underlying algorithms. In contrast, work on mu...
172. Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes ​
Author: Luc Hazenoot, Zhaochun Ren, Amirhossein Zohrehvand
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, econ.GN, q-fin.EC
arXiv:2608.07208v1 Announce Type: cross Abstract: Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding similarities. They score the words a text uses, not the judgment a reader forms about it. Recent work has shown...
173. Toward a Causal Data Management Ecosystem for Decision Making and Agentic AI ​
Author: Dazhuo Qiu, Yingli Zhou, Amedeo Pachera, Angela Bonifati, Andrea Mauri
Published: 8/10/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.SY, eess.SY
arXiv:2608.07214v1 Announce Type: cross Abstract: Modern AI is no longer a single model but an ecosystem: classical ML predictors, deep and multimodal models, large language models, and agents, each trained and tuned over different data sources and each producing outputs at scale that become inputs ...
174. Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications ​
Author: Kaela Kokkas, Hairong Wang, Richard Klein, Nazir A. Ismail, Natalie Irwin, Mohammad Z. Moonsamy, Kubendran Naidoo, Jeremy Nel, Ekene E. Nweke, Raveen Parboosing, Emmanuel K. Sekyi, Rebecca T. van Dorsten, Bruce A. Bassett, Robert F. Breiman
Published: 8/10/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI, cs.CL
arXiv:2608.07250v1 Announce Type: cross Abstract: Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will reduce disease burdens. However, relevant evidence is dispersed and infeasible for humans to comprehensi...
175. Reading Copom's Tone: A Weighted LLM Framework for Hawkish-Dovish Sentiment, Forward Guidance, and Uncertainty ​
Author: Gabriel de Macedo Santos
Published: 8/10/2026, 4:00:00 AM
Categories: econ.GN, cs.AI, q-fin.EC
arXiv:2608.07251v1 Announce Type: cross Abstract: This paper documents an applied natural-language-processing framework for measuring the tone of Brazilian Monetary Policy Committee (Copom) statements. The project is explicitly inspired by iSent, Ita'u's Central Bank sentiment classifier, particula...
176. SCALE: Scientific Concept Aggregation via LLMs and Embeddings for Fine-Grained Taxonomy Extension ​
Author: Daniele Raimondi, Feichi Lu, Oliver Grun, Mariia Eremina, Andrea Perlato
Published: 8/10/2026, 4:00:00 AM
Categories: cs.DL, cs.AI
arXiv:2608.07254v1 Announce Type: cross Abstract: The increasing specialization of scientific research challenges existing classification systems, which provide effective representations of broad disciplines and research topics but often fail to capture the fine-grained conceptual structure of conte...
177. TOFD: Target-Oriented Feature Decoupling against Poisoning Attacks in Split Federated Learning ​
Author: Yuhan Xie, Jingrong Huang, Chen Lyu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.07274v1 Announce Type: cross Abstract: Split Federated Learning (SFL) facilitates privacy-preserving collaborative training with reduced client-side overhead. However, its split architecture introduces unique attack surfaces, rendering it vulnerable to diverse poisoning attacks. Most exis...
178. A Finite E-Group of Nilpotency Class Three ​
Author: Xinan Dai, Wenhao Deng, Yidong Shi, Tailin Wu, Yuchen Yang
Published: 8/10/2026, 4:00:00 AM
Categories: math.GR, cs.AI
arXiv:2608.07275v1 Announce Type: cross Abstract: A group is an E-group if every element commutes with each of its endomorphic images. Caranti asked whether a finite E-group can have nilpotency class three. We prove that the $3$-group of order $3^{84}$ introduced by Abdollahi, Faghihi, and Mohammadi...
179. How Much AI Is in This Track? Quantifying the Proportion of AI-Generated Stems in Hybrid Music Mixtures ​
Author: Fernando Garcia de la Cruz, David L'opez-Ayala, Pablo Zinemanas, Emilio Molina, Mart'in Rocamora
Published: 8/10/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, eess.SP
arXiv:2608.07285v1 Announce Type: cross Abstract: AI-generated music is increasingly used at the stem level, with producers integrating synthetic drums, basslines, or vocals alongside human-performed instruments. However, current AI music detection systems are binary, treating tracks as either fully...
180. FUSE: Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabular Flow Matching ​
Author: Suman Cha, Seongchan Lee, Dohyun Ko, Hyunjoong Kim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.07294v1 Announce Type: cross Abstract: Generating mixed-type tabular data requires jointly modeling diverse feature distributions and their complex cross-column dependencies. Variational flow matching handles distinct endpoints via factorized distributions, yet leaves feature-specific pro...
181. EliSeg: Verified Target Construction for Report-Grounded Abnormality Segmentation ​
Author: Chengyi Peng, Haoyu Yang, Meixing Shi, Yuxiang Cai, Yankai Jiang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.07299v1 Announce Type: cross Abstract: Radiology reports describe clinical observations but do not specify executable segmentation targets. They may contain present, negated, prior,uncertain, or irrelevant findings, while multiple valid abnormalities may coexist. Existing segmentation met...
182. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination ​
Author: Zichuan Wang, Songlin Yang, Bo Peng, Zhenchen Tang, Yang Li, Beibei Dong, Jing Dong
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.07302v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work largely attributes this to insufficient visual attention. However, we find that both real and hallucinated obje...
183. Natural Language Processing Psychometrics ​
Author: Edoardo Sebastiano De Duro, Emma Franchino, Massimo Stella
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SI
arXiv:2608.07316v1 Announce Type: cross Abstract: Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure. NLP Psychometrics treats psychological prediction from text as a psychometri...
184. Towards Assurance Closure in AI-Native Large-Scale Agile Software Development ​
Author: Ricardo Britto
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.07317v1 Announce Type: cross Abstract: The AI-Native Manifesto envisions large-scale agile software development in which humans increasingly govern intent, risk, and exceptions while agents execute more of the engineering process. Realizing that end-state requires more than better code ge...
185. Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks ​
Author: Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.07335v1 Announce Type: cross Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parallelized Q-Network (PQN) algorithm achieves stable off-policy learning without relying on computationally expensi...
186. H2AL: Hyperbolic Hierarchy-aware Aggregative Learning for Registration-based Few-shot Medical Image Segmentation ​
Author: Jia Wang, Jiaming Cai, Zunying Hu, Zhanjie Wu, Jinyuan Liu, Hua Cheng, Yun Peng
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.07340v1 Announce Type: cross Abstract: Registration-based Few-shot medical image segmentation (RFMIS) aims to generate pseudo-labels for unlabeled images by warping a labeled image through registration. However, existing methods primarily perform pixel-level optimization and inference in ...
187. Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination ​
Author: Ruijie Hou, Yueyang Jiao, Zhao Wang, Yingming Li
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.07341v1 Announce Type: cross Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contamination mitigation evaluation} intervenes in the decoding process to suppress memorization and restore a contaminate...
188. Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding ​
Author: Karim Radouane, Jose G Moreno, Lynda Tamine
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR, cs.LG
arXiv:2608.07353v1 Announce Type: cross Abstract: Understanding concepts is fundamental to generalization. Despite their impressive performance on a wide range of tasks, Large Language Models (LLMs) still struggle with genuine concept understanding. Prior work has evaluated conceptual understanding ...
189. Assessing AI-generated music detection in real-world broadcast monitoring ​
Author: David L'opez-Ayala, Fernando Garc'ia de la Cruz, Pablo Zinemanas, Emilio Molina, Mart'in Rocamora
Published: 8/10/2026, 4:00:00 AM
Categories: eess.AS, cs.AI
arXiv:2608.07359v1 Announce Type: cross Abstract: The proliferation of AI-generated music in broadcast media raises concerns about transparency and fair compensation, but reliable detection under real broadcast conditions remains unresolved. Existing studies report substantial performance degradatio...
190. Measurements Automatically Extracted from Zero Echo Time MRI Using Deep Learning Image Segmentation and Geometric Modeling Agree with Expert Manual Readings ​
Author: Jack Consolini, Eric A. Bogner, Meghan Sahr, Matthew F. Koff, Kevin M. Koch, Hollis G. Potter
Published: 8/10/2026, 4:00:00 AM
Categories: physics.med-ph, cs.AI, cs.CV
arXiv:2608.07368v1 Announce Type: cross Abstract: Computed tomography (CT) remains the reference for 3D osseous morphometry in femoroacetabular impingement (FAI) but requires ionizing radiation and manual measurement. Zero echo time (ZTE) MRI visualizes cortical bone and yields FAI angles that agree...
191. LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer's Disease Screening ​
Author: Xin Wang, Yingchao Huang, Yuhan Su, Shanshan Yao, Wei Peng
Published: 8/10/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.LG
arXiv:2608.07378v1 Announce Type: cross Abstract: Early diagnosis of Alzheimer's disease (AD) is critical for enabling timely interventions that may slow disease progression and improve patient outcomes. There is a growing need for AD detection methods that are non-invasive and cost-effective, espec...
192. Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations ​
Author: Ioannis Ziogas, Ensieh Khazaei, Bilal Taha, Aamna Al Shehhi, Ahsan H. Khandoker, Leontios J. Hadjileontiadis, Dimitrios Hatzinakos
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.AS, eess.SP, stat.ML
arXiv:2608.07385v1 Announce Type: cross Abstract: Learning disentangled representations is a key requirement for developing versatile, general-purpose, and sustainable models in multi-modal wearable computing. However, existing approaches do not operate as full-stack wearable processors, i.e., they ...
193. PACE: Primitive-Aware Code Evolution for Automated Algorithm Design ​
Author: Zhuoliang Xie, Ruihao Zheng, Xiang Xu, Genghui Li, Zhengkun Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.07395v1 Announce Type: cross Abstract: Large Language Model (LLM)-based automated algorithm design typically evolves algorithms as complete, indivisible programs. While this whole-program perspective simplifies the search space, it fundamentally couples the useful local logic to its host ...
194. GeoDistill-Refine: Silhouette-First Geometry Distillation for Annotation-Free Spacecraft Segmentation ​
Author: Yonglong Zhang, Zongwu Xie, Yang Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.07405v1 Announce Type: cross Abstract: Foundation segmentation models can provide supervision for spacecraft imagery without manual training masks, but their predictions vary with textual prompts and may contain geometric errors that are amplified during distillation. This paper presents ...
195. I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning ​
Author: Shibo Gao, Chongxiao Wang, Chenglong Huang, Jie Ma, Haolin Shi, Fei Ding, Jing Li, Qiang Lyu, Yangyang Liu, Yang Liu, Jun Liu, Linlin Huang, Peipei Yang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.07417v1 Announce Type: cross Abstract: Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simplified video-text setting, limiting identity matching and person-centric reasoning. To bridge this gap, we introd...
196. Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits ​
Author: Elena Dumitrescu, Gert Lek, Lydia Y. Chen, J'er'emie Decouchant
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.07430v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as targets and as adversari...
197. SABRE: Scalable and Automated Benchmarking of VLMs under Stress ​
Author: Zixuan Lan, Luzhe Sun, Matthew R. Walter, Jiawei Zhou
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and challenge current mode...
198. Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools ​
Author: Afreen Alam, Evgenija Popchanovska, Ana Gjorgjevikj, Maryan Rizinski, Lubomir T. Chitkushev, Irena Vodenska, Dimitar Trajanov
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CY
arXiv:2608.07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As generative AI applications move from pilot to production, manual harm identification and mitigation are becoming diff...
199. Strategy-first synthesis planning for complex natural products ​
Author: Daniel Armstrong, Xuan-Vu Nguyen, Octavian Susanu, Gabriel Gibberd, Th'eo A. Neukomm, Tadd"aus Strunden, Dan Forster, Morgane Delattre, Shawn Teh, Cl'ement Rols, John Federice, Hayden Leatherwood, M. Lavelle Barnes, Maarten R. Dobbelaere, Peter Wipf, Jon T. Njardarson, Jieping Zhu, Philippe Schwaller
Published: 8/10/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2608.07454v1 Announce Type: cross Abstract: The total synthesis of a complex molecule is among the most demanding intellectual and experimental feats in chemistry: a chemist must plan many steps ahead for how to assemble simple building blocks into an intricate target, devise backup strategies...
200. CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG ​
Author: Gyuwan Kim, Cheoneum Park, Tao Yang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR, cs.LG
arXiv:2608.07458v1 Announce Type: cross Abstract: Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the co...
201. CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity ​
Author: Ananya Sahu, Mohit Bansal, Elias Stengel-Eskin
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.07460v1 Announce Type: cross Abstract: While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as well as those that req...
202. Boundary Density Likelihood for Direct Event-Time Supervision ​
Author: Clark Peng, Tolga Din\c{c}er
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ML
arXiv:2408.12792v2 Announce Type: replace Abstract: Event detection turns long recordings into a sparse set of ranked timestamps. Yet many sequence models are trained for samplewise segmentation and only convert predicted states into events after training. We ask whether training directly for the ev...
203. Serious Games: Human-AI Interaction, Evolution, and Coevolution ​
Author: Nandini Doreswamy (Southern Cross University, Lismore, New South Wales, Australia, National Coalition of Independent Scholars), Louise Horstmanshof (Southern Cross University, Lismore, New South Wales, Australia)
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.GT
arXiv:2505.16388v3 Announce Type: replace Abstract: The serious games between humans and AI have only just begun. Evolutionary Game Theory (EGT) models the competitive and cooperative strategies of biological entities. EGT could help predict the potential evolutionary equilibrium of humans and AI. T...
204. Social World Models ​
Author: Xuhui Zhou, Jiarui Liu, Akhila Yerukola, Hyunwoo Kim, Maarten Sap
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited information. In contrast, AI systems struggle to structure and reason about implicit social contexts, as the...
205. "LLM Agent Performance" Is Not a Single Evaluation Target ​
Author: Pengyu Zhu, Li Sun, Philip S. Yu, Sen Su
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2602.03238v3 Announce Type: replace Abstract: LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget. Unified execution controls these non-model factors by evaluating candidate models under the same configuration,...
206. Counterfactual Simulation Training for Chain-of-Thought Faithfulness ​
Author: Peter Hase, Christopher Potts
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2602.20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output. But well-known problems with CoT faithfulness severely limit what insights can be gained from this practice. In this paper, we int...
207. AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation ​
Author: Sukriti Manna, Henry Chan, Subramanian K. R. S. Sankaranarayanan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mes-hall
arXiv:2603.20986v2 Announce Type: replace Abstract: Phase-field modeling links thermodynamics and kinetics to microstructural evolution, but multiphysics frameworks such as MOOSE require expertise to construct inputs, manage campaigns, diagnose failures, and validate results. We introduce AutoMOOSE,...
208. INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation ​
Author: Alexandra Bazarova, Andrei Volodichev, Daria Kotova, Alexey Zaytsev
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.21607v2 Announce Type: replace Abstract: While retrieval-augmented generation (RAG) significantly improves the factual reliability of LLMs, it does not eliminate hallucinations, so robust uncertainty quantification (UQ) remains essential. In this paper, we reveal that standard entropy-bas...
209. MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large Language Models ​
Author: Farhad Abtahi, Abdolamir Karbalaie, Eduardo Illueca-Fernandez, Fernando Seoane
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.16009v2 Announce Type: replace Abstract: Most large language model benchmarks evaluate final-answer quality but reveal little about how models revise beliefs under disagreement or conflicting evidence. We introduce MEDLEY-BENCH, an open benchmark comparing structured private self-review a...
210. Alignment has a Fantasia Problem ​
Author: Nathanael Jo, Zoe De Simone, Mitchell Gordon, Ashia Wilson
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2604.21827v2 Announce Type: replace Abstract: In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.g., from brainstorming ideas to writing an essay). With the advent of highly capable AI assistants, people now offload various parts of their task to ...
211. DATAREEL: Automated Data-Driven Video Story Generation with Animations ​
Author: Ridwan Mahbub, Syem Aziz, Mizanur Rahman, Mahir Ahmed, Shadikur Rahman, Shafiq Joty, Enamul Hoque
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.25220v2 Announce Type: replace Abstract: Data videos combine animated visualizations with synchronized narration to communicate quantitative information and are widely used in journalism, education, and public communication. Automatically generating them requires deciding what story to te...
212. In-Context Examples Suppress Scientific Knowledge Recall in LLMs ​
Author: Chaemin Jang, Woojin Park, Hyeok Yun, Dongman Lee, Jihee Kim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.27540v2 Announce Type: replace Abstract: Scientific reasoning rarely stops at what is directly observable; it often requires uncovering hidden structure from data. From estimating reaction constants in chemistry to inferring demand elasticities in economics, this latent structure recovery...
213. Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models ​
Author: Shubham Kumar, Narendra Ahuja
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.00123v3 Announce Type: replace Abstract: Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts. Because we lack a robust understanding of why LLMs are susceptible to jailbreaks, future frontier models operating more autonomou...
214. Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On ​
Author: Yixiang Yao, Yuhang Yao, Xinyi Fan, Jiechao Gao, Jie Wang, Minjia Zhang, Srivatsan Ravi, Carlee Joe-Wong
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.19035v2 Announce Type: replace Abstract: The rapid advancement of Large Language Models has given rise to autonomous LLM-based agents capable of complex reasoning and execution. As these agents transition from isolated operation to collaborative ecosystems, we witness the emergence of the...
215. Ratchet: How Reliable Must an LLM Judge Be to Retire a Skill? ​
Author: Xing Zhang, Yanwei Cui, Guanghui Wang, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2605.22148v3 Announce Type: replace Abstract: A large language model (LLM) agent that writes and edits its own skill library must also decide which skills to keep, from one noisy scalar per skill. The answer is exact: a judge scoring failures as passes at rate $(1-\tau)/2$ or above retires not...
216. Same Answer, Different Confidence: Protocol Sensitivity in LLM Confidence Calibration ​
Author: Hankyeol Kim, Pilsung Kang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.27752v3 Announce Type: replace Abstract: Is verbalized confidence better calibrated than token likelihood? The answer depends on how the token likelihood is measured: which answer is scored, and under which prompt. Published comparisons diverge on this, and in a twelve-study audit five ne...
217. Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation ​
Author: Haocheng Luo, Jiahui Liu, Ruicheng Zhang, Zhizhou Zhong, Jiaqi Huang, Zunnan Xu, Quan Shi, Jun Zhou, Shuiyang Mao, Wei Liu, Xiu Li
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2606.06076v3 Announce Type: replace Abstract: While Vision-Language Models excel at general multimodal understanding, they still struggle with visual spatial planning. We attribute this limitation to a perception--reasoning modality gap. Visual planning requires models to infer latent state st...
218. ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents ​
Author: Lu Jia, Haibo Tong, Feifei Zhao, Jindong Li, Dongqi Liang, Ping Wu, Qian Zhang, Yi Zeng
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.08531v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks. As their capabilities and autonomy expand, the s...
219. Semantic Adapter Routing with Fine-Tuning Task Embeddings ​
Author: Enrico Cassano, Micha{\l} Brzozowski, Paolo Mandica, Zuzanna Dubanowska, Neo Christopher Chung
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.19079v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning (PEFT) has led to model ecosystems in which a single backbone is paired with many task-specialized adapters. Given such a library, routing aims to select the most appropriate adapter for a user query. While existing ...
220. SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data ​
Author: Zenghui Zhou, Xiaoyang Li, Xiaoxuan Qiao, Zhilang Wei, Tianming Lei
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2607.19949v4 Announce Type: replace Abstract: Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share. To address this challenge, ...
221. OpenForgeRL: Train Harness-native Agents in Any Environment ​
Author: Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou, Qianhui Wu, Hao Cheng, Wenlin Yao, Nikhil Singh, Zhou Yu, Jianfeng Gao
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.21557v3 Announce Type: replace Abstract: Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-e...
222. Property-driven Causal Abstractions for Markov Decision Processes ​
Author: Jule Schmidt, Maximilian Weininger, Clemens Dubslaff, David Parker, Nils Jansen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LO
arXiv:2607.26787v3 Announce Type: replace Abstract: Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and their valuations. The exponential blowup in the number of states renders many reasoning tasks in M...
223. Can AI agents conduct open-ended AI research? Early evidence from two case studies ​
Author: Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG
arXiv:2607.27191v2 Announce Type: replace Abstract: Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended...
224. H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases ​
Author: Shusen Zhang, Junyi Hu, Ye Feng, Ziteng Wang, Zhaoyuan Pan, Xiaojun Yuan, Jiangshou Hong, Guosheng Dong, Xiangzhi Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.00065v3 Announce Type: replace Abstract: Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constraints, and compositional concepts. However, existing representations lie at two extremes: single-vector retri...
225. Homebot: A Personal AI Agent for Conversational Home Assistance and Automation ​
Author: Shengyuan Ye, Yixin Zhang, Han Liang, Liekang Zeng, Jiangsu Du, Mu Yuan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02254v2 Announce Type: replace Abstract: \texttt{Homebot} is a locally deployable AI agent for conversational household assistance and automation. It accepts voice and instant-messaging requests through a shared runtime that combines language-model responses with registered tools and task...
226. LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment ​
Author: Linhan Xia, Rui Liu, Zhaofeng Zhang, Yihao Wang, Binrui Shen, Shengxin Zhu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.03020v2 Announce Type: replace Abstract: Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation through the frozen backbone. Every adaptation step therefore needs backward-capable hardware and must store or reco...
227. Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model ​
Author: Abdallah Khemais
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.03629v2 Announce Type: replace Abstract: A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream. For the one composition in that model where two carriers ...
228. SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse ​
Author: Jialuo Chen, Minghe Wang, Lingqi Jiang, Jianan Ma, Xinhao Deng, Xiaohu Du, Ruixiao Lin, Yunhao Feng, Linkang Du, Jingyi Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.05204v2 Announce Type: replace Abstract: LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. As skills become marketplace artifacts, auditing their reuse is...
229. Recursive Synthesis for Long-Horizon Terminal Tasks ​
Author: Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.05466v2 Announce Type: replace Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consis...
230. CourseGraph: Finding overlaps and differences in Computer Science courses across universities ​
Author: Arthur Nijdam, Paul Stankovski Wagner, Sara Ramezanian
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.05910v2 Announce Type: replace Abstract: Student mobility programs such as Erasmus+ enable students to take courses at other universities, broadening their academic and cultural horizons. However, this flexibility also leads to a practical challenge: ensuring that students do not take cou...
231. Contextual Information Policy Optimization for Search Agents ​
Author: Xingyu Guo, Wei Chen, Linlin Yang, Baochang Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06128v2 Announce Type: replace Abstract: Search agents extend large language models beyond static parametric memory by enabling them to acquire and use external evidence during multi-step reasoning. For knowledge-intensive tasks involving complex or evolving information, their reliability...
232. DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models ​
Author: ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun Guo, Wenbin Hu, Jinqiao Wang, Yafeng Deng
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.06243v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-dist...
233. Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations ​
Author: Sagar Tamang, Ayush Vyas, Tabarakul Hazarika
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR
arXiv:2608.06305v2 Announce Type: replace Abstract: Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents -- financial statements, audit...
234. Towards a Theoretical Understanding of Two Tower Recommendation Models ​
Author: Amit Kumar Jaiswal
Published: 8/10/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2403.00802v2 Announce Type: replace-cross Abstract: Production-grade recommender systems rely heavily on a large-scale corpus used by online media services, including Netflix, Pinterest, and Amazon. These systems enrich recommendations by learning users' and items' embeddings projected in a lo...
235. Harnessing the Synergy between LLM Agents and Knowledge Graphs for Urban Socioeconomic Prediction ​
Author: Zhilun Zhou, Jingyang Fan, Yu Liu, Fengli Xu, Depeng Jin, Yong Li
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.SI
arXiv:2411.00028v3 Announce Type: replace-cross Abstract: Socioeconomic prediction aims to leverage various urban data to predict the socioeconomic indicators of regions such as population and commercial activity level, which plays an important role in understanding urban regions and supporting deci...
236. Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription ​
Author: Benjamin Gutteridge, Matthew Thomas Jackson, Toni Kukurin, Xiaowen Dong
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2502.20295v3 Announce Type: replace-cross Abstract: Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical to obtain for real-world problems, or rely on zero-shot tools such as OCR engines and multi-modal LLM...
237. A primer on optimal transport for causal inference with observational data ​
Author: Florian F Gunsilius
Published: 8/10/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, econ.EM
arXiv:2503.07811v3 Announce Type: replace-cross Abstract: The theory of optimal transportation has developed into a powerful and elegant framework for comparing probability distributions, with wide-ranging applications in all areas of science. The fundamental idea of analyzing probabilities by compa...
238. PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks ​
Author: Ziyuan Li, Uwe Jaekel, Babette Dellen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV
arXiv:2505.04397v3 Announce Type: replace-cross Abstract: Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexplored. Product units offer a direct approach to modeling such interactions, but their use in deep archit...
239. Minimal Ingredients for Reward Assignment from Expert Demonstrations ​
Author: Zixuan Dong, Yumi Omori, Keith Ross
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2506.06793v2 Announce Type: replace-cross Abstract: Reward assignment from scarce demonstrations is a key challenge in both offline and online imitation learning. A common and intuitive strategy assigns rewards according to how closely learner trajectories match expert demonstrations. Although...
240. Learning to Walk With Less: A Dyna-Style Approach to Quadrupedal Locomotion ​
Author: Francisco Affonso, Felipe Tommaselli, Jo~ao H. Al'essio, Vivian S. Medeiros, Mateus V. Gasparino, Girish Chowdhary, Marcelo Becker
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2509.06296v2 Announce Type: replace-cross Abstract: Traditional on-policy reinforcement learning (RL) controllers for quadrupedal locomotion often suffer from low data efficiency, requiring millions of interactions with simulated environments to achieve stable control. We integrate model-based...
241. Evaluating Useful Surrogate Models for Configuration Tuning Beyond Accuracy: A Fitness Landscape Analysis Perspective ​
Author: Pengzhou Chen, Hongyuan Liang, Tao Chen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2509.21945v2 Announce Type: replace-cross Abstract: To efficiently tune configuration for better software system performance (e.g., latency) at the deployment and maintenance stage, many tuners have leveraged a surrogate model to expedite the process instead of solely relying on the profoundly...
242. Provable Training Data Identification for Large Language Models ​
Author: Zhenlong Liu, Hao Zeng, Weiran Huang, Hongxin Wei
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2510.09717v3 Announce Type: replace-cross Abstract: Identifying training data of large-scale models is critical for copyright litigation, privacy auditing, and ensuring fair evaluation. However, existing works typically treat this task as an instance-wise identification without controlling the...
243. Stability of Transformers under Layer Normalization ​
Author: Kelvin Kan, Xingjian Li, Benjamin J. Zhang, Tuhin Sahai, Stanley Osher, Krishna Kumar, Markos A. Katsoulakis
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC
arXiv:2510.09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable. Layer normalization, a standard component, improves training stability, but its placement has often been ad-hoc. In this paper, we conduct a principled study on the for...
244. In Situ Training of Implicit Neural Compressors for Scientific Simulations via Sketch-Based Regularization ​
Author: Cooper Simpson, Stephen Becker, Alireza Doostan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, cs.NA, math.NA
arXiv:2511.02659v4 Announce Type: replace-cross Abstract: Focusing on implicit neural representations, we present a novel in situ training protocol that employs limited memory buffers of full and sketched data samples, where the sketched data are leveraged to prevent catastrophic forgetting. The the...
245. Intelligence per Watt: Measuring Intelligence Efficiency of Local AI ​
Author: Jon Saad-Falcon, Avanika Narayan, Hakki Orhun Akengin, J. Wes Griffin, Herumb Shandilya, Adrian Gamarra Lafuente, Medhya Goel, Rebecca Joseph, Shlok Natarajan, Etash Kumar Guha, Shang Zhu, Ben Athiwaratkun, John Hennessy, Azalia Mirhoseini, Christopher R'e
Published: 8/10/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.CL, cs.LG
arXiv:2511.07885v5 Announce Type: replace-cross Abstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strains this paradigm faster than providers can scale. Two advances create an opportunity to rethink it: smal...
246. MetaSICL: Globalizing Auditory LLMs for Underserved Speakers and Languages via Meta Speech In-Context Learning ​
Author: Haolong Zheng, Siyin Wang, Zengrui Jin, Mark Hasegawa-Johnson
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL
arXiv:2601.18904v3 Announce Type: replace-cross Abstract: Generative AI for speech and audio is increasingly expected to serve users across languages, cultures, and communities, yet current auditory Large Language Models (LLMs) are still largely trained and evaluated on high-resource data. Globalizi...
247. Kimi K2.5: Visual Agentic Intelligence ​
Author: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, S. H. Cai, Yuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Cheng Chen, Guanduo Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kefan Chen, Liang Chen, Ruijue Chen, Xinhao Chen, Yanru Chen, Yanxu Chen, Yicun Chen, Yimin Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Ziwei Chen, Dazhi Cheng, Yean Cheng, Minghan Chu, Jialei Cui, Jiaqi Deng, Muxi Diao, Hao Ding, Mengfan Dong, Mengnan Dong, Yuxin Dong, Yuhao Dong, Angang Du, Chenzhuang Du, Dikang Du, Lingxiao Du, Yulun Du, Yu Fan, Shengjun Fang, Qiulin Feng, Yichen Feng, Garimugai Fu, Kelin Fu, Hongcheng Gao, Tong Gao, Yuyao Ge, Shangyi Geng, Chengyang Gong, Xiaochen Gong, Zhuoma Gongque, Qizheng Gu, Xinran Gu, Yicheng Gu, Longyu Guan, Shuhao Guan, Yuanying Guo, Xiaoru Hao, Dailan He, Tianhong He, Weiran He, Wenyang He, Yibo He, Yunjia He, Chao Hong, Hao Hu, Jiaxi Hu, Yangyang Hu, Zhenxing Hu, Ke Huang, Ruiyuan Huang, Weixiao Huang, Zhiqi Huang, Chaobo Jia, Tao Jiang, Zhejun Jiang, Xinyi Jin, Yu Jing, Guokun Lai, Aidi Li, C. Li, Cheng Li, Fang Li, Guanghe Li, Guanyu Li, Haitao Li, Haoyang Li, Jia Li, Jingwei Li, Junxiong Li, Lincan Li, Mo Li, Weihong Li, Wentao Li, Xinhang Li, Xinhao Li, Yang Li, Yanhao Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zhaoxi Li, Zheming Li, Weilong Liao, Jiawei Lin, Xiaohan Lin, Yibo Lin, Zhishan Lin, Zichao Lin, Cheng Liu, Chenyu Liu, Hongzhang Liu, Liang Liu, Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Tianyu Liu, Weizhou Liu, Xiangyan Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yuanxin Liu, Zhengying Liu, Zhongnuo Liu, Enzhe Lu, Haoyu Lu, Zhiyuan Lu, G. Luo, Junyu Luo, Tongxu Luo, Yashuo Luo, Long Ma, Shaoguang Mao, Yuan Mei, Xin Men, Fanqing Meng, Zhiyong Meng, Yibo Miao, Minqing Ni, Kun Ouyang, Siyuan Pan, Bo Pang, Yuchao Qian, Ruoyu Qin, Zeyu Qin, Jiezhong Qiu, Bowen Qu, Zeyu Shang, Youbo Shao, Tianxiao Shen, Zhennan Shen, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Feifan Song, Pengwei Song, Tianhui Song, Xiaoxi Song, Hongjin Su, Jianlin Su, Zhaochen Su, Lin Sui, Jinsong Sun, Junyao Sun, Tongyu Sun, Flood Sung, Yunpeng Tai, Chuning Tang, Heyi Tang, Xiaojuan Tang, Zhengyang Tang, Jiawen Tao, Shiyuan Teng, Chaoran Tian, Pengfei Tian, Bowen Wang, Chensi Wang, Chuang Wang, Congcong Wang, Dingkun Wang, Dinglu Wang, Dongliang Wang, Feng Wang, Hailong Wang, Haiming Wang, Hao Wang, Hengzhi Wang, Huaqing Wang, Hui Wang, Jiahao Wang, Jinhong Wang, Jiuzheng Wang, Kaixin Wang, Linian Wang, Qibin Wang, Shengjie Wang, Shuyi Wang, Si Wang, Wei Wang, Xiaochen Wang, Xinyuan Wang, Yao Wang, Yejie Wang, Yipu Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhexu Wang, Zifan Wang, Zihan Wang, Zizhe Wang, Chu Wei, Ming Wei, Chuan Wen, Zichen Wen, Chengjie Wu, Haoning Wu, Junyan Wu, Rucong Wu, Wenhao Wu, Yuefeng Wu, Yuhao Wu, Yuxin Wu, Zijian Wu, Chenjun Xiao, Jin Xie, Xiaotong Xie, Yuchong Xie, Bowei Xing, Boyu Xu, Jianfan Xu, Jing Xu, Jinjing Xu, L. H. Xu, Lin Xu, Suting Xu, Weixin Xu, Xinbo Xu, Xinran Xu, Yangchuan Xu, Yichang Xu, Yuemeng Xu, Zelai Xu, Ziyao Xu, Junjie Yan, Yuzi Yan, Guangyao Yang, Hao Yang, Junwei Yang, Kai Yang, Ningyuan Yang, Xiaofei Yang, Xinlong Yang, Xinyu Yang, Ying Yang, Yi Yang, Yi Yang, Zhen Yang, Zhilin Yang, Zonghan Yang, Haotian Yao, Dan Ye, Haoran Ye, Wenjie Ye, Zhuorui Ye, Peng Yebo, Bohong Yin, Chengzhen Yu, Longhui Yu, Tao Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Xiaokun Yuan, Yang Yue, Weihao Zeng, Dunyuan Zha, Haobing Zhan, Dehao Zhang, Hao Zhang, Jin Zhang, Puqi Zhang, Qiao Zhang, Rui Zhang, Xiaobin Zhang, Xiaoyun Zhang, Y. Zhang, Yadong Zhang, Yangkun Zhang, Yichi Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yushun Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Chenguang Zhao, Feifan Zhao, Jinxiang Zhao, Shuai Zhao, Xiangyu Zhao, Xuanle Zhao, Yikai Zhao, Zijia Zhao, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Junfeng Zhong, Longguang Zhong, Weiming Zhong, M. Zhou, Runjie Zhou, Xinyu Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yuxuan Zhu, Zhen Zhu, Jingze Zhuang, Weiyu Zhuang, Ying Zou, Xinxing Zu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2602.02276v2 Announce Type: replace-cross Abstract: We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of t...
248. Optimizing Spectral Prediction in MXene-Based Metasurfaces Through Multi-Channel Spectral Refinement and Savitzky-Golay Smoothing ​
Author: Shujaat Khan, Waleed Iqbal Waseer, Muhammad Shahid Jabbar
Published: 8/10/2026, 4:00:00 AM
Categories: physics.optics, cs.AI, eess.SP
arXiv:2602.08406v2 Announce Type: replace-cross Abstract: The prediction of electromagnetic spectra for MXene-based solar absorbers, where MXenes are a family of two-dimensional transition metal carbides and nitrides, is a computationally intensive task traditionally addressed using full-wave solver...
249. SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization ​
Author: Sunghwan Kim, Wooseok Jeong, Serin Kim, Sangam Lee, Dongha Lee
Published: 8/10/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2602.12187v2 Announce Type: replace-cross Abstract: Search-Augmented Generative Engines (SAGE) have emerged as a new paradigm for information access, bridging web-scale retrieval with generative capabilities to deliver synthesized answers. This shift has fundamentally reshaped how web content ...
250. MAC: A Conversion Rate Prediction Benchmark Featuring Labels Under Multiple Attribution Mechanisms ​
Author: Jinqi Wu, Sishuo Chen, Zhangming Chan, Yong Bai, Lei Zhang, Sheng Chen, Chenghuan Hou, Xiang-Rong Sheng, Han Zhu, Jian Xu, Bo Zheng, Chaoyou Fu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2603.02184v2 Announce Type: replace-cross Abstract: Multi-attribution learning (MAL), which enhances model performance by learning from conversion labels yielded by multiple attribution mechanisms, has emerged as a promising learning paradigm for conversion rate (CVR) prediction. However, the ...
251. Deterministic Preprocessing and Interpretable Fuzzy Banding for Cost-per-Student Reporting from Extracted Records ​
Author: Shane Lee, Stella Ng
Published: 8/10/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2603.04905v2 Announce Type: replace-cross Abstract: Administrative extracts are often exchanged as spreadsheets and may be read as reports in their own right during budgeting, workload review, and governance discussions. When an exported workbook becomes the reference snapshot for such decisio...
252. Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving ​
Author: Nikos Theodoridis, Reenu Mohandas, Ganesh Sistu, Anthony Scanlan, Ciar'an Eising, Tim Brophy
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.06054v2 Announce Type: replace-cross Abstract: The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios. However, these models often ...
253. Seeking SOTA: Time-Series Forecasting Must Adopt Taxonomy-Specific Evaluation to Dispel Illusory Gains ​
Author: Raeid Saqur, Christoph Bergmeir, Blanka Horvath, Daniel Schmidt, Frank Rudzicz, Terry Lyons
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2603.15506v2 Announce Type: replace-cross Abstract: We argue that the current practice of evaluating AI/ML time-series forecasting models, predominantly on benchmarks characterized by strong, persistent periodicities and seasonalities, obscures real progress by overlooking the performance of e...
254. Improving Attributed Long-form Question Answering with Intent Awareness ​
Author: Xinran Zhao, Aakanksha Naik, Jay DeYoung, Joseph Chee Chang, Jena D. Hwang, Tongshuang Wu, Varsha Kishore
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.27435v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly being used to generate comprehensive, knowledge-intensive reports. However, while these models are trained on diverse academic papers and reports, they are not exposed to the reasoning processes a...
255. CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment ​
Author: Anurag Kumar, Raghuveer Peri, Jon Burnsky, Alexandru Nelus, Rohit Paturi, Srikanth Vishnubhotla, Yanjun Qi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.00310v2 Announce Type: replace-cross Abstract: Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Models aligned on text alone show a higher rate of successful attacks when extended to two or more moda...
256. Cluster Attention for Graph Machine Learning ​
Author: Oleg Platonov, Liudmila Prokhorenkova
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.07492v2 Announce Type: replace-cross Abstract: Message Passing Neural Networks have recently become the most popular approach to graph machine learning tasks; however, their receptive field is limited by the number of message passing layers. To increase the receptive field, Graph Transfor...
257. GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking ​
Author: Yunqiang Wang, Hengyuan Na, Di Wu, Miao Hu, Guocong Quan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2604.09222v2 Announce Type: replace-cross Abstract: Audio Large Language Models (ALLMs) enable spoken interaction but introduce new jailbreak vulnerabilities. Existing perturbation-based jailbreaks do not explicitly control which frequency bands carry the perturbation. Although such perturbati...
258. From Plan to Action: How Well Do Agents Follow the Plan? ​
Author: Shuyang Liu, Saman Dehghan, Jatin Ganhotra, Martin Hirzel, Reyhaneh Jabbarvand
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL
arXiv:2604.12147v3 Announce Type: replace-cross Abstract: Agents are commonly instructed to follow a task-specific plan for guidance. However, it is unknown to what extent agents actually follow instructed plans. Without such an analysis, determining the extent agents comply with a given plan, it is...
259. Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference ​
Author: Sanjeev Rao Ganjihal
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.DC, cs.PF
arXiv:2604.26968v2 Announce Type: replace-cross Abstract: Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference serving. Current systems suffer from three compounding inefficiencies: (1) the absence of unified KV cache s...
260. SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages ​
Author: Sen Fang, Hongbin Zhong, Yanxin Zhang, Dimitris N. Metaxas
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2605.01720v3 Announce Type: replace-cross Abstract: Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in laboratory settings. While such resources are important for semantic understanding, they do not...
261. Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages ​
Author: Kevin Guan, Happy Buzaaba, Christiane Fellbaum
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2605.02608v2 Announce Type: replace-cross Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers---the Biaffine LSTM,...
262. Playing Games with My Heart: An Evaluation of AI Companion Apps ​
Author: Maribeth Rauh, Dick A. H. Blankvoort, Matias Duran, Caoilfhionn N'i Dheor'ain, Harshvardhan J. Pandit, Syrine Enneifer, Siddharth D. Jaiswal, Anthony Ventresque, Abeba Birhane
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2605.08093v2 Announce Type: replace-cross Abstract: The use of chatbots for various forms of companionship is growing rapidly, raising a myriad of questions about simulated relationships, emotional dependence, and psychological harm. While major platforms such as ChatGPT, Grok, and Character A...
263. On Seeding Watermarks to Detect Verbatim LLM Copy-Paste Responses ​
Author: Aizierjiang Aiersilan, Artin Yousefi, Robert Pless
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CY
arXiv:2605.16336v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have made fluent essay writing, code drafting, and quiz answering instantly available to students at every level, from secondary school through graduate study. Many educators do not object to LLM use \emph{per~se}...
264. PULSE: Agentic Investigation with Passive Sensing for Proactive Affective Intervention in Cancer Survivorship ​
Author: Zhiyuan Wang, Subigya Nepal, Ariful Islam, Indrajeet Ghosh, Xinyu Chen, Katharine E. Daniel, Laura E. Barnes, Philip Chow
Published: 8/10/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2605.17679v2 Announce Type: replace-cross Abstract: Cancer survivors face elevated rates of depression, anxiety, and emotional distress, yet self-report may be unavailable at some moments when support is relevant, a challenge we term the diary paradox. We present PULSE, a system for agentic se...
265. Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions ​
Author: Volodymyr Ovcharov
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.29738v2 Announce Type: replace-cross Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual comparison impossible. We introduce Multi-Legal-Bench, the first cross-jurisdictional legal...
266. Rethinking Evaluation Paradigms in IBP-based Certified Training ​
Author: Konstantin Kaulen, Hadar Shavit, Holger H. Hoos
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2606.02134v2 Announce Type: replace-cross Abstract: Deep neural networks achieve strong performance on many supervised learning tasks but remain vulnerable to adversarial perturbations. Neural network verification provides mathematically rigorous robustness guarantees, yet at substantial compu...
267. Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path ​
Author: Thomas Sesmat, Gabriel Meseguer-Brocal, Geoffroy Peeters
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SD
arXiv:2606.07271v2 Announce Type: replace-cross Abstract: Understanding memorization in generative models remains challenging, with implications for copyright and privacy. Beyond verbatim reproduction, models can encode subtler traces of their training data that never surface in their outputs yet re...
268. The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products ​
Author: Hao-Ping Lee, Jessica He, David Piorkowski, Thomas Serban von Davier, Jodi Forlizzi, Sauvik Das
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC, cs.LG, cs.SE
arXiv:2606.15485v2 Announce Type: replace-cross Abstract: Agentic AI systems act autonomously, use tools, adapt to context, and operate in complex real-world environments. However, these same characteristics can create or exacerbate product risks. We studied how industry developers (n=35) perceive, ...
269. An Empirical Study of openPangu Quantization on Ascend NPUs ​
Author: Tong Shi, Jiacheng Wang, Hui Xie, Ying Li, Aishan Liu, Jinyang Guo, Xianglong Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.21257v4 Announce Type: replace-cross Abstract: openPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive post-training quantization on Ascend NPUs has not been systematically characterized. This paper conducts a...
270. Compositional Behavioral Semantics for State Abstraction in Reinforcement Learning ​
Author: Yivan Zhang, Ziyan Luo, Manuel Baltieri
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.CT
arXiv:2606.25357v2 Announce Type: replace-cross Abstract: State abstraction plays a key role in scaling reinforcement learning to complex but structured systems. In studying such systems, a wide range of behavioral structures have been studied in reinforcement learning, including value functions, in...
271. LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models ​
Author: Sojung An, Junha Lee, Sujeong You, Nam Ik Cho, Donghyun Kim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.06918v2 Announce Type: replace-cross Abstract: Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM adaptation stems from the prohibitive costs of full fine-tuning and catastrophic forgetting. To address t...
272. MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs ​
Author: Hantao Zhang, Jinru Sui, Ed Li, Dirk Bergemann, Zhuoran Yang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.08970v2 Announce Type: replace-cross Abstract: Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce M...
273. A Physics-Inspired Classical Digital Twin of Cortical Dynamics: A Band-Stratified Metriplectic Port-Hamiltonian Neural Network Learned from Brain-Computer-Interface EEG ​
Author: Dibakar Sigdel
Published: 8/10/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI
arXiv:2607.10439v3 Announce Type: replace-cross Abstract: We present a physics-inspired classical digital twin of brain-computer- interface (BCI) data: a graph neural network constrained to a band-stratified, metriplectic port-Hamiltonian form, with parameters learned from scalp EEG recorded during ...
274. DeepLoop: Depth Scaling for Looped Transformers ​
Author: Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.13491v2 Announce Type: replace-cross Abstract: Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied...
275. Counterfactual Shapley Credit Assignment ​
Author: Mingxuan Li, Kai-Zhan Lee, Elias Bareinboim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.16999v2 Announce Type: replace-cross Abstract: The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing frameworks, whether relying on temporal contiguity or hindsight-conditioned reward reweighting, frequently...
276. The Ethics of Autonomous AI Agents for Offensive Security ​
Author: Andreas Happe, J"urgen Cito, Jasmin Wachter
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2607.20255v2 Announce Type: replace-cross Abstract: LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling - deterministic, narrowly scoped, and operated by trained practitioners - agentic security tools exhibit indeterminacy along three i...
277. Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention ​
Author: Aryan Sood, Shantanu Acharya, Gaurav Kumar Nayak
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.20457v2 Announce Type: replace-cross Abstract: Inference with large language models (LLMs) on long sequences is computationally expensive due to the quadratic complexity of self-attention. Distributed blockwise methods such as Star Attention reduce this cost by sharding context across hos...
278. IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning ​
Author: Wei Zhang, Xinwu Liu, Yihang Cheng
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22251v2 Announce Type: replace-cross Abstract: Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT) of LLMs whose effectiveness depends on rank allocation. Existing adaptive LoRA methods derive ranks from local gradient, activation, or matrix stat...
279. Automated Numerical Stability Analysis of Deep Learning Operators ​
Author: Xinye Chen
Published: 8/10/2026, 4:00:00 AM
Categories: math.NA, cs.AI, cs.NA
arXiv:2607.25494v3 Announce Type: replace-cross Abstract: Finite-precision arithmetic unavoidably introduces numerical approximation errors. Numerical computations may use insufficient precision or an improper formulation, which leads to numerical instability. In this paper, we introduce a unified s...
280. F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill ​
Author: Florian Krebs
Published: 8/10/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.SI
arXiv:2607.25637v2 Announce Type: replace-cross Abstract: F(AI)2R is FAIR research with AI in the loop, twice: an AI-assisted authoring pass and a machine-readable audit pass over every artefact. AI systems now draft, refactor, and verify research artefacts, yet their contributions are rarely record...
281. FinanceHarness: Autonomous Financial Deep Research Framework ​
Author: Yijia Xiao, Rujun Han, Yanfei Chen, Zifeng Wang, Ke Jiang, Zhongying CuiZhu, Vishy Tirumalashetty, Wei Wang, Burak Gokturk, Tomas Pfister, Chen-Yu Lee
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, q-fin.CP
arXiv:2607.27853v2 Announce Type: replace-cross Abstract: Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. ...
282. Topology-Aware Data Movement for Disaggregated GPU Inference ​
Author: Sanjeev Rao Ganjihal
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.PF
arXiv:2607.28633v2 Announce Type: replace-cross Abstract: Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run on separate GPU pools, the KV cache must be transferred between them. For a 70B model this is 1.3 GB per...
283. A Fortran General-Purpose Transpiler: Proof of Concept ​
Author: Shivamshan Sivanesan, Kazem Ardaneh
Published: 8/10/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.CL, cs.MS, cs.SE
arXiv:2608.00130v2 Announce Type: replace-cross Abstract: Fortran has been the cornerstone of high-performance computing for decades and remains unmatched in many domains. Yet the language faces an expertise gap: a new generation of scientists is barely familiar with it, while many experienced Fortr...
284. Rethinking and formalising the state across languages: a unified computational learning theory account ​
Author: Mohamed El Idrissi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.00523v2 Announce Type: replace-cross Abstract: The linguistic notion of state has traditionally been restricted to the construct (annexation) state of Afroasiatic languages and treated as a language-specific morphosyntactic phenomenon. This article argues instead that the state is a syste...
285. WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA ​
Author: Zhihao Zhu, Hanlin Shang, Mingwang Xu, Feipeng Cai, Zhuolin He, Yaoyi Li, Jianhua Han, Hang Xu, Siyu Zhu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV
arXiv:2608.01035v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high computational latency and exposure bias arising from sequential au...
286. Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Measure for Infrared-Visible Fusion Assessment ​
Author: Haoran Liu, Mingzhe Liu, Peng Li, Guibin Zan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.01301v3 Announce Type: replace-cross Abstract: Infrared-visible image fusion (IVIF) has no ideal fused reference, so algorithms are ranked by scalar objective metrics that formalize proxies for information transfer, structure, or source similarity. These proxies often disagree with the ju...
287. A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation ​
Author: Abdallah Khemais
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.03620v2 Announce Type: replace-cross Abstract: Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different objects: one forward pass versus the parameters behind every forward pass. We ask when they agree. We study...
288. SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant ​
Author: Adel Javanmard, David P. Woodruff, Vahab Mirrokni
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2608.05127v2 Announce Type: replace-cross Abstract: Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD, use high-dimensional geometric constructions but incur unfavor...
289. A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper ​
Author: Ali Shendabadi, Parnia Izadirad, Mostafa Salehi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.SD
arXiv:2608.05165v2 Announce Type: replace-cross Abstract: Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we study the use of Whisper for Persian SER with a particular focus on representation dimensionality reduction...
290. Challenges for Musical Education in the Age of AI and Digital Transformation ​
Author: Jean-Pierre Briot
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG
arXiv:2608.05176v2 Announce Type: replace-cross Abstract: Music education has never been a static discipline. Each major technological shift has forced educators and institutions to reconsider what they teach, how they teach it, and why. We now stand at what may be the most consequential of such tur...
291. PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis ​
Author: Xiaomin He, Dongling Xiao, Jiahao Xie, Ruiqi Lu, Qianle Wang, Zhongbin Guo, Wanxuan Sun
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.05249v2 Announce Type: replace-cross Abstract: Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering one self-contained question. We study this gap through rubric com...
292. When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems ​
Author: Jialuo Chen, Lingqi Jiang, Xinhao Deng, Xiaohu Du, Jianan Ma, Yunhao Feng, Yuqi Qing, Zhihao Yuan, Linkang Du, Jingyi Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.05563v2 Announce Type: replace-cross Abstract: Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We introduce PoisonedEvolution, a trajectory-poisoning attack on this promotion process. Our ski...
293. HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection ​
Author: Aohua Li, Jin Kuang, Yubing Lu, Pingping Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.05771v2 Announce Type: replace-cross Abstract: Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance often degrades markedly when generalizing to unseen infrared domains. Existing methods primarily improve de...
294. Is Self-Pretraining really useful to improve diagnosis in medical Time Series? ​
Author: Omar Coser, Antonio Orvieto, Paolo Soda, Loredana Zollo
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06122v2 Announce Type: replace-cross Abstract: Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series...
295. Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset ​
Author: Eoin Cummins, Zhongyi Huang, Alexandre D'Hooge, Zhuoru Mo, Yaolong Ju
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM
arXiv:2608.06165v2 Announce Type: replace-cross Abstract: Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first presents the new SheetSage-A2S Dataset, which includes 61 hours of audio with \texttt{**ker...