arXiv cs.AI - 2026-07-20 ​
201 items collected.
1. GraphDx: A Cost-Aware Knowledge-Enhanced Multi-Agent Framework for Sequential Diagnosis ​
Author: Shaoting Tan, Ning Liu, Yuntao Du, Shuyue Wei, Wu Shuai, Qian Li, Yanyu Xu, Wei Zhang, Lizhen Cui, Haitao Yuan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15280v1 Announce Type: new Abstract: Sequential diagnosis requires balancing diagnostic accuracy against resource costs through iterative information gathering. Existing Large Language Model (LLM) approaches exhibit a critical knowledge-reasoning gap: despite encoding extensive medical kn...
2. Causal-Audit: Explicit and Auditable Graph-based Reasoning via Target-Aware Causal Chain Construction ​
Author: Su Lan, Xuefei Yin, Yanming Zhu, Alan Wee-Chung Liew
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15281v1 Announce Type: new Abstract: Causal and intervention-based question answering is fundamental to advancing large language models (LLMs) toward reasoning beyond surface-level correlations and understanding underlying causal mechanisms. However, existing LLM-based methods often rely ...
3. Cura 1T: Specialized Model for Agentic Healthcare ​
Author: actAVA AI, :, Haolin Chen, Leon Qi, Steve Brown, Deon Metelski, Tao Xia, Joonyul Lee, Qixuan Wang, Kevin Riley, Frank Wang, Weiran Yao
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15314v1 Announce Type: new Abstract: Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images,...
4. AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery ​
Author: Raunak B Sinha
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15367v1 Announce Type: new Abstract: Desktop voice assistants are still dominated by cloud pipelines that ship raw audio off the machine and expose a fixed set of skills. We describe AnovaX, a small local-first assistant that runs entirely on the user's computer and treats the desktop its...
5. Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning ​
Author: Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan, Cheng-Hau Yang, Huihuo Zheng, Le Chen, Eliu A. Huerta, Venkatram Vishwanath, Ian T. Foster, Rajeev Thakur
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15388v1 Announce Type: new Abstract: Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates into correct ones. We test this assumption on 4,181 verifier-grounded Omni-...
6. DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings ​
Author: Yoonhwa Jung, Junryu Fu, Mani Golparvar-Fard
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2607.15418v1 Announce Type: new Abstract: We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices. Unlike natural images or schema...
7. Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3? ​
Author: Sergey Rodionov
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15439v1 Announce Type: new Abstract: Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea accounted for its performance. We address this attribution question with four nested Codex-based agents:...
8. Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes ​
Author: Shanhong Liu, Pai Chet Ng, De Wen Soh, Malika Meghjani, Konstantinos N. Plataniotis
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15442v1 Announce Type: new Abstract: Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist. These complexities highlight the need for explainable meme un...
9. From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems ​
Author: Eduardo C. Garrido-Merch'an
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15459v1 Announce Type: new Abstract: A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer...
10. A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chasms ​
Author: Michael Papademas, Xenia Ziouvelou, Kostas Karpouzis, Vangelis Karkaletsis
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15480v1 Announce Type: new Abstract: As artificial intelligence (AI) systems increasingly impact society, ensuring their ethical and trustworthy deployment has become a global priority. While a myriad of high-level ethical guidelines have emerged, criticism persists that these frameworks ...
11. Logic, Optimization, and Artificial Intelligence ​
Author: J. N. Hooker
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, math.OC
arXiv:2607.15532v1 Announce Type: new Abstract: Logic and optimization can, in combination, make valuable contributions to rule-based AI. Logic is the obvious medium for encoding a rule base and drawing inferences from it, while optimization provides a powerful technology for computing inferences. T...
12. SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction ​
Author: Xue Yu, Bo Yuan, Pengshuai Yang, Kailin Zhao, Hong Hu, Junlan Feng
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15550v1 Announce Type: new Abstract: Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks where a single erroneous action can lead to irreversible consequences. Existing safety mechanis...
13. MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion ​
Author: Xu Hou, Meiyu Liang, Wei Huang, Yawen Li, Zhe Xue, Wu Liu, Guanhua Ye, Lei Shi, Kangkang Lu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15592v1 Announce Type: new Abstract: Multimodal Knowledge Graph Completion (MKGC) requires inferring missing entities from structural, textual, and visual cues. Existing diffusion-based MKGC methods usually denoise directly on raw multimodal features. Such a design forces the denoiser to ...
14. Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts ​
Author: Aritro De (The University of Texas at Austin), Juliana Felkner (The University of Texas at Austin)
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15647v1 Announce Type: new Abstract: LEED v4.1 BD+C certification remains a document-intensive process that requires reviewers to read hundreds of pages of project evidence and apply credit-specific threshold logic by hand. This paper investigates whether small, locally deployed language ...
15. ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning ​
Author: Shuaiyu Zhou, Fengpeng Yue, Zengjie Hu, Yuanzhe Shen, Chenyang Zhang, feng hong, Cao Liu, Ke Zeng
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15660v1 Announce Type: new Abstract: While LLM agents demonstrate strong reasoning abilities in compact and well-defined scenarios, they struggle to maintain robustness and effectiveness when faced with large-scale, diverse, and dynamic real-world environments that demand seamless tool in...
16. S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation ​
Author: Jiahao Zhao, Junyi Liu, Lifeng Xu, Nan Xu, Qingli Wang, Qingxiao Li, Tianle Chen, Xiaoyu Wu, Yawen Zheng, Zikai Wang, Guanming Liu, Hequn Zhou, Jingyi Wang, Jingyuan Shu, Keqi Wang, Li He, Songyang Diao, Wenhui Xu, Xinyu Ren, Yaqin Fan, Yujin Zhou, Zhanao Yao
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15686v1 Announce Type: new Abstract: We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific language models. How...
17. Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents ​
Author: Lujia Zhang, Xingzhou Chen, Hongwei Feng
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15715v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components such as reflection and memory lead to observable and controllable improvements over fixed LLM workflows....
18. NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning ​
Author: Hui Yang, Jiaoyan Chen, Yiping Song, Renate Schmidt, Wen Zhang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15776v1 Announce Type: new Abstract: OWL ontologies provide a formal knowledge representation framework that enables semantic reasoning, and have been widely adopted across domains such as healthcare and bioinformatics. In practice, however, real-world ontologies are often incomplete, whi...
19. AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets ​
Author: Ming Chen, Pranav Pai
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, cs.MA
arXiv:2607.15781v1 Announce Type: new Abstract: Geospatial datasets support applications from urban planning to climate modeling, yet consistent assessment of FAIR compliance is difficult. Existing evaluators use different rubrics and evidence sources and may fail on JavaScript-rendered pages or rep...
20. Knowledge-Centric Agents for Workflow Generation ​
Author: Zhendong Li, Lei Sun, Ruibo Ming, He Zhang, Danda Pani Paudel, Luc Van Gool, Jinjin Gu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15845v1 Announce Type: new Abstract: Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large language model (LLM) approaches often treat this as a direct text-to-JSON gener...
21. DSWorld: A Data Science World Model for Efficient Autonomous Agents ​
Author: Zherui Yang, Fan Liu, Hao Liu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15901v1 Announce Type: new Abstract: Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottleneck motivates models that can anticipate the effects o...
22. A Formally Grounded ODRL Evaluator: Implementation and Comparison ​
Author: Jaime Osvaldo Salas, Paolo Pareti, Adeel Aslam, Christopher Maidens, George Konstantinidis
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15987v1 Announce Type: new Abstract: The ODRL policy language is emerging as the de-facto standard for policy modelling data access and usage preferences, AI governance policies and data workflows in European dataspaces. The current standard has no mathematical formal semantics to describ...
23. Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI ​
Author: Trisevgeni Papakonstantinou, Cansu Canca, Farah Nanji, Waheedullah Pardess, Jen Weedon, Jasmijn Remmers, Eliza Krigman, Matthew Ball, Yalda Daryani, Kiran Iqbal, Francielle Vargas, Mar'ia Llorente S'anchez, Joe Humphreys, Fendi Tsim, Kelly Fitzpatrick, Jeff Dunn, Catherine Feldman
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15992v1 Announce Type: new Abstract: Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in high-stakes settings. Yet this work has not produced a market that rewards trustworthiness. Firms that invest se...
24. SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery ​
Author: SciForge Team, Zhangyang Gao, Minghao Fang, Yifei Liu, Hanhui Yang, Xinyu Gu, Shixiang Tang, Siqi Sun, Lei Bai, Cheng Tan, Mengdi Liu, Hao Wu, Shuizhou Chen
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.16038v1 Announce Type: new Abstract: Scientific work increasingly spans heterogeneous artifacts -- papers, code, datasets, scientific file formats, model outputs, figures, manuscripts, and team decisions -- yet general-purpose AI assistants rarely preserve these objects as a coherent, aud...
25. Harmonizing AI Safety Thresholds ​
Author: Wilber Sean Anterola, Matthew Ball, Luis F. Lafuerza, Markov Grey
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.16112v1 Announce Type: new Abstract: Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thr...
26. CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data ​
Author: Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru, MohammadHossein Rezaei, Aakash Sabharwal, Yunzhong He
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.16122v1 Announce Type: new Abstract: Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak examples, topics, or c...
27. Empathy as Predictive Misalignment Tolerance: A Co-Regulation Framework and the Regime Structure of Dialogue Repair ​
Author: Molood Arman
Published: 7/20/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL
arXiv:2607.15282v1 Announce Type: cross Abstract: Empathy is most often theorized as resonance: a mirroring of another's present emotional or cognitive state. This synchronic framing has shaped artificial systems, where empathic behavior is defined as affect recognition and response alignment. We ar...
28. How Does Empowering Users with Greater System Control Affect News Filter Bubbles? ​
Author: Ping Liu, Karthik Shivaram, Aron Culotta, Matthew Shapiro, Mustafa Bilgic
Published: 7/20/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2607.15284v1 Announce Type: cross Abstract: While recommendation systems enable users to find articles of interest, they can also create ``filter bubbles'' by presenting content that reinforces users' pre-existing beliefs. Users are often unaware that the system placed them in a filter bubble ...
29. Structure of the Circular-Dyadic Convolution Error ​
Author: Ben Fauber, Alireza Moradzadeh
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NA, math.NA
arXiv:2607.15293v1 Announce Type: cross Abstract: Dyadic and circular convolution can both be computed in $O(N\log N)$ time using the Hadamard transform and the FFT-computed discrete Fourier transform (DFT), respectively. The Hadamard transform is preferable for its real-valued sign flips, yet its s...
30. AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning ​
Author: Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin
Published: 7/20/2026, 4:00:00 AM
Categories: cs.MM, cs.AI, cs.LG, cs.SD
arXiv:2607.15295v1 Announce Type: cross Abstract: We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of global and per-modality ...
31. Data-driven Video Codec with Implicit Neural Representations ​
Author: Nishan Khanal, Saugat Neupane, Abhinav Chalise, Nimesh Gopal Pradhan, Dinesh Baniya Kshatri
Published: 7/20/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.MM, cs.SD
arXiv:2607.15298v1 Announce Type: cross Abstract: A conventional codec stores a video as compressed pixel data. We instead store the video, together with its audio track, as the weights of a single sinusoidal representation network (SIREN) that maps space-time coordinates to RGB values and audio amp...
32. Lazy Arithmetic using Systolic Arrays for Closing the Verification Gap on Embedded Systems ​
Author: Taisa Kushner (Galois Inc), Ryan McCleeary (Galois Inc), Martin Brain (City St George University of London)
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.AR
arXiv:2607.15328v1 Announce Type: cross Abstract: Complex algorithms such as deep neural networks are increasingly being deployed on embedded, resource constrained platforms. However, existing hardware and software schemes for implementing these models on the edge fall short, particularly for safety...
33. Large Language Models as Unified Multimodal Learners for Clinical Prediction ​
Author: Ajay Madhavan Ravichandran, Bilgin Osmandoja, Klemens Budde, Klaus Netter, Tobias Strapatsas, Aljoscha Burchardt, Sebastian M"oller, Roland Roller
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.15380v1 Announce Type: cross Abstract: Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities. Yet most clinical prediction systems still rely on task-specific fusion architectures, pairing ded...
34. Partial Information Decomposition as a Multi-Contrast 3D MRI Selection Strategy for Resource-Constrained Deep Neural Network Training in Brain Tumor Segmentation ​
Author: Agamdeep Chopra, Mehmet Kurt
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.15396v1 Announce Type: cross Abstract: Multi-contrast 3D MRI segmentation can be computationally demanding when all available sequences are used. We evaluate a pre-training Partial Information Decomposition framework that ranks input pairs according to their redundant, unique, and synergi...
35. AI Trading: Evaluating Large Language Models for Technical Market Analysis ​
Author: Geofrey Ntale
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-fin.CP
arXiv:2607.15414v1 Announce Type: cross Abstract: Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial markets. This paper presents a systematic, comparative evaluation of five prominent LLMs: GPT-4 Turbo, Claude 3 ...
36. Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation ​
Author: Jasmine Brazilek, Maheep Chaudhary, Zoe Lu, Miles Tidmarsh
Published: 7/20/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CR
arXiv:2607.15434v1 Announce Type: cross Abstract: Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No bench...
37. Design-Based Supervised Learning with Noisy Human Labels ​
Author: Robert Chew, Matthew R. Williams
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, stat.AP
arXiv:2607.15455v1 Announce Type: cross Abstract: Researchers increasingly use automated classifiers to label unstructured data for statistical analysis. Existing rectification methods can correct errors in these automated labels using a probability-sampled audit set, but they usually treat the audi...
38. FLINT: Fingerprinting Federated Learning Architectures from 5G PHY-Layer Side Channels ​
Author: Md Nahid Hasan Shuvo, Mahmudul Hassan Ashik, Moinul Hossain
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2607.15469v1 Announce Type: cross Abstract: Federated Learning (FL) over 5G cellular networks protects raw data but remains vulnerable to side-channel leakage. Prior fingerprinting attacks assume packet-level network visibility, an assumption that does not hold at the 5G Physical (PHY) layer, ...
39. Verbalizable Representations Form a Global Workspace in Language Models ​
Author: Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, Jack Lindsey
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.15495v1 Announce Type: cross Abstract: Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous function...
40. LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4 ​
Author: Mobina Kashaniyan, Amirhossein Ghassemi, Nasser Mozayani
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.15509v1 Announce Type: cross Abstract: We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition. Each large language model independently ge...
41. An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism ​
Author: Mobina Kashaniyan, Mehrdad Ashtiani, Amirhossein Ghassemi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.PF
arXiv:2607.15511v1 Announce Type: cross Abstract: Serverless computing provides automatic resource management and pay-per-use execution, but effective autoscaling remains challenging because of dynamic workloads, cold-start latency, and dependencies among functions. We present a dependency-aware aut...
42. Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching ​
Author: Yan Song
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15516v1 Announce Type: cross Abstract: Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent). The compression literature has standardized on query-aware methods that produc...
43. SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification ​
Author: Bibesh Pyakurel, M. G. Sarwar Murshed
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.15517v1 Announce Type: cross Abstract: Four-finger SLAP fingerprints are flat live-scan impressions of the index, middle, ring, and little fingers of one hand, used for identity verification in border control and law enforcement. No benchmark has evaluated whether multimodal large languag...
44. Recursive Harness Self-Improvement ​
Author: Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Matei Zaharia, Yujin Tang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15524v1 Announce Type: cross Abstract: Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both ...
45. Kolmogorov--Arnold Networks for Small Language Models ​
Author: Felippe Alves, Renato Vicente
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15525v1 Announce Type: cross Abstract: Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks. We test these claims sep...
46. CoWeaver: A Bi-directional, Learnable and Explainable Matching Engine for Mixed Human-Agent Science Collaboration ​
Author: Jiayao Gu, Kexin Chu, Peidong Liu, Yue Yang, Lynn Ai, Qi Zhang, Ling Yang, Tianyu Shi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CL
arXiv:2607.15545v1 Announce Type: cross Abstract: LLM-based agents excel at writing articles, coding and information retrieval. However, they fail to form strong collaborations within the scientific community due to the bidirectional, dynamic nature of the problem and a high demand of decision inter...
47. From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation ​
Author: Himel Dev, Tanmoy Sen, Madhusudan Basak, Bashima Islam
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15552v1 Announce Type: cross Abstract: Generating personalized trip itineraries is a complex planning task and involves a tension between hard combinatorial feasibility and soft latent desirability. Classical optimization enforces constraints but fails to capture subjective traveler prefe...
48. Evolutionary Algorithm-Guided LLMs for Physics-Informed Neural Network Design ​
Author: Xu Yang, Mingyang Yu, Jing Xu, Keqian Li
Published: 7/20/2026, 4:00:00 AM
Categories: cs.NE, cs.AI
arXiv:2607.15560v1 Announce Type: cross Abstract: Physics-informed neural networks (PINNs) are unusually sensitive to interacting choices of architecture, activation, loss weighting, collocation, optimization, and constraint enforcement. Large language models (LLMs) can propose these choices, but in...
49. Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation ​
Author: Himel Dev, Madhusudan Basak, Tanmoy Sen, Paromita Shome, Bashima Islam
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15562v1 Announce Type: cross Abstract: Packing for air travel is recurring and error-prone: the checklist must be personal and context-aware, yet feasible under safety rules, item dependencies, and luggage limits. Existing packing assistants are template-driven and generic, or recommendat...
50. Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models ​
Author: Rakshanda Hassan Abhinandan, John Galeotti, Deva Ramanan, Gautam Rajendrakumar Gare
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV
arXiv:2607.15565v1 Announce Type: cross Abstract: Where should the question go in a vision-language model (VLM) prompt: before the image or after it? Intuition says before: knowing what is asked should tell the model where to look. Yet across visual question answering benchmarks, question-first prom...
51. Information-Directed Sampling for Causal Bandits ​
Author: Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15577v1 Announce Type: cross Abstract: Causal bandits exploit structural relationships among variables to share information across interventions and accelerate the identification of high-reward decisions. In many applications, however, some variables cannot be directly manipulated, even t...
52. MemoGuard: An Adaptive Runtime for Guarding Against Memory Traps in Communication-Limited Robot Navigation ​
Author: Rajat Bhattacharjya, Hyeonjong Ju, Sing-Yao Wu, Eli Bozorgzadeh, Nikil Dutt
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.NI, cs.SY, eess.SY
arXiv:2607.15589v1 Announce Type: cross Abstract: Communication-limited robots in mission-critical scenarios such as disaster inspection and search-and-rescue must make reliable onboard decisions without access to remote operators or high-capacity reasoning services. Episodic memory reuse is an attr...
53. Field-Aware RankMixer with Dual-Stream Bilinear Fusion for the Tencent UNI-REC Challenge ​
Author: Yufeng Zhang, Zhengqi Xu, Jiajun Cui
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15590v1 Announce Type: cross Abstract: This paper presents our solution to the KDD Cup 2026 Tencent UNIREC Challenge. The task requires joint modeling of multi-domain user behavior sequences and non-sequential multi-field features for target-ad pCVR prediction. We develop a Field-Aware Ra...
54. Scalable LLM Agent Tool Access in the Cloud ​
Author: Mingxin Li, Enge Song, Yueshang Zuo, Xiaodong Liu, Rong Wen, Qiang Fu, Gianni Antichi, Jian He, Jing Tie, Zhou Shao, Xiaobo Xue, Xiong Xiao, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Zihao Fan, Haonan Li, Tian Pan, Xiaomin Wu, Yang Song, Xing Li, Biao Lyu, Meng Li, Haipeng Dai, Guihai Chen, Shunmin Zhu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.NI
arXiv:2607.15593v1 Announce Type: cross Abstract: LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de facto interface. Operating MCP at cloud scale, however, becomes difficult. On the tool provider side, legacy servi...
55. Process Reward Informed Tree Rollout for Effective Multi-Turn RL ​
Author: Xintong Li, Sha Li, Yuwei Zhang, Changlong Yu, Rongmei Lin, Hongye Jin, Shuyi Guan, Xin Liu, Linwei Li, Qingyu Yin, Jingbo Shang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.15610v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple independently sampled complete trajectories for advantage estimation. In long-horizon agentic tasks, such a uniform ...
56. AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots ​
Author: Priyanka V. Setty, Arvind Ramanathan, Ian Foster, Rick Stevens
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.15620v1 Announce Type: cross Abstract: Self-driving laboratories increasingly rely on low-cost liquid handlers such as the Opentrons OT-2, which ship without the pressure-based aspiration monitoring of Hamilton or Tecan systems and are typically run open-loop. Two failure modes go undetec...
57. Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving ​
Author: Yun Li, Jiachen Gong, Simon Thompson, Ehsan Javanmardi, Qunli Zhang, Zifan Zeng, Shiming Liu, Peng Wang, Zixuan Guo, Manabu Tsukada
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.15621v1 Announce Type: cross Abstract: Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires. Existing closed-loop agents hide this gap by invoking the model on alternate s...
58. A cubical formalisation of topos causal models: intervention, sheaf gluing, and the intuitionistic do-calculus ​
Author: Karen Sargsyan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, math.CT
arXiv:2607.15629v1 Announce Type: cross Abstract: Topos causal models recast causal inference inside a topos: a causal world is a presheaf, an intervention is a characteristic map into the subobject classifier, and reasoning is carried out in the intuitionistic internal language. We give the first m...
59. IMBench: A Benchmark for Intuitive Robotic Manipulation ​
Author: Anurag Maurya, Sukhvansh Jain, Prajwal Avhad, Gautham Balachandran, Ziyi Zhou, Atharva Kshirsagar, Satyam Singh, Bowen Li. Rishabh Mukund, Ritul Singh, Jatin Vira, Suvonil Chatterjee, Devesh K. Jha
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.15641v1 Announce Type: cross Abstract: Humans combine reasoning and motor control to solve complex manipulation tasks under diverse constraints. They build an understanding of the physical world that helps them convert reasoning into actions and quickly adapt to new scenes, tasks, and rul...
60. On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous Levels ​
Author: Taiga Mori, Koji Inoue, Divesh Lala, Tatsuya Kawahara
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2607.15648v1 Announce Type: cross Abstract: In multi-party dialogues between a dialogue system and multiple users, identifying to whom an utterance is addressed is a key challenge. Prior work has typically treated addressee detection as a multi-class classification task, selecting a single lab...
61. Toward a mechanistic understanding of inference in visual cortex and diffusion models ​
Author: Zeyu Yun, Alexander Belsten, Dasheng Bi, Zahra Kadkhodaie, Yubei Chen, Bruno A. Olshausen
Published: 7/20/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI
arXiv:2607.15693v1 Announce Type: cross Abstract: We describe a model of perceptual inference in primary visual cortex (V1) equivalent to a minimal diffusion model whose function can be readily understood from its parameters. The model is based on sparse coding with a non-factorial prior over latent...
62. Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution ​
Author: Xue Wu, Kang Zhao, Kafeng Wang, Jianfei Chen, Jingwei Xin, Nannan Wang, Xinbo Gao
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.15711v1 Announce Type: cross Abstract: Diffusion-based methods have achieved impressive performance in real-world image super-resolution (Real-ISR) by leveraging large pre-trained stable diffusion (SD) models as powerful generative priors. However, these methods still face two key limitat...
63. Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization ​
Author: Yong Chu, Xun Zhou, Zenglin Xu, Hui Wang, Yue Yu
Published: 7/20/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG
arXiv:2607.15713v1 Announce Type: cross Abstract: Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing. Despite its importance, achieving precise localization across diverse environmen...
64. Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling ​
Author: Bo-An Chang, Yu-Chih Chen
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.MM
arXiv:2607.15740v1 Announce Type: cross Abstract: As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI. Existing T2I evaluation metrics and multimodal judges often rely on ...
65. AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis ​
Author: Zhenqi Jia, Yuan Zhao, Aruukhan, Rui Liu, Haizhou Li
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2607.15755v1 Announce Type: cross Abstract: Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predefined emo...
66. GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing ​
Author: Yujie Li, Jiancheng Pan, Zhiwei Wei, Jiuniu Wang, Mugen Peng, Wenjia Xu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.15768v1 Announce Type: cross Abstract: Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated moments, but also track changes, memorize evolution histories, and reason ...
67. Scaling Time Series Classification via XAI-Driven Data Reduction ​
Author: Davide Italo Serramazza, Thach Le Nguyen, Georgiana Ifrim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15774v1 Announce Type: cross Abstract: Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI, a novel methodology...
68. AquaAugmentor: A Novel Feature Augmentation Algorithm for Water Potability Prediction ​
Author: Muntasir Tabasum, Al Zadid Sultan Bin Habib, Tanpia Tasnim, Md. Ekramul Islam, Md Younus Ahamed, Md Asif Bin Syed
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE
arXiv:2607.15775v1 Announce Type: cross Abstract: Access to potable water is crucial for health, economic development, and sustainability. However, accurately classifying water quality remains a significant challenge due to the complexity and variability of water source data. This paper addresses th...
69. Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding ​
Author: Wei Feng, Xin Wang, Yu-Wei Zhan, Yuwei Zhou, Wenwu Zhu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.15778v1 Announce Type: cross Abstract: Video Large Language Models (Video LLMs) have made significant advancements in various video understanding tasks. However, long-video scenarios remain challenging due to the tension between limited visual token budgets and the need to capture multipl...
70. On the Geometry of Learned Representations in Event-Based Multi-Modal Egomotion Estimation ​
Author: Stefano Silvestrini, Michele Ceresoli
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.15794v1 Announce Type: cross Abstract: Classical approaches to event-based egomotion estimation, including those adopted by the top-performing teams of the ELOPE challenge, rely on geometric optimization frameworks such as contrast maximization, homography estimation, or dense optical flo...
71. Knowledge-Assisted Multi-Graph Dependency Learning for Multivariate Time Series Anomaly Detection in Multi-Stage Industrial Processes ​
Author: Jaeyeong Lee, Taeseong Yoon, Wonmo Koo, Heeyoung Kim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15799v1 Announce Type: cross Abstract: Industrial processes often generate complex, interdependent time-series data from multiple sensors across multiple stages, forming complex dependencies among variables and process stages. Effective monitoring and timely anomaly detection of these tim...
72. In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention ​
Author: Katsuyuki Hagiwara
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15819v1 Announce Type: cross Abstract: In-context learning is a remarkable property of transformers and has recently received a lot of interest. In many studies of in-context learning, it has been shown that transformers are capable of implementing solver for linear and non-linear regress...
73. RTL-Sequencer: Towards Scalable RTL Timing Prediction with the Sequence-based Paradigm ​
Author: Ziyan Guo, Wenji Fang, Wenkai Li, Yuchao Wu, Shang Liu, Zhiyao Xie
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.LG
arXiv:2607.15830v1 Announce Type: cross Abstract: Accurate timing prediction at the register-transfer level (RTL) is a longstanding challenge in design automation. Existing graph-based methods struggle with limited receptive fields, high complexity, and a lack of signal directionality. We present RT...
74. CAMMAR: Culture-Aware Matryoshka for Metaphorical Arabic Representations ​
Author: Suzan Awinat, Alfonso Ortega del Puente
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.15847v1 Announce Type: cross Abstract: Metaphor in Arabic is a culturally grounded mechanism for constructing meaning, encoding cultural knowledge that shapes interpretation. Yet current Arabic language models typically collapse lexical, cultural, and metaphorical information into a singl...
75. Test-Time Noise Guided Adaptation for Realistic Autoregressive Video Generation ​
Author: Dimitrios Karageorgiou, Symeon Papadopoulos, Ioannis Kompatsiaris, Efstratios Gavves
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.15849v1 Announce Type: cross Abstract: Autoregressive video diffusion models have enabled the generation of arbitrarily long videos by removing conditioning on future frames, thus greatly improving computational efficiency. Yet, they suffer from error accumulation over time, as the denois...
76. Agentic Synthesis against Counterexample-Supplemented Sketches ​
Author: Muness Castle, Eric Rubeck
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.15854v1 Announce Type: cross Abstract: Coding agents can fix a failing example without preserving the domain rule that made it fail, so later generations can repeat the same plausible mistake. We present agentic synthesis against counterexample-supplemented sketches, a repository-native m...
77. Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection ​
Author: Indraveni Chebolu, Rohan Singh, Arnab Mallick, Harmesh Rana
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.15861v1 Announce Type: cross Abstract: Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch. We study the \emph{conditional reliability} of toxicity priors in Indian multilingual an...
78. EgoExoMoCap: Distributed Ego-Exo Human Motion Capture ​
Author: Jiaxi Jiang, Bharat Lal Bhatnagar, Nan Yang, Lingni Ma, Sebastian Starke, Robin Kips, Nadine Bertsch, Christian Holz, Federica Bogo
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.GR, cs.HC, cs.RO
arXiv:2607.15868v1 Announce Type: cross Abstract: Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which is crucial for applications in embodied AI and VR/AR. Existing approaches focus on either egocentric body track...
79. DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods ​
Author: Jens Frankenreiter
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2607.15879v1 Announce Type: cross Abstract: Much empirical legal research depends on translating unstructured text into structured variables. In corporate governance research as elsewhere, this translation has traditionally relied on human coding of documents such as charters and bylaws, a pro...
80. Perceived AGI: Believability as Dimensional Completeness, Not Capability ​
Author: Sebastian Cochinescu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2607.15883v1 Announce Type: cross Abstract: Large language models are broadly capable, yet in sustained one-to-one conversation they still read as flat: competent, responsive, and somehow not quite the presence of a mind. We hypothesize that a central missing ingredient is not more capability ...
81. Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models ​
Author: Andy Catruna, Emilian Radoi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.15893v1 Announce Type: cross Abstract: While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is known about diffusion language models (DLMs), an emerging alternative that generates text by iterative denoising. In this work, we study how...
82. Orbis 2: A Hierarchical World Model for Driving ​
Author: Sudhanshu Mittal, Arian Mousakhan, Silvio Galesso, Karim Farid, Jonannes Dienert, Rajat Sahay, Thomas Brox
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.RO
arXiv:2607.15898v1 Announce Type: cross Abstract: Current world models operate at a single level of abstraction, with most prioritizing perceptual fidelity while lacking the spatial reasoning and semantic understanding required for real-world downstream tasks. We present a hierarchical driving world...
83. On the Failure of Boundary-Seeking Distillation in Bottlenecked Generative Architectures ​
Author: Mohamed Amine Kina
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.15919v1 Announce Type: cross Abstract: Data-free knowledge distillation transfers the knowledge encoded in a teacher model to a student model without access to the original training data. Prior work such as Contrastive Abductive Knowledge Extraction (CAKE) achieves this for classifiers by...
84. When Not to Automate: A Formal Protocol for Human Preservation in AI-Optimized Organizations ​
Author: Jose Manuel de la Chica Rodriguez, Jairo Rodriguez Arias, Spyridon Chouliaras
Published: 7/20/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CY, cs.HC
arXiv:2607.15944v1 Announce Type: cross Abstract: Standard automation ROI misses four categories of systemic risk -- tacit knowledge erosion, resilience reduction, regulatory exposure, and socio-institutional capital degradation -- that affect long-term organizational performance. PHP-AIO (Protocol ...
85. Sociocultural Influences on Opinion Formation: Word of Mouth Dynamics, Mass Media and Behavioural Development ​
Author: Elpida Tzafestas
Published: 7/20/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI
arXiv:2607.15968v1 Announce Type: cross Abstract: We study a society of agents belonging to a number of occupational or cultural groups that form opinions about others' situation in the same or different group. Opinions develop either by observation within own group or by directly interacting with m...
86. Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage Grids ​
Author: Josef Hoppe, Sarra Bouchkati, Farah Nasr, Jonathan Krapp, Alexander Och, Maximilian Wirth, Jan Schiefelbein-Lach, Oliver Pohl, Andreas Ulbig, Michael T. Schaub
Published: 7/20/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY
arXiv:2607.16004v1 Announce Type: cross Abstract: Increases in photovoltaic generation, charging of electric vehicles and heat-pump demand challenge operating limits in low-voltage distribution grids. This requires curative curtailment methods that can operate under sparse observability, noisy measu...
87. DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction ​
Author: Jehun Kang, Jungha Wang, Youngjun Hwang, David Hyunchul Shim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO
arXiv:2607.16012v1 Announce Type: cross Abstract: Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation. While Vision Foundation Models (VFMs) are increasingly adopted as robust feature...
88. Candidate Attended Dialogue State Tracking Using BERT ​
Author: Junyuan Zheng, Onkar Salvi, John Chan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.16021v1 Announce Type: cross Abstract: Dialogue state tracking (DST) is one of the core components in task-oriented dialogue systems. At each turn in a conversation, DST estimates the user belief or dialogue state, which is used as input for downstream modules to predict system actions an...
89. Rethinking Quantum Continual Learning with Quantum Fisher Information ​
Author: Yu-Chao Hsu, Yu-Cheng Lin, Tai-Yue Li, Nan-Yow Chen, En-Jui Kuo
Published: 7/20/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG
arXiv:2607.16030v1 Announce Type: cross Abstract: Quantum continual learning aims to train quantum models on sequential tasks without losing previously learned knowledge. However, variational quantum classifiers (VQCs) are prone to catastrophic forgetting under nonstationary task distributions. We p...
90. Revisiting data-driven dynamic security assessment with a tabular foundation model ​
Author: Olayiwola Arowolo, Maosheng Yang, Jochen Cremer
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.16031v1 Announce Type: cross Abstract: Data-driven pre-fault dynamic security assessment (DSA) rapidly evaluates the dynamic risk of credible contingencies on a power system using machine learning. Existing approaches face two limitations. First, they require a large labelled database for...
91. Loop the Loopies! ​
Author: Zitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang, Ran Tao, Joey Zhou, Bryan Dai
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.16051v1 Announce Type: cross Abstract: We present Loopie, the most powerful looped Transformer to date. The Loopie series consists of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6Bparameter model with 0.6B active parameters. Looped Transforme...
92. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning ​
Author: Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani, Mitch Weiss
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.16057v1 Announce Type: cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. W...
93. When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis ​
Author: S. Aaron McClendon
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.16062v1 Announce Type: cross Abstract: Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is essentially never tested against the baseline it claims to replace: methods merge independently released agents pr...
94. Spatial Normalization for Cross-Domain Retinal Layer Segmentation in Optical Coherence Tomography ​
Author: Iker Moran-Cavero, Monica Hernandez, Elvira Mayordomo, Naiara Artiaga, Beatriz Pardi~nas, Beatriz Cordon, Elena Garcia-Martin
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.16065v1 Announce Type: cross Abstract: Retinal layer segmentation in Optical Coherence Tomography (OCT) is a fundamental step for extracting quantitative biomarkers of retinal structure. Indeed, there is a growing interest in the analysis of OCTs in the context of neurodegenerative diseas...
95. LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization ​
Author: Mazene Ameur, Abdelkader Mekrache, Bouziane Brik, Adlen Ksentini
Published: 7/20/2026, 4:00:00 AM
Categories: cs.NI, cs.AI
arXiv:2607.16066v1 Announce Type: cross Abstract: Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation toward autonomous, goal-driven control of Next-Generation Networks (NGNs). Existing surveys treat the two domains in isolation, leaving p...
96. JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models ​
Author: Haoran Sun, Wentao Zhang, Junyang Hua, Hedan Yang, Yongjian Guo, Yifei Zhang, Xiaolong Xiang, Mingxi Luo, Jing Long, Chen Zhao, Chen Zhou, Wanting Xu, Qiming Yang, Hui Zhang, Song Wang, Xiaodong Bai, Shuai Di, Xu Chu, Xiaotie Deng, Yicheng Gong, Junwu Xiong
Published: 7/20/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.SE
arXiv:2607.16074v1 Announce Type: cross Abstract: The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, ty...
97. HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection ​
Author: Bhavana Verma, Priyanka Meel, Dinesh Kumar Vishwakarma
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2607.16076v1 Announce Type: cross Abstract: Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textual and visual information rather than from either modality alone. Existing multimodal approaches primarily rely...
98. DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning ​
Author: Hanyang Chen, Anirudh Satheesh, Longchao Da, Hua Wei
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.16090v1 Announce Type: cross Abstract: Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains. In this paper, we consider the setting of online dynamics adaptation, where policies are train...
99. Understanding Reasoning from Pretraining to Post-Training ​
Author: Jingyan Shen, Ang Li, Salman Rahman, Yifan Sun, Micah Goldblum, Matus Telgarsky, Pavel Izmailov
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.16097v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain o...
100. A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance ​
Author: Andrea Ferrario
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.SY, eess.SY
arXiv:2607.16130v1 Announce Type: cross Abstract: AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time, whether observed changes are tolerable, and how such judgments should be documented in a transparent and contestable way. Yet existing ...
101. ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning ​
Author: Binglin Zhou, Peng Shi, Ryo Kamoi, Nan Zhang, Rui Zhang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.16131v1 Announce Type: cross Abstract: Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and textual context. However, existing methods often fail because they strugg...
102. When Do Multi-Agent Systems Help? An Information Bottleneck Perspective ​
Author: Wendi Yu, Lianhao Zhou, Xiangjue Dong, Sai Sudarshan Barath, Declan Staunton, Byung-Jun Yoon, Xiaoning Qian, James Caverlee, Shuiwang Ji
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.16133v1 Announce Type: cross Abstract: LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with performance varying inconsistently across settings. Here, we provide an infor...
103. An Exam for Active Observers ​
Author: Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG
arXiv:2607.16165v1 Announce Type: cross Abstract: Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. W...
104. When Does Muon Help Agentic Reinforcement Learning? ​
Author: Kai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.16169v1 Announce Type: cross Abstract: Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld ...
105. Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities ​
Author: Md Erfan, Ahmed Ryan, Md Kamal Hossain Chowdhury, Md Rayhanur Rahman
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2607.16175v1 Announce Type: cross Abstract: Connected and Autonomous Vehicles (CAVs) rely on interconnected software and hardware components, including sensors, Electronic Control Units, in-vehicle infotainment systems, and telematics units, where vulnerabilities can compromise assets, users, ...
106. RAD: Retrieval High-quality Demonstrations to Enhance Decision-making ​
Author: Lu Guo, Yixiang Shan, Zhengbang Zhu, Qifan Liang, Lichang Song, Ting Long, Weinan Zhang, Yi Chang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2507.15356v3 Announce Type: replace Abstract: Offline reinforcement learning (RL) learns policies from fixed datasets, thereby avoiding costly or unsafe environment interactions. However, its reliance on finite static datasets inherently restricts the ability to generalize beyond the training ...
107. A Neuro-Symbolic Approach for Probabilistic Reasoning on Graph Data ​
Author: Raffaele Pojer, Andrea Passerini, Kim G. Larsen, Manfred Jaeger
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2507.21873v2 Announce Type: replace Abstract: Graph neural networks (GNNs) excel at predictive tasks on graph-structured data but often lack the ability to incorporate symbolic domain knowledge and perform general reasoning. Relational Bayesian Networks (RBNs), in contrast, enable fully genera...
108. Human-Aligned Procedural Level Generation Reinforcement Learning via Text-Level-Sketch Shared Representation ​
Author: In-Chang Baek, Seoyoung Lee, Sung-Hyun Kim, Geumhwan Hwang, KyungJoong Kim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2508.09860v2 Announce Type: replace Abstract: Human-aligned AI is a critical component of co-creativity, as it enables models to accurately interpret human intent and generate controllable outputs that align with design goals in collaborative content creation. This direction is especially rele...
109. RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs ​
Author: Ruike Hu, Shulei Wu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2512.00319v3 Announce Type: replace Abstract: The Structure Gap between probabilistic LLM generation and deterministic schema requirements hinders automated workflows. We propose RL-Struct, a lightweight framework using Gradient Regularized Policy Optimization (GRPO) with a hierarchical reward...
110. Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy ​
Author: Bianca Raimondi, Maurizio Gabbrielli
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2602.17229v2 Announce Type: replace Abstract: The black-box nature of Large Language Models necessitates novel evaluation frameworks that transcend surface-level performance metrics. This study investigates the internal neural representations of cognitive complexity using Bloom's Taxonomy as a...
111. The AI Fiction Paradox ​
Author: Katherine Elkins
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY
arXiv:2603.13545v2 Announce Type: replace Abstract: AI development has a fiction dependency problem. Developers have treated large corpora of modern books, including fiction, as valuable enough to accept substantial cost and legal risk, yet current models still struggle to generate compelling long-f...
112. SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents ​
Author: Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Nathaniel Gorski, Jianxin Sun, Guoxi Liu, Helgi I. Ingolfsson, David Lenz, Hanqi Guo, Hongfeng Yu, Teja Leburu, Michael Molash, Bei Wang, Tom Peterka, Chaoli Wang, Shusen Liu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.GR, cs.HC
arXiv:2603.29139v3 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visualization (SciVis) tasks. Despite rapid progress, the community lacks a principled and reproducible ben...
113. FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment ​
Author: Sophie Chiang, Tom Brennan, Fethiye Irmak Dogan, Jiaee Cheong, Hatice Gunes
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2604.23786v2 Announce Type: replace Abstract: In recent years, the integration of multimodal machine learning in wellbeing assessment has offered transformative potential for monitoring mental health. However, with the rapid advancement of Vision-Language Models (VLMs), their deployment in cli...
114. Towards a General Intelligence and Interface for Wearable Health Data ​
Author: Girish Narayanswamy, Maxwell A. Xu, A. Ali Heydari, Samy Abdel-Ghaffar, Marius Guerard, Kara Vaillancourt, Zhihan Zhang, Jake Garrison, Levi Albuquerque, Dimitris Spathis, Hong Yu, Hamid Palangi, Xuhai "Orson" Xu, David G. T. Barrett, Joseph Breda, Jed McGiffin, Yubin Kim, Yuwei Zhang, Naghmeh Rezaei, Samuel Solomon, Karan Ahuja, Tim Althoff, Jake Sunshine, Ming-Zher Poh, Benjamin Yetton, Ari Winbush, Nicholas B. Allen, James M. Rehg, Isaac Galatzer-Levy, Yun Liu, John Hernandez, Anupam Pathak, Conor Heneghan, Yuzhe Yang, Ahmed A. Metwally, Pushmeet Kohli, Mark Malhotra, Shwetak Patel, Xin Liu, Daniel McDuff
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.22759v3 Announce Type: replace Abstract: While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into personalized health insights is challenging. Specifically, converting low-level sensor data into representat...
115. Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields ​
Author: Liya Zhu, Jingzhe Ding, Jian Zhang, Jianbo Xue, Shihao Liang, Ge Zhang, Yi Zhu, Duju Zeng, Xiang Gao, Qingshui Gu, Mailun Gao, Huimin Che, Yan Zhao, Peiheng Zhou, Haojun Wang, Chaobo Xian, Lili Le, Chi Wu, Yiwei Liu, Shengda Long, Jiale Yang, Fangzhi Xu, Sijin Wu, Haodong Duan, Chao He, Zhaojian Li, Minchao Wang, Huan Zhou, Jiani Hou, Chuqian Yu, Weiran Shi, Hongwan Gao, Jiamin Chen, Guanhong Chen, Tingqin Luo, Kaiyuan Zhang, Zhixin Yao, Qing Hua, Yuhao Jiang, Jin Chen, Pu Chen, Zhenyu Hu, Xingyu Li, Zhengxuan Jiang, Meng Cao, Tianfeng Long, Haozhe Wang, Mingzhang Wang, Yichen Zhang, Yiming Dai, Chenchen Zhang, Jiaying Wang, Xinying Liu, Xingzu Liu, Lingling Zhang, Xinjie Chen, Yujia Qin, Wangchunshu Zhou, Zhiyong Wu, Yang Liu, Jiaheng Liu, Lei Zhang, Shen Yan, Wenhao Huang, Zaiyuan Wang, Xiaolong Chang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.11042v4 Announce Type: replace Abstract: Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-va...
116. Agents-K1: Towards Agent-native Knowledge Orchestration ​
Author: Zongsheng Cao, Bihao Zhan, Jinxin Shi, Jiong Wang, Fangchen Yu, Zhijie Zhong, Yingnan Han, Zijie Guo, Tianshuo Peng, Zhuo Liu, Yi Xie, Xiang Zhuang, Shengji Tang, Yue Fan, Runmin Ma, Shiyang Feng, Xiangchao Yan, Anran Liu, Peng Ye, Wenlong Zhang, Xiaosong Wang, Shufei Zhang, Chunfeng Song, Fenghua Ling, Jie Zhou, Liang He, Bo Zhang, Lei Bai
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.13669v3 Announce Type: replace Abstract: Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration. Existing works often reduce papers to abstracts, surface mentions, and flat \texttt{cites} edges, omitting key ent...
117. GA-VINO: A Geometry-Aware Variational Physics-informed Neural Operator for Mindlin-Reissner Plates ​
Author: Siqi Wang, Daobo Sun, Yizheng Wang, Yilong Zhang, Yabin Jin, Xiaoying Zhuang, Timon Rabczuk
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.16624v2 Announce Type: replace Abstract: Plate and shell structures are widely used in engineering fields. Rapid response prediction for such structures under complex geometries, heterogeneous materials, and varying loads is important for engineering design, but conventional numerical met...
118. UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation ​
Author: Songjun Tu, Chengdong Xu, Qichao Zhang, Yiwen Ma, Yaocheng Zhang, Linjing Li, Dong Li, Xiangyuan Lan, Dongbin Zhao
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2606.29502v2 Announce Type: replace Abstract: Skill memories can improve agentic reinforcement learning by reusing past experience as textual guidance, but retrieved skills are not oracular: they may help in one state while misleading the same policy in another. This makes the common privilege...
119. ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning ​
Author: Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara, Scott Sanner, Pascal Poupart, Jongmin Lee, Kee-Eung Kim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.30072v2 Announce Type: replace Abstract: Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centralized Training with Decentralized Execution (CTDE) paradigm, policy gradients have remained difficult to compute...
120. MirrorCode: AI can rebuild entire programs from behavior alone ​
Author: Tom Adamczewski, David Owen, David Rein, Florian Brand, Giles Edkins, Allen Hart, Daniel O'Connell
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.30182v2 Announce Type: replace Abstract: AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C compiler. However, existing coding benchmarks tend to focus on shorter tasks, and one-off demonstrations are...
121. Internal Pluralism and the Limits of Pairwise Comparisons ​
Author: Bailey Flanigan, Michelle Si
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2607.02672v2 Announce Type: replace Abstract: Local pairwise comparisons are a standard tool for learning how people want decision rules to work, e.g., in participatory design or alignment. However, their use builds in two strong assumptions: that local comparisons are sufficient evidence abou...
122. Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces ​
Author: Andrew Zhang, Chengzhan Li
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.04419v4 Announce Type: replace Abstract: Pooling, substituting, or reusing evaluator-derived step rewards assumes that their direction survives a change of evaluation channel. The same frozen transition can violate that assumption. Process rewards vary agent states, while evaluator audits...
123. OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice ​
Author: Qian Jiang, Zhecheng Shi, Jingpu Yang, Zirui Song, Miao Fang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.08423v2 Announce Type: replace Abstract: The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare and dietary management. However, in the domain of food systems, autonomous agents face a unique and persiste...
124. Evidence-Aware MapReduce for Forkable Compute ​
Author: Yossi Eliaz
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, math.PR, math.ST, stat.TH
arXiv:2607.09689v3 Announce Type: replace Abstract: Snapshot-backed sandboxes make branching cheap while leaving evidence dependence unchanged. Branches can reuse a model, prompt, repository, tests, observations, or execution ancestor, so counting outputs can amplify one repeated error into high-con...
125. Length Penalties Make Chain-of-Thought Less Monitorable ​
Author: Bryce Little
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2607.09786v2 Announce Type: replace Abstract: Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In our experiments, training with length penalties does not stop misleading hints from steering models, even th...
126. ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory ​
Author: Jiayi Tian, Shiao Liu, Yuting Xu, Jia Lu, Zihao Guan, Honglin Han, Di Yang, Minqi Gu, Yifei Qian, Tianlin Zhang, Yanqing Zhu, Zeqian Ye, Menglin Yang, Fei Wang, Xu Hu, Xiuxian Li, Wei Zhang, Shihui Su, Yiyan Ji, Jingbo Wang, Ziteng Feng, Jiaheng Liu, Zhaoxiang Zhang, Xiaolong Wu, Zixiao Tang, Zhining Gu, Yang Cai, Linbo Zheng, Jingjing Ma, Mingyang Yin, Zedong Chu, Wenbin Tang, Mu Xu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2607.10350v3 Announce Type: replace Abstract: Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot...
127. First-Order Modal Logic in HOL: Deep and Shallow Embeddings with Automated Faithfulness (Extended Preprint) ​
Author: Christoph Benzm"uller, Daniel Kirchner
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, math.LO
arXiv:2607.10880v2 Announce Type: replace Abstract: We extend, in Isabelle/HOL, the deep-and-shallow embedding methodology of our prior work from propositional to first-order modal logic (FML) with constant-domain Kripke semantics. Three embeddings of FML into classical higher-order logic (HOL) are ...
128. The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning ​
Author: Wencheng Ye, Yi Bin, Yujuan Ding, Hongye Fang, Zheng Wang, Xing Xu, Jingkuan Song, Yun Zhang, Sirui Da, Heng Tao Shen
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.11436v2 Announce Type: replace Abstract: Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening evidence-grounded reasoning. To understand this fragility, we examine the ...
129. MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku ​
Author: Pedro Orvalho, Guillem Aleny`a, Felip Many`a
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LO
arXiv:2607.12711v2 Announce Type: replace Abstract: Vision--Language Models (VLMs) have recently demonstrated promising performance on structured visual reasoning tasks, including grid-based puzzles. However, despite strong perceptual capabilities, these models lack explicit mechanisms for enforcing...
130. Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents ​
Author: Shiyu Ying, Xuejie Cao, Yingfan Ma, Yuanhao Dong, Wenyu Chen, Bowen Song, Lin Zhu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2607.14573v2 Announce Type: replace Abstract: Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows, verify payment outcomes, and preserve consistency between transaction and business states. We intro...
131. BrainPilot: Automating Brain Discovery with Agentic Research ​
Author: Haoxuan Li, Tianci Gao, Jianhe Li, Yang Fan, Runze Shi, Weiran Wang, Tianxiang Zhao, Zezhao Wu, Xiaoyang Jiang, Qihui Zhang, Jia Li, Xiao Xiao, Kai Du, Xiaoxuan Jia, Chao Xie, Lu Mi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15079v2 Announce Type: replace Abstract: Understanding the brain increasingly depends on integrating evidence across scales, modalities, and disciplines. Addressing a single research question therefore requires a coordinated sequence of operations, from surveying prior work to executing a...
132. Long-Context Fine-Tuning with Limited VRAM ​
Author: Vladimir Fedosov, Aleksandr Sazhin, Artemiy Grinenko, Frank Woernle
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15105v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning reduces model and optimizer memory, but dense attention still makes long training sequences expensive. We combine Hierarchical Global Attention (HGA) with segment-wise backpropagation and tiered KV storage. Only the ...
133. Can We Trust Item Response Theory for AI Evaluation? ​
Author: Han Jiang, Sunbeom Kwon, Jinwen Luo, Ziang Xiao, Susu Zhang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.15190v2 Announce Type: replace Abstract: AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often ...
134. Perception-Aligned AI Outputs: End-to-End Visual Prediction for Uncertainty Communication in Clinical Decision-Making ​
Author: Mohammad Eslami, Solale Tabarestani, Saber Kazeminasab, Ehsan Adeli, Glyn Elwyn, Tobias Elze, Mengyu Wang, Nazlee Zebardast, Lucia Sobrin, Nassir Navab, Daniel Shu Wei Ting, Malek Adjouadi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2205.04599v2 Announce Type: replace-cross Abstract: Explainable Artificial Intelligence (XAI) is essential for trustworthy AI in healthcare, yet many existing methods rely on technical explanations that are difficult for clinicians and patients to interpret. We introduce Visualized Learning fo...
135. Why do CNNs excel at feature extraction? A mathematical explanation ​
Author: Vinoth Nandakumar, Arush Tagade, Tongliang Liu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2307.00919v2 Announce Type: replace-cross Abstract: Over the past decade deep learning has revolutionized the field of computer vision, with convolutional neural network models proving to be very effective for image classification benchmarks. However, a fundamental theoretical questions remain...
136. Decoupled Alignment for Robust Plug-and-Play Adaptation ​
Author: Haozheng Luo, Jiahao Yu, Wenxin Zhang, Jialong Li, Chenghao Qiu, Yimin Wang, Eric Hanchen Jiang, Jerry Yao-Chieh Hu, Yan Chen, Binghui Wang, Xinyu Xing, Han Liu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CR
arXiv:2406.01514v5 Announce Type: replace-cross Abstract: We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback. Our main idea is to provide a robust plug-and-play app...
137. Derivation of effective gradient flow equations and dynamical truncation of training data in Deep Learning ​
Author: Thomas Chen
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.AP, math.OC, stat.ML
arXiv:2501.07400v3 Announce Type: replace-cross Abstract: We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient descent for the Euclidean loss in the input layer, and under the assumption that the weights are, in a ...
138. CTC: The Composite Task Challenge for Cooperative Multi-Agent Reinforcement Learning ​
Author: Yurui Li, Yuxuan Chen, Xiaoli Yang, Shijian Li, Gang Pan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA
arXiv:2502.00345v2 Announce Type: replace-cross Abstract: The critical role of division of labor (DOL) in enhancing cooperation is well-recognized in real-world applications. Consequently, many cooperative multi-agent reinforcement learning (MARL) methods have incorporated DOL mechanisms to improve ...
139. MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation ​
Author: Haonan Yu, Junhao Liu, Xin Zhang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2502.11068v3 Announce Type: replace-cross Abstract: Anchors is a popular local model-agnostic explanation technique whose applicability is limited by its computational inefficiency. To address this limitation, we propose a memorization-based framework that accelerates Anchors while preserving ...
140. AuditVotes: Elevating Provable Defense for GNNs with Efficient Augmentation and Conditional Smoothing ​
Author: Yuni Lai, Yulin Zhu, Yixuan Sun, Yulun Wu, Bin Xiao, Gaolei Li, Jianhua Li, Qi Xie, Kai Zhou
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2503.22998v2 Announce Type: replace-cross Abstract: Despite advancements in Graph Neural Networks (GNNs), adaptive attacks continue to challenge their robustness. Certified robustness via randomized smoothing offers provable guarantees but suffers from a severe accuracy-robustness trade-off, l...
141. A Scaffolded GenAI Lab in Early Undergraduate CS: A Mixed-Methods, Multi-Course Evaluation ​
Author: Ethan Dickey, Andres Bejarano, Rhianna Kuperus, B'arbara Fagundes
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.ET
arXiv:2505.00100v2 Announce Type: replace-cross Abstract: Background and Context. Generative AI (GenAI) tools are increasingly used in programming courses, but we have limited evidence about how brief instruction can foster responsible, learning-oriented use. Objectives. We evaluate "AI-Lab", a scaf...
142. SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training ​
Author: Jiaheng Hu, Peter Stone, Roberto Mart'in-Mart'in
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2506.04147v5 Announce Type: replace-cross Abstract: Building capable household and industrial robots requires mastering the control of versatile, high-degree-of-freedom (DoF) systems such as mobile manipulators. While reinforcement learning (RL) holds promise for autonomously acquiring robot c...
143. A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model ​
Author: Andris Ambainis, Joao F. Doriguello, Debbie Lim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC, quant-ph, stat.ML
arXiv:2507.22854v3 Announce Type: replace-cross Abstract: We propose novel classical and quantum online algorithms for learning finite- and infinite-horizon Markov Decision Processes (MDPs). Our algorithms are based on a hybrid online-offline reinforcement learning model wherein the agent can, from ...
144. Acoustic Imaging for UAV Detection: Dense Beamformed Energy Maps and U-Net SELD ​
Author: Belman Jahir Rodriguez, Sergio F. Chevtchenko, Marcelo Herrera Martinez, Yeshwanth Bethi, Saeed Afshar
Published: 7/20/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD, eess.SP
arXiv:2508.00307v4 Announce Type: replace-cross Abstract: We introduce a U-net model for 360{\deg} acoustic source localization formulated as a spherical semantic segmentation task. Rather than regressing discrete direction-of-arrival (DoA) angles, our model segments beamformed audio maps (azimuth &...
145. Unsupervised Deep Learning for Inverse Problems in Computed Tomography ​
Author: Laura Hellwege, Johann Christopher Engster, Moritz Schaar, Thorsten M. Buzug, Maik Stille
Published: 7/20/2026, 4:00:00 AM
Categories: physics.med-ph, cs.AI
arXiv:2508.05321v4 Announce Type: replace-cross Abstract: Assume you encounter an inverse problem that shall be solved for a large number of data, but no ground-truth data is available. To emulate this, in this study we assume it is unknown how to solve the imaging problem of Computed Tomography. We...
146. Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs ​
Author: Wenpeng Xing, Bohan Yang, Mohan Li, Chunqiang Hu, Haitao Xu, Ningyu Zhang, Bo Lin, Meng Han
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CR
arXiv:2508.10029v3 Announce Type: replace-cross Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations. We introduce Latent Fusion Jailbreak (LFJ), which works by pairing a harmful query with a structurally s...
147. Poison to Detect: Detection of Targeted Overfitting in Federated Learning ​
Author: Soumia Zohra El Mestari, Maciej Krzysztof Zuziak, Gabriele Lenzini
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2509.11974v2 Announce Type: replace-cross Abstract: Federated Learning (FL) enables collaborative model training among clients without centralising data, making it a widely adopted privacy-enhancing technology (PET). Despite its privacy benefits, FL remains vulnerable to orchestrator-driven pr...
148. A Systematic Study of Large Language Models for Task and Motion Planning With PDDLStream ​
Author: Jorge Mendez-Mendez
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2510.00182v2 Announce Type: replace-cross Abstract: While we know that large language models (LLMs) can solve some planning problems, we do not understand the extent of these capabilities for robotics. One promising direction is to integrate the semantic knowledge of LLMs with the formal reaso...
149. Are Heterogeneous Graph Neural Networks Truly Effective for Node Classification? A Causal Perspective ​
Author: Xiao Yang, Xuejiao Zhao, Zhiqi Shen
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2510.05750v2 Announce Type: replace-cross Abstract: Graph neural networks (GNNs) have achieved remarkable success in node classification. Building on this progress, heterogeneous graph neural networks (HGNNs) integrate relation types and node and edge semantics to leverage heterogeneous inform...
150. Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability ​
Author: Bianca Raimondi, Daniela Dalbagno, Maurizio Gabbrielli
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.12229v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have been shown to internalize human-like biases during finetuning, yet the mechanisms by which these biases manifest remain unclear. In this work, we investigated whether the well-known Knobe effect, a moral bias...
151. Human-Inspired Neuro-Symbolic World Modeling and Logic Reasoning for Interpretable Safe UAV Landing Site Assessment ​
Author: Weixian Qian, Tianyi Yang, Sebastian Schroder, Yao Deng, Jiaohong Yao, Xiao Cheng, Richard Han, Xi Zheng
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2510.22204v3 Announce Type: replace-cross Abstract: Reliable assessment of safe landing sites in unstructured environments is essential for deploying Unmanned Aerial Vehicles (UAVs) in real-world applications such as delivery, inspection, and surveillance. Existing learning-based approaches of...
152. DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone ​
Author: Vaibhav Singh, Oleksiy Ostapenko, Pierre-Andr'e No"el, Eugene Belilovsky, Torsten Scholak
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2511.15927v4 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency due to quadratic attention or KV-cache overhead. We introduce ...
153. 3D Motion Perception of Binocular Vision Target with PID-CNN ​
Author: Jiazhao Shi, Pan Pan, Haotian Shi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2511.20332v3 Announce Type: replace-cross Abstract: This article trained a network for perceiving three-dimensional motion information of binocular vision target, which can provide real-time three-dimensional coordinate, velocity, and acceleration, and has a basic spatiotemporal perception cap...
154. Hybrid coupling with operator inference and the overlapping Schwarz alternating method ​
Author: Irina Tezaur, Eric Parish, Anthony Gruber, Ian Moore, Christopher Wentland, Alejandro Mota
Published: 7/20/2026, 4:00:00 AM
Categories: math.NA, cs.AI, cs.NA, math-ph, math.MP
arXiv:2511.20687v4 Announce Type: replace-cross Abstract: This paper presents a novel hybrid approach for coupling subdomain-local non-intrusive Operator Inference (OpInf) reduced order models (ROMs) with each other and with subdomain-local high-fidelity full order models (FOMs) with using the overl...
155. Energy-Efficient Federated Learning via Adaptive Encoder Freezing for MRI-to-CT Conversion: A Green AI-Guided Research ​
Author: Ciro Benito Raggio, Lucia Migliorelli, Nils Skupien, Mathias Krohmer Zabaleta, Oliver Blanck, Francesco Cicone, Giuseppe Lucio Cascini, Paolo Zaffino, Maria Francesca Spadea
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, cs.DC, physics.med-ph
arXiv:2512.03054v3 Announce Type: replace-cross Abstract: Federated Learning (FL) holds the potential to advance equality in health by enabling diverse institutions to collaboratively train deep learning (DL) models, even with limited data. However, the significant resource requirements of FL often ...
156. Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length ​
Author: Zhiyu Xu, Jia Liu, Yixin Wang, Yuqi Gu
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, stat.AP, stat.ML
arXiv:2512.07019v3 Announce Type: replace-cross Abstract: The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream applications and actionable future improvements. The Item Response Theory (IRT) model with Computerized Adaptive T...
157. PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning ​
Author: Zhiyan Hou, Haiyun Guo, Haokai Ma, Yandu Sun, Yonghui Yang, Jinqiao Wang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.13020v2 Announce Type: replace-cross Abstract: Continual instruction tuning (CIT) requires multimodal large language models (MLLMs) to adapt to a stream of tasks without forgetting prior capabilities. A common strategy is to isolate updates by routing inputs to different LoRA experts. How...
158. Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models ​
Author: Charles Westphal, Keivan Navaie, Fernando E. Rosas
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2601.22818v2 Announce Type: replace-cross Abstract: Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels. Prior work demonstrated this threat but relied on trivially recoverable encodings. We formalize payload recoverability via classifier accuracy and sh...
159. Inelastic Constitutive Kolmogorov-Arnold Networks: A generalized framework for automated discovery of interpretable inelastic material models ​
Author: Chenyi Ji, Kian P. Abdolazizi, Hagen Holthusen, Christian J. Cyron, Kevin Linka
Published: 7/20/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, physics.comp-ph
arXiv:2602.17750v3 Announce Type: replace-cross Abstract: A key problem of solid mechanics is the identification of the constitutive law of a material, that is, the relation between strain history and stress. Machine learning has lead to considerable advances in this field lately. Here we introduce ...
160. Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking ​
Author: Zhicheng Fang, Jingjie Zheng, Chenxu Fu, Wei Xu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG
arXiv:2602.24009v4 Announce Type: replace-cross Abstract: Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datasets, harnesses, and judging protocols. We introduce JAILBREAK F...
161. KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models ​
Author: Songming Zhang, Xue Zhang, Tong Zhang, Bojie Hu, Yufeng Chen, Jinan Xu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2603.01875v3 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is an essential technique to compress large language models (LLMs) into smaller ones. However, despite the distinct roles of the student model and the teacher model in KD, most existing frameworks still use a homog...
162. Interaction-Aware Whole-Body Control for Compliant Object Transport ​
Author: Hao Zhang, Yves Tseng, Ding Zhao, H. Eric Tseng
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2603.03751v2 Announce Type: replace-cross Abstract: Cooperative object transport in unstructured environments remains challenging for assistive humanoids because strong, time-varying interaction forces can make tracking-centric whole-body control unreliable, especially in close-contact support...
163. CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation ​
Author: Mahmoud Ibrahim, Bart Elen, Chang Sun, Gokhan Ertaylan, Michel Dumontier
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.16551v3 Announce Type: replace-cross Abstract: Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that generators produce equally high-quality images across demographic groups. Models trained on imbalanced...
164. When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models ​
Author: Juan Gabriel Kostelec, Qinghai Guo
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.26556v2 Announce Type: replace-cross Abstract: Converting a pretrained Transformer into a more efficient hybrid model through distillation offers a promising approach to reducing inference costs. However, achieving high-quality generation in distilled models requires careful joint design ...
165. Ruling Out to Rule In: Contrastive Hypothesis Retrieval for Medical Question Answering ​
Author: Byeolhee Kim, Min-Kyung Kim, Young-Hak Kim, Tae-Joon Jeon
Published: 7/20/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2604.04593v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds large language models in external medical knowledge, yet standard retrievers frequently surface hard negatives that are semantically close to the query but describe clinically distinct conditions. ...
166. Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios ​
Author: Ruida Hu, Xinchen Wang, Chao Peng, Cuiyun Gao, David Lo
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2604.06742v2 Announce Type: replace-cross Abstract: The evolution of Large Language Models (LLMs) has catalyzed a paradigm shift towards intent-driven software development, where autonomous agents are expected to design and deliver complete, runnable software systems from scratch. However, exi...
167. LVSum: A Benchmark for Timestamp-Aware Long Video Summarization ​
Author: Alkesh Patel, Melis Ozyildirim, Ying-Chang Cheng, Ganesh Nagarajan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2604.10024v2 Announce Type: replace-cross Abstract: Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally gro...
168. Robust Explanations for User Trust in Enterprise NLP Systems ​
Author: Guilin Zhang, Kai Zhao, Jeffrey Friedman, Xu Chu, Amine Anoun, Jerry Ting
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2604.12069v4 Announce Type: replace-cross Abstract: Robust explanations are increasingly required for user trust in enterprise NLP, yet pre-deployment validation is difficult in the common case of black-box deployment (API-only access) where representation-based explainers are infeasible and e...
169. Soft $Q(\lambda)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces ​
Author: Pranav Mahajan, Ben Seymour
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.13780v2 Announce Type: replace-cross Abstract: Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a penalty on the divergence from a reference policy. Despite its success, the multi-step extens...
170. What Is the Minimum Architecture for Prolepsis? Early Irrevocable Commitment Across Tasks in Small Transformers ​
Author: 'Eric Jacopin
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2604.15010v2 Announce Type: replace-cross Abstract: When do transformers commit to a decision, and what prevents them from correcting it? We introduce prolepsis: a transformer commits early, task-specific attention heads sustain the commitment, and no layer corrects it. Replicating Lindsey et ...
171. Brain-CLIPLM: Semantic Compression for EEG-to-Text Decoding ​
Author: Xiaoli Yang, Huiyuan Tian, Yurui Li, Jianyu Zhang, Shijian Li, Gang Pan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV
arXiv:2604.16370v3 Announce Type: replace-cross Abstract: Decoding natural language from non-invasive electroencephalography (EEG) remains constrained by low signal-to-noise ratio and limited information bandwidth. This raises a central question: can sentence-level language be reliably recovered fro...
172. RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization ​
Author: Qiming Bao, Juho Leinonen, Paul Denny, Michael J. Witbrock
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.04539v4 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO), the efficient alternative to PPO-based RLHF, falls short on knowledge-intensive generation: standard preference signals from human annotators or LLM judges exhibit a systematic verbosity bias that rewards...
173. Energy-based Transport for Amortized Bayesian Inference ​
Author: Ricardo Baptista, Hojjat Kaveh, Andrew M. Stuart
Published: 7/20/2026, 4:00:00 AM
Categories: math.NA, cs.AI, cs.NA
arXiv:2605.15407v3 Announce Type: replace-cross Abstract: We consider amortized Bayesian inference for nonlinear inverse problems using only samples from the joint distribution of parameters and observations, including problems with unknown functions in a Banach space. Classical methods such as Mark...
174. Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory ​
Author: Bole Ma, Jan Eitzinger, Harald Koestler, Gerhard Wellein
Published: 7/20/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG
arXiv:2605.20982v2 Announce Type: replace-cross Abstract: AlltoAll dispatch is the dominant bottleneck of MoE expert parallelism, and the interconnect community has responded with four families of mitigations: predictive sample placement, adaptive expert relayout, hierarchical collectives, and EP-aw...
175. The Terminal Representation in Reinforcement Learning ​
Author: Amir Esterhuysen, Anders Jonsson
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.31289v2 Announce Type: replace-cross Abstract: Representation learning is a powerful tool for spatio-temporal abstraction within reinforcement learning (RL). Two well established approaches are through the successor representation (SR) and the default representation (DR). The SR encodes s...
176. memorywire: A Vendor-Neutral Wire Format for Agent Memory Operations ​
Author: Thamilvendhan Munirathinam
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.DC
arXiv:2606.01138v3 Announce Type: replace-cross Abstract: Agent-memory frameworks -- mem0, Letta/MemGPT, Cognee, Zep/Graphiti, MemoryOS, MemTensor -- each ship their own SDK, storage layout, and operational vocabulary. There is no shared wire format: every integration is bespoke, every migration reb...
177. AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations ​
Author: Hiskias Dingeto, William Leeney
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.ET
arXiv:2606.02240v3 Announce Type: replace-cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither writ...
178. AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models ​
Author: Andreas Einwiller, Max Klabunde, Florian Lemmerich
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2606.16127v2 Announce Type: replace-cross Abstract: The worldwide rise of authoritarianism and the growing role of Large Language Models (LLMs) in users' everyday lives raise the question of whether specific models exhibit or promote authoritarian attitudes. We introduce AuAu, a comprehensive ...
179. GeoRouteNet: A Geometry-Aware Non-Autoregressive Neural Solver for the Euclidean Traveling Salesman Problem ​
Author: Xiang Li
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.22776v2 Announce Type: replace-cross Abstract: Non-autoregressive neural solvers amortize computation across traveling salesman problem (TSP) instances, but models trained on random Euclidean instances can degrade when the number or spatial distribution of nodes changes. We study whether ...
180. HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization ​
Author: Kuangshi Ai, Patrick Phuoc Do, Chaoli Wang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.GR
arXiv:2606.26614v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis). Still, prior systems have essentially prioritized autonomy over human analytical control, thereby limiting transparency and human ove...
181. RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources ​
Author: Yijia Fan, Zonglin Di, Zimo Wen, Yifan Yang, Mingxi Cheng, Qi Dai, Bei Liu, Kai Qiu, Yue Dong, Ji Li, Chong Luo
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2606.29538v4 Announce Type: replace-cross Abstract: Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or derived from agent traces, leaving tutorial vid...
182. Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph ​
Author: Duen Horng Chau, Donghao Ren, Fred Hohman, Dominik Moritz
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DS, cs.HC
arXiv:2607.08746v2 Announce Type: replace-cross Abstract: While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally. This graph encodes the dat...
183. Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations ​
Author: Nada Zine, Tristan Coignion, Vincenzo Stoico, Cl'ement Quinton, Ivano Malavolta, Romain Rouvoy, Patricia Lago
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.PF
arXiv:2607.09172v2 Announce Type: replace-cross Abstract: Large Language Models are reshaping how software is developed and maintained. They are typically deployed in production using inference engines such as vLLM, which can efficiently serve pre-trained, highly configurable models. While prior wor...
184. ABot-N1: Toward a General Visual Language Navigation Foundation Model ​
Author: Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Yang Cai, Jingjing Ma, Shihui Su, Zixiao Tang, Linbo Zheng, Zedong Chu, Xiaolong Wu, Wenbin Tang, Mu Xu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO
arXiv:2607.10383v3 Announce Type: replace-cross Abstract: Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map ...
185. Learning the Brain's Dynamics as a Port-Hamiltonian System: A GNN-Surrogate Metriplectic Twin for Non-Equilibrium Cortical Dynamics and Closed-Loop Neuromodulation ​
Author: Dibakar Sigdel
Published: 7/20/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI
arXiv:2607.10439v2 Announce Type: replace-cross Abstract: We model human motor cortex, recorded during rest and motor-imagery BCI conditions, as a port-Hamiltonian system: a conservative interconnection (skew-symmetric coupling between band-limited neural phasors) together with a dissipative port wh...
186. Scaling Point-in-Time Language Models ​
Author: Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.11889v2 Announce Type: replace-cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences. Point-...
187. From Reconstruction to Interpretation: Zero-Setup Multi-Phase Segmentation of X-ray Tomography Data ​
Author: Pradyumna Elavarthi, Arun J. Bhattacharjee, Harrison Lisabeth, Anca Ralescu, Petrus H. Zwart, Dilworth Parkinson, Elizabeth G. Clark
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.12175v3 Announce Type: replace-cross Abstract: X-ray tomography enables nondestructive characterization of material microstructures, while advances in micro-CT imaging have accelerated volumetric data acquisition and reconstruction. However, rapid interpretation remains limited by image s...
188. Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs ​
Author: Xiaoning Ren, Yinxing Xue, Lei Ma, Yuheng Huang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL
arXiv:2607.12273v2 Announce Type: replace-cross Abstract: As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, where even minor errors can lead to severe functional, security, or safety consequences. Reli...
189. Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference ​
Author: Zebin Yang, Qi Wang, Yunhe Wang, Xiurui Guo, Bo Yu, Shaoshan Liu, Jiafeng Xu, Hao Dong, Meng Li
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.12659v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-power onboard devices, such as the Jetson Orin, remains challenging due to their high computational compl...
190. Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models ​
Author: Chen Li, Jiexiong Liu, Yi Li
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2607.13093v2 Announce Type: replace-cross Abstract: On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers strong computing power but exposes user prompts and dialogue data, while standalone on-device inference i...
191. What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors ​
Author: Winston Zeng, Ali Emami, Jinho D. Choi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.13162v3 Announce Type: replace-cross Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this...
192. Faithful Autoformalization of Natural Language Assertions ​
Author: Hongyi Liu, Madhusudan Parthasarathy, Adithya Murali
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.13303v2 Announce Type: replace-cross Abstract: Formal contracts are essential for software testing and verification, yet writing them remains labor-intensive and error-prone. LLMs offer a promising path toward autoformalization: synthesizing executable assertions from natural-language spe...
193. LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition ​
Author: Donghwan Kim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.13347v2 Announce Type: replace-cross Abstract: LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a controlled tes...
194. MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model ​
Author: Charilaos Papaioannou, Ioannis Tsantilas, Dimitris Giannakakos, Vasilis Michalakopoulos, Sotiris Pelekis, Vangelis Marinakis, Arsam Aryandoust, Antonello Monti, Ricardo J. Bessa, Perdo P. Vergara, Jochen Cremer, Elissaios Sarmas
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.13763v2 Announce Type: replace-cross Abstract: Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the lowest in-distribution error degrade the most under topology shift. We term this topology overfitting:...
195. Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code ​
Author: Niels M"undler-Sasahara, Hristo Venev, Dawn Song, Martin Vechev, Jingxuan He
Published: 7/20/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.LG
arXiv:2607.13921v2 Announce Type: replace-cross Abstract: Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generation more difficult. Off-the-shelf compilers can provide useful feedback post-generation, but does not gui...
196. NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs ​
Author: Jiarong Zhao, Zhikai Lei, Zhiheng Xi, Rui Zheng, Hang Yan, Jie Zhou, Qin Chen, Liang He
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2607.14186v2 Announce Type: replace-cross Abstract: Synthesizing training data to scale agent capabilities in LLM post-training is bottlenecked by substrate-bound task synthesis: tasks are generated from fixed tools, repositories, or skill graphs, so expanding coverage requires manual substrat...
197. Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values ​
Author: Jan Betley, Johannes Treutlein, Jan Dubi'nski, Harry Mayne, Karol Ga{\l}\k{a}zka, Niels Warncke, Anna Sztyber-Betley, Owain Evans
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2607.14345v2 Announce Type: replace-cross Abstract: People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed...
198. Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction ​
Author: Ryuichi Sumida, Mao Saeki, Masaki Eguchi, Sadahiro Yoshikawa, Koji Inoue, Tatsuya Kawahara, Yoichi Matsuyama
Published: 7/20/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL
arXiv:2607.14593v2 Announce Type: replace-cross Abstract: As conversational AI systems are designed for repeated use, a central question is how a series of interactions becomes a relationship. We present a longitudinal multimodal study of a memory-augmented conversational agent (24 participants x 10...
199. Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents ​
Author: Dylan Van Mulders, Matthias Bogaert, Dirk Van den Poel
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA
arXiv:2607.15095v2 Announce Type: replace-cross Abstract: The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convictions. While Large Language Models (LLMs) open new avenues for computational political science, the neu...
200. T^2MLR: Transformer with Temporal Middle-Layer Recurrence ​
Author: Ziyang Cai, Xingyu Zhu, Yihe Dong, Yinghui He, Sanjeev Arora
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.15178v2 Announce Type: replace-cross Abstract: Transformer reasoning is limited by autoregressive decoding, which repeat edly compresses rich hidden computation through token space and makes it difficult for intermediate reasoning states to persist across time. We in troduce Transformers ...
201. Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents ​
Author: Paul Kassianik, Blaine Nelson, Yaron Singer
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2607.15263v2 Announce Type: replace-cross Abstract: Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomp...