arXiv cs.AI - 2026-07-27 ​
223 items collected.
1. FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills ​
Author: Zeyu Ren, Ling Yue, Ran Li, Yishu Wang, Shengxiang Xu, Hanmo Liu, Shaowu Pan, Shimin Di
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21596v1 Announce Type: new Abstract: Large language model agents increasingly solve complex tasks by constructing inference-time workflows that combine reasoning, tool use, and code execution. While such workflows enable flexible problem solving, the useful procedures discovered during ex...
2. Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals ​
Author: Nicolas Caron, Christophe Guyeux, Maxime Coulmeau, Benjamin Aynes
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21597v1 Announce Type: new Abstract: Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event prediction accuracy, not the operational coherence of a continuous risk signal. This work proposes a no...
3. Securing Multimodal AI through Internal Information Decomposition ​
Author: Jehyeok Yeon, Hyeonjeong Ha, Qiusi Zhan, Heng Ji
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21600v1 Announce Type: new Abstract: Multimodal large language models introduce attack surfaces absent in unimodal systems: adversaries can distribute malicious intent across modalities to evade unimodal safeguards. This motivates using cross-modal consistency as a detection signal rather...
4. From Frame-Level Recognition to Event-Level Confirmation: Repair Traces and Runtime Failure Analysis of Public-Space Gesture Interaction ​
Author: M. Meng, Yansong Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21601v1 Announce Type: new Abstract: Public-space gesture interaction is often evaluated as a frame-level recognition problem, but deployed systems expose a different failure boundary. In scenic kiosks, exhibition halls, and service terminals, users experience whether an intended action b...
5. Transferable Latency Prediction for Fast LLM Screening on Heterogeneous Edge Devices ​
Author: Xiaolong Tu, Vinod K. Mishra, Venkat R. Dasari, Anu G. Bourgeois, Haoxin Wang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21602v1 Announce Type: new Abstract: Accurate latency prediction is critical for deploying large language models (LLMs) on heterogeneous edge devices, where inference latency is affected by model architecture, prompt behavior, runtime backend, hardware utilization, dynamic voltage and fre...
6. AgentKVShift: Efficient KV Cache Reuse for Agentic Memory Systems ​
Author: Nilesh Prasad Pandey, Jason Kong, Lanxiang Hu, Quanling Zhao, Yujie Zhao, Onat Gungor, Hao Zhang, Tajana Rosing
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21604v1 Announce Type: new Abstract: Memory-augmented LLM agents maintain context across hundreds of interactions through agentic memory systems that actively curate retrieved content with LLM-generated metadata such as summaries, keywords, and tags. From an inference cost standpoint, eve...
7. TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward ​
Author: Debottam Dutta, Jaehoon Hahm, Jianchong Chen, Romit Roy Choudhury
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21606v1 Announce Type: new Abstract: Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a traini...
8. Spectral Flow Certificates for Depth-Aware Long-Range Propagation in Graph Neural Networks ​
Author: Ranjan Veerabhadraswamy, Ajith Jubilson Emerson
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.21607v1 Announce Type: new Abstract: Graph Neural Networks propagate information through local message passing, but the graph topologies themselves can silently prevent any amount of training from solving long-range tasks. When we deploy GNNs on new graphs, there is currently no inexpensi...
9. Coupled Hierarchical Search over Topology and Execution for Agentic Workflow Synthesis ​
Author: Dong Li, Yanchi Liu, Xujiang Zhao, Wei Cheng, Zhengzhang Chen, Xintao Wu, Zhong Chen, Chen Zhao, Haifeng Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21609v1 Announce Type: new Abstract: Although structured workflows empower Large Language Models (LLMs) to tackle complex problems, automating their creation is severely hindered by a vast combinatorial search space, frequently resulting in inflexible and resource-heavy offline training d...
10. SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text ​
Author: Miaobo Hu, Xiaobo Guo, Shuhao Hu, Bokun Wang, Rui Chen, Xin Wang, Daren Zha, Jun Xiao
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.21610v1 Announce Type: new Abstract: Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assume the schema is already available. We introduce SCOPE (Schema Construction and Ontology-induction Pipe...
11. Procedural Knowledge Is Not Low-Rank: Why LoRA Fails to Internalize Multi-Step Procedures ​
Author: Simon Dennis, Kevin Shabahang, Hao Guo, Rivaan Patil
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.21612v1 Announce Type: new Abstract: Parameter-efficient fine-tuning methods like LoRA have become the default for adapting large language models, succeeding across instruction following, style transfer, and factual adaptation. We show that for procedural knowledge--the ability to follow ...
12. The Hard Decision Layer: Evidence for Committed Inference in Transformers ​
Author: Ashwath Vaithinathan Aravindan, Mayank Kejriwal
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.21613v1 Announce Type: new Abstract: We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the Hard Decision Layer (HDL), a natural architectural property where answer option rankings stabilize abruptly du...
13. Household Movement Detection in Mixed-Format Occupancy Data Using LLM-Based Entity Resolution ​
Author: Sasirekha Oguri, John R. Talburt, Mert Can Cakmak
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21614v1 Announce Type: new Abstract: Entity resolution (ER) typically relies on pairwise similarity comparisons between records, which limits its ability to capture indirect relationships present in demographic occupancy data. An important indirect pattern arises from household movement, ...
14. FrED: External Data Influence Estimation via Domain Knowledge Graph Grounding ​
Author: Theodoros Aivalis, Iraklis A. Klampanos, Antonis Troumpoukis, Joemon M. Jose
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.21615v1 Announce Type: new Abstract: The rapid deployment of generative AI has amplified the critical need for Training Data Attribution to ensure transparency and accountability. However, current parametric approaches require computationally prohibitive access to model weights, while sim...
15. Lost in Context: Addressing Context Anxiety in Large Language Models ​
Author: Ifueko Igbinedion, Jillian Ross, Etienne Ricardez, Sertac Karaman, Eric So
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21616v1 Announce Type: new Abstract: Conventional wisdom suggests that reasoning models fail when problems exceed their capabilities. However, we find that frontier reasoning models sometimes possess the necessary capabilities to solve problems but fail due to premature self-doubt -- a ph...
16. Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models ​
Author: Gwang Gook Lee, Kenan Emir Ak, Jay Mohta, Yan Xu, Dimitrios Dimitriadis
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.21617v1 Announce Type: new Abstract: Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful transcribers: when text is imperfect, they often tend to rewrite it into a m...
17. LeafData: An Agentic System for Data Migration ​
Author: Sadanand Katukuri, Rajasekhar Bada, Navya Induri, Rohit Gandham, Lynette Pinto, Joses Selvan, Abishek Krishnamoorthy, Joseph Rozario, Pu Tian, Pavan Poudel, Yalong Wu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21618v1 Announce Type: new Abstract: Modern data migration relies on JSON configuration to define data connection, pipeline logic, and orchestration behavior. This requires domain knowledge from users and is time-consuming and error-prone. In this paper, we present LeafData, an agentic sy...
18. From Profiles to Steering Vectors: Global Sparse Priors and Local Semantic Calibration for Personalized Text Generation ​
Author: Liuji Chen, Zeyu Zhang, Xinyuan Zhang, Shuai Nie, Qiang Liu, Shu Wu, Liang Wang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21620v1 Announce Type: new Abstract: Personalized text generation requires models to capture user-specific writing styles from historical data. Existing approaches based on retrieval, parameter-efficient fine-tuning, or activation steering either introduce inference and storage overhead o...
19. FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs ​
Author: Kahou Tam, Wei Niu, Yu Bao, Xiaomin Ouyang, Chengzhong Xu, Li Li
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.LG
arXiv:2607.21624v1 Announce Type: new Abstract: Transformer-based models have enabled unprecedented capabilities across language, vision, and multimodal tasks. On-device fine-tuning of transformer models offers a privacy-preserving path to personalized AI, yet remains inefficient on mobile GPUs due ...
20. Trajectory-Aware Retrieval Agents for Temporal Decision- Making ​
Author: Jing Wang, Jie Shen, Xing Niu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21625v1 Announce Type: new Abstract: We study the problem of decision-making from long-form, temporally structured text using large language model (LLM) agents. Standard retrievalaugmented generation (RAG) pipelines fragment chronological context into isolated snippets, discarding the tem...
21. Discrete Action Space as a Prerequisite for GRPO Convergence in Small-Model Continuous Control ​
Author: Dmytro Filatov, Valentyn Fedorov, Vira Filatova
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21626v1 Announce Type: new Abstract: We study whether Group Relative Policy Optimization (GRPO) can fine-tune small language models for simulated quadrotor continuous-control tasks. In our benchmark, vanilla GRPO fine-tuning of Qwen-0.5B for 25 Hz quadrotor velocity control collapses to t...
22. Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems ​
Author: Xiaoyang Cao, Siddarth Srinivasan, Michiel A. Bakker
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.21627v1 Announce Type: new Abstract: End-to-end reinforcement learning can improve the accuracy of compound LLM systems, but it does not constrain how modules divide labor internally. We identify Role Drift, a failure mode in which modules preserve or improve end-task performance while de...
23. Wavelet Phase Diffusion for Structurally and Semantically Consistent Sim-to-Real Translation ​
Author: Kaiwen Wang, Frank Bieder, Yinzhe Shen, Carlos Fernandez, Jan-Hendrik Pauls, Omer Sahin Tas
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, eess.IV
arXiv:2607.21628v1 Announce Type: new Abstract: Simulation-to-reality translation must bridge the appearance gap between synthetic and real domains while preserving structural and semantic consistency. Conditioning-based methods achieve spatial alignment but introduce computationally expensive contr...
24. Defining AI-Native Systems: Autonomy as Revision Authority ​
Author: Cheng Tan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.OS
arXiv:2607.21659v1 Announce Type: new Abstract: AI has begun to write systems code: agents now synthesize, verify, and deploy system components. Despite this shift, "AI-native" remains a marketing term with no precise technical definition. This paper gives it one. We define AI-nativeness along a sin...
25. Persistent Computational State: A Session-Centric Runtime for Generative World Models ​
Author: Zhen Lin
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21686v1 Announce Type: new Abstract: Generative world models are increasingly driven as simulators: a planner forks a state, rolls out futures, backtracks, and returns to a visited viewpoint. Recent benchmarks establish that current video world models fail this usage, and attribute it to ...
26. What AI Red-Team Evaluations Can and Cannot Prove ​
Author: Bandana Kaur
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2607.21735v1 Announce Type: new Abstract: Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result ca...
27. From Seasonality to Semantics: Benchmarking a Hybrid Probabilistic Forecasting System for Roadblocks in Bolivia ​
Author: Rodrigo Vargas Sainz, Christian Ber'on Curti
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2607.21785v1 Announce Type: new Abstract: Roadblocks in Bolivia are a social conflict phenomenon with devastating economic impacts, estimated at losses equivalent to 4% of the national Gross Domestic Product. Despite their recurrence and impact, there is a lack of local predictive systems to a...
28. QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization ​
Author: Siwei Chen, Siqi Chen, Xupeng Miao, Bin Cui
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21793v1 Announce Type: new Abstract: Recent large reasoning models often develop long chain-of-thought responses during reinforcement learning (RL), resulting in high inference latency and deployment cost. Existing methods for response length control typically rely on explicit length pena...
29. DAGForge: Auditable Causal DAG Authoring with Biomedical Literature ​
Author: Yi-han Sheu, Michael R. Steigman, Yu Zhou, Bo Wang, Fan-Yu Yen, Jordan W. Smoller
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21859v1 Announce Type: new Abstract: Constructing causal directed acyclic graphs (DAGs) is a core step in biomedical causal analysis, yet it remains a largely manual process. Analysts must connect study variables to prior literature, evaluate uncertain causal claims, and preserve sufficie...
30. When Is a Learned Command Adapter Worth It? Closed-Loop Identification and Counterfactual Auditing of Frozen Locomotion Policies ​
Author: Zongtan Li
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY
arXiv:2607.21867v1 Announce Type: new Abstract: Adding a learned adapter to a frozen, command-conditioned locomotion policy is worthwhile only if the interface exposes improvements that are both real and recoverable from deployment-time observations. We introduce an adapter necessity audit that sepa...
31. Multi-Agent System-driven Digital Twins for predictive maintenance: architectures, technologies and open research challenges ​
Author: Korota Ars`ene Coulibaly, Mohamed Hamlich
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21873v1 Announce Type: new Abstract: Digital twins have emerged as a foundational technology within the context of Industry 4.0, offering a paradigm for the real-time virtual representation of physical systems. However, managing their growing complexity, particularly in distributed indust...
32. TRW: TRACE-RealWorld---An Auditable Consistency Contract for World Models as Materialized Views ​
Author: Edward Y. Chang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.DB
arXiv:2607.21910v1 Announce Type: new Abstract: TRACE-RealWorld addresses a core data-management problem: maintaining an actionable materialized view over a continuously changing physical world when reads of the base state are priced, delayed, heterogeneous, and fallible. Its data-management contrib...
33. Semiotic logical hexagon theory for LLM logical reasoning ​
Author: Yunyao Zhang, Xinglang Zhang, Zeliang Chen, Junqing Yu, Zikai Song
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21933v1 Announce Type: new Abstract: Large language models (LLMs) have become powerful tools for language understanding and logical reasoning. However, they still make mistakes when a problem requires both understanding meaning and following logic. A key reason is that natural-language st...
34. Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning ​
Author: Heyang Jiang, Henry Liu, Baharan Mirzasoleiman
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22002v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a highly effective framework for improving LLM reasoning, with methods such as GRPO among its most successful instantiations. However, GRPO relies on repeated generation of long chain...
35. Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents ​
Author: Suman Navaratnarajah, Taehyoung Kim, Jona Ruthardt, Ishaan Bhimwal, Ryousuke Yamada, Yannik Blei, Wolfram Burgard, Yuki M Asano
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.RO
arXiv:2607.22014v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level instruction. We introduce Mission...
36. Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode ​
Author: Nanbeige Lab, :, Chen Yang, Chengrui Huang, Fufeng Lan, Hanhui Chen, Hao Zhou, Huatong Song, Jiaqi Cao, Jiaying Zhu, Jinlin Niu, Kai Wang, Lisheng Huang, Qiliang Liang, Ran Le, Ruixiang Feng, Shuang Sun, Tao Gu, Tao Zhang, Tianyu Luo, Yang Song, Yun Xing, Yuntao Wen, Ziyao Xu, Zongchao Chen, Zongqiang Li
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.22083v1 Announce Type: new Abstract: We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathe...
37. Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models ​
Author: Junlin Fang, Do Nguyen-Thanh, Xiaogang Xu, Zhen Fang, Sean Du
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.22098v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long reasoning traces before producing final answers. While these traces may contain useful signals for hallucination detection, harnessing them is non-trivial because long trajectories often include noisy steps t...
38. Industrial Tokenization for LLM-Based Health Intelligence: A Federated Architecture for Industrial Evidence Integration ​
Author: Deshui Li, Xiao-Ming Yuan, Zishun Wang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.22153v1 Announce Type: new Abstract: Industrial health management increasingly relies on heterogeneous information sources, including condition monitoring systems, supervisory control and data acquisition systems, maintenance records, inspection results, and prognostic models. Although la...
39. Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents ​
Author: Valentin Tablan, Scott Taylor, Kristoffer Bernhem
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22157v1 Announce Type: new Abstract: AI agents encounter learning opportunities in every episode they run, and discard nearly all of them: the underlying models are frozen at deployment, so an agent that resolves a difficult request today starts from zero when it recurs tomorrow. Yet ordi...
40. Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning ​
Author: Guanqun Zhao, Zijun Xie, Binbin Zheng, Enlei Gong, Jiafeng Lu, Yehan Yang, Aoqi Hu, Zeyu Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22186v1 Announce Type: new Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause policy ...
41. AI4PLE: A Methodology for Integrating AI into Product Line Engineering ​
Author: Bedir Tekinerdogan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22260v1 Announce Type: new Abstract: Reuse-based development has become increasingly important in the creation of complex systems, offering significant opportunities to reduce costs, improve quality, and accelerate time-to-market. Product Line Engineering (PLE) provides a systematic appro...
42. A Roadmap to Impactful Pluralistic Alignment Research ​
Author: Elinor Poole-Dayan, Jillian Fisher, Atoosa Kasirzadeh, Jacob Andreas, Mitchell Gordon, Michiel A. Bakker
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22305v1 Announce Type: new Abstract: Pluralistic value alignment---the goal of building AI systems that represent and serve diverse human values and perspectives---has emerged as an active research agenda. Yet, there's no public evidence that it has shaped the training or evaluation of th...
43. Learning Structural Convergence: A Neuro-Symbolic Benchmark for Temporal Reasoning ​
Author: Michael Romei De Socio, Gian Luca Pozzato, Alessio Merlo
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22365v1 Announce Type: new Abstract: High-complexity operational environments require methods that detect and anticipate temporally distributed patterns rather than classify isolated events. This paper introduces TRACTA (Temporal Reasoning and Capability-Trajectory Analysis), a controlled...
44. Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI ​
Author: Jiaqi Shao, Hanck Chen, Wei Zhang, Maxm Pan, Bing Luo
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22368v1 Announce Type: new Abstract: Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent re...
45. IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation ​
Author: Varun Gumma, Navonil Majumder, Soumitra Sinhahajari, Soujanya Poria
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22375v1 Announce Type: new Abstract: Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize ideas independently for either Quality or Diversity. ...
46. Agentic Root Cause Analysis through Evidence-Grounded Reasoning ​
Author: Amaury Wei, Olga Fink
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.22385v1 Announce Type: new Abstract: Diagnosing the root cause of anomalies is essential for safe industrial operation. Despite extensive sensor instrumentation, formulating hypotheses and gathering evidence remains a manual process, creating a major operational bottleneck. While existing...
47. SceneActBench: Can Agents Act on the 3D Scenes They See? ​
Author: Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang, Pu Jian, Huanjin Yao, Jiarui Yao, Haowei Lin, Chunchao Guo, Zhuo Chen, Wenkai Lyu, Jianzhu Ma, Xueqian Wang, Wenxi Zhu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2607.22393v1 Announce Type: new Abstract: Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under eval...
48. Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture ​
Author: Halil Burak Noyan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CR
arXiv:2607.22445v1 Announce Type: new Abstract: Enterprise AI agents are typically granted static credential sets at configuration time, holding every tool the role might need for every task they perform. This persistent over-privilege expands the attack surface. We argue that capability scoping mus...
49. TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI ​
Author: Ritik Raj, Souvik Kundu, Sarbartha Banerjee, Dheemanth Joshi, Ishita Vohra, Tushar Krishna
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2607.22465v1 Announce Type: new Abstract: Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, agentic appli...
50. The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents ​
Author: Darshan Tank, Baran Nama
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22520v1 Announce Type: new Abstract: Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this metric hides an important cost: skills can also make agents worse. We measure both sides by comparing agents with and without skills a...
51. Explainable Reinforcement Learning for assisting Air Traffic Controllers ​
Author: Anduel Mehmeti, Gabriella Gigante, Salvatore Venticinque
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22525v1 Announce Type: new Abstract: To effectively integrate AI into high-stakes, critical environments such as healthcare, autonomous driving, and aviation--and to advance toward higher levels of automation and seamless human-AI collaboration--building trust in AI-driven solutions is es...
52. Do emulated quantum circuits change what CNNs look at? Performance and explainability comparison in medical image classification ​
Author: Guillermo Rubi~nos Rodr'iguez, Mart'in Ottavianelli, Mateo Alonso, Gonzalo Bl'azquez Gil, Boris-Stephan Rauchmann, Pablo D'iez-Valle, Sergio Altares-L'opez
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG
arXiv:2607.21186v1 Announce Type: cross Abstract: Numerous studies have analyzed the use of hybrid quantum-classical convolutional neural networks as a promising alternative to classical deep learning. However, network components on quantum hardware impose fundamental limitations, while the scalabil...
53. Control panels to clarify user intent with Large Language Models ​
Author: Ben Shneiderman
Published: 7/27/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2607.21598v1 Announce Type: cross Abstract: Typical user interfaces for Large Language Models present a blank prompt window that invites a natural language query by users, but offers little guidance. This paper proposes a visual control panel interface that would provide more cues to the seman...
54. Decoupled Attention Fusion: Accelerating RAG with Efficient KV Cache Reuse ​
Author: Xiabao Wu, Wentao Liu, Yongchao Liu, Jiajun Zheng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.PF, cs.AI
arXiv:2607.21599v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) effectively mitigates hallucinations in Large Language Models (LLMs) but suffers from prohibitive Time-To-First-Token (TTFT) latency in long-context scenarios. Reusing pre-computed document KV caches addresses thi...
55. Analyzing Middle School Students' Dialogue and Behaviors during Collaborative AI Chatbot Development Using Ordered Network Analysis ​
Author: Shan Zhang, Andres Felipe Zambrano, Xiaoyi Tian, Yukyeong Song, Anthony F. Botelho, Kristy Elizabeth Boyer, Maya Israel, Shiyan Jiang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2607.21603v1 Announce Type: cross Abstract: As Artificial Intelligence (AI) education has become a key component of K-12 curricula, activities such as designing and developing conversational agents are increasingly used as instructional practice. Prior work has primarily examined these activit...
56. From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE ​
Author: T. Y. Emmy Lai, Sven Giesselbach, Matthias Koch, H'ector Allende-Cid
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL, cs.CY
arXiv:2607.21608v1 Announce Type: cross Abstract: With the EU AI Act entering into force, organizations developing or operating AI systems face new obligations on transparency, risk management, and traceability. For Requirements Engineering (RE), these obligations must be translated into testable, a...
57. A Systematic Survey on Image Description Techniques for STEM Domains ​
Author: Marco Cardia, Letizia Angileri, Marina Buzzi, Giulio Galesi, Barbara Leporini
Published: 7/27/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2607.21611v1 Announce Type: cross Abstract: The proliferation of visual data in Science, Technology, Engineering, and Mathematics (STEM) fields presents accessibility barrier for individuals with blindness or visual impairments. While recent advances in Artificial Intelligence (AI) offer new o...
58. Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization ​
Author: Bingjun Luo, Jialin Guo, Yue Yao, Xinpeng Ding
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.21619v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Existing content-based jailbreaks are often inconsistent and show unsatisfying performance against the ...
59. Local Synaptic Rules Can Implement a SIGReg Gradient Without Backpropagation ​
Author: Martin Andrews
Published: 7/27/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CV, cs.LG
arXiv:2607.21622v1 Announce Type: cross Abstract: We prove that two canonical local synaptic learning rules, the potentiation arm of spike-timing-dependent plasticity (STDP$^+$) and homeostatic plasticity (instantiated here via flashlight granule-cell-like neurons), together can implement the exact ...
60. On the Depth Scalability of Logic Gate Networks ​
Author: Taegun An, Dohun kim, Haebeom Lee, Changhee Joo
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO
arXiv:2607.21633v1 Announce Type: cross Abstract: Logic Gate Networks (LGNs) implement computation through compositions of Boolean operations, yet unlike classical Boolean circuits, existing LGNs do not reliably benefit from increased depth. We identify two distinct causes: optimization collapse in ...
61. MotifRole-Diff: Risk-Optimal Role-Aware Corruption for Masked Molecular Graph Diffusion ​
Author: Tasfia Nuzhat Ornee, Elias Hossain, Ivan Garibay, Niloofar Yousef
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21634v1 Announce Type: cross Abstract: Masked discrete diffusion for molecular graph generation typically applies a uniform corruption schedule to all tokens in a lossless graph-to-sequence representation, implicitly treating structurally heterogeneous molecular components as equally diff...
62. Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models ​
Author: Jie Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21636v1 Announce Type: cross Abstract: Synthetic tabular data is valued for preserving not only each column's marginal distribution but the dependencies between columns -- structure that carries much of the discriminative signal for minority classes in imbalanced domains such as fraud and...
63. Tool-Guided Retrieval-Augmented Repair for Securing LLM-Generated C Code ​
Author: Vidyut Sriram, Saatvik Pradhan, Suman Saha
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.21641v1 Announce Type: cross Abstract: Large language models can generate C code from natural-language descriptions, but resulting programs often contain security vulnerabilities and compilation errors, posing risks for embedded and resource-constrained systems. This work investigates how...
64. Multi-Horizon Consistency as Geometry: When Latent Dynamics Contract, and When They Do Not ​
Author: Kavya Bhand, Aadi Joshi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21645v1 Announce Type: cross Abstract: Multi-horizon latent consistency is a common training knob in video predictors and world models, but practitioners rarely know what it does to transition geometry. We treat lambda, the weight on multi-step latent agreement, as a diagnostic control an...
65. A Drift Stable Quantum Federated Learning for Intelligent Services ​
Author: Shanika Iroshi Nanayakkara, Shiva Raj Pokhrel
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21647v1 Announce Type: cross Abstract: Quantum federated learning enables distributed clients to train quantum neural networks without sharing local data, making it promising for privacy-aware intelligent services. Intelligent services in this context refer to privacy-sensitive distribute...
66. Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa? ​
Author: Zuodong Xiang, Yike Zhang, YueMing Zhang, Hailu Xu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.21656v1 Announce Type: cross Abstract: Developers increasingly use two coding agents together: one writes a draft, and the other reviews it. However, it is not clear whether the pairing is worth its cost and time, or whether the order of the pairing matters. We run a controlled experiment...
67. Generative and multimodal AI for materials prediction and design: Progress, challenges, and perspectives ​
Author: Xianyuan Liu, Charles Anjah, Benjamin E. Jolly, Jonathon F. S. Markanday, Joshua Berry, Haolin Wang, Nicola A. Morley, Robert D. J. Oliver, Alexandra J. Ramadan, Delvin Ce Zhang, Katerina A. Christofidou, Haiping Lu
Published: 7/27/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, cs.LG
arXiv:2607.21660v1 Announce Type: cross Abstract: Artificial intelligence (AI) is accelerating materials prediction and design by enabling efficient exploration of chemical and structural spaces, with particular promise for novel materials discovery. However, novelty in materials discovery encompass...
68. Ordered Action Tokens for Visuomotor Policy Learning ​
Author: Chaoqi Liu, Yue Zhao, Haonan Chen, Xiaoshen Han, Jiawei Gao, Ehsan Adeli, Yilun Du
Published: 7/27/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2607.21670v1 Announce Type: cross Abstract: Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies. Existing approaches either rely on analytical discretization methods that produce prohibitively long toke...
69. Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images ​
Author: Ronak Bhalgami
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CV
arXiv:2607.21672v1 Announce Type: cross Abstract: Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whether models can still solve code tasks after this transformation. We examine a different systems ques...
70. Self-Poisoning in Adaptive Out-of-Distribution Detection: A Sharp-Threshold Theory and Certified Label-Free Calibration ​
Author: Vishnu Bindu Balachandran
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR, cs.CV, stat.ML
arXiv:2607.21673v1 Announce Type: cross Abstract: Test-time adaptive out-of-distribution (OOD) detectors update a memory bank from the unlabelled stream. We show this adaptation obeys a provable dynamical law. Modelling bank impurity as a generalized P'olya urn, we prove almost-sure convergence to ...
71. Enhancing SLMs for Sustainable Code Optimization in Radio-Astronomy ​
Author: Elisa Chiarotto, Jingbo Li, P. Chris Broekema, Rob V. van Nieuwpoort
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.21677v1 Announce Type: cross Abstract: Recent Large Language Models (LLMs) can produce and optimize complex code. We investigate the use of LLMs to generate and optimize code for large-scale sciences, focusing on radio astronomy and sustainability. The LOFAR telescope is currently being u...
72. A Defense of the Quadratic Model ​
Author: Alexandru Meterez, Pranav Ajit Nair, Depen Morwani, Cengiz Pehlevan, Sham Kakade, Alex Damian
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC, stat.ML
arXiv:2607.21716v1 Announce Type: cross Abstract: Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimizat...
73. Deep Sigma Point Processes for RCS Modeling in Spaceborne SAR Imagery ​
Author: Khalid El-Darymli, Christoph H. Gierull, Katerina Biron, Weimin Huang
Published: 7/27/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG
arXiv:2607.21745v1 Announce Type: cross Abstract: Radar cross-section (RCS) modeling is foundational to advancing the utility and sensitivity of spaceborne radar systems. This study introduces a deep sigma-point process (DSPP) model for predicting RCS in synthetic aperture radar (SAR) imagery using ...
74. Co-design of LLM-based preference agents: participation may drive overtrust ​
Author: Michael J. Fell
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2607.21757v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising w...
75. Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks ​
Author: Michael Kouremetis, Ads Dawson, Raja Sekhar Rao Dheekonda, Brian Greunke
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2607.21763v1 Announce Type: cross Abstract: Large language model (LLM) agents routinely cheat on cybersecurity benchmarks, inflating reported pass rates far beyond genuine capability. Prior audits of Cybench found cheating in 0.3-3.4% of traces, implicating only a handful of models. We present...
76. Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders ​
Author: Pablo Santiago Potes Velasco, Mar'ia del Mar Garc'ia Matabanchoy, 'Oscar Juli'an P'erez Ladino, Jhoan Stevan Mosquera Ortiz, Nicol'as Lozano Mazuera, Gilber Alexis Corrales Gallego
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.21774v1 Announce Type: cross Abstract: Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated. This pilot study examines whether Qwen2.5-7B-Instruct internally represents Colombian identity, socioeconomic sta...
77. AI-Integrated Scientific Inquiry: A Practice-Centered Vision for Science Education ​
Author: Arne Bewersdorff, Matias Rojas, Xiaoming Zhai
Published: 7/27/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2607.21777v1 Announce Type: cross Abstract: Artificial intelligence (AI) has become part of scientific inquiry. Scientists use AI to observe and measure phenomena, to identify patterns in data, and to build models. As AI moves into scientific inquiry, it gains relevance for science education: ...
78. Graph-Theoretic Neural Network Fragmentation with Covariant Direct Molecular Force Learning: Enabling Coupled-Cluster Accuracy AIMD for Fluxional Systems ​
Author: Xiao Zhu, Srinivasan S. Iyengar
Published: 7/27/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.AI, physics.comp-ph
arXiv:2607.21779v1 Announce Type: cross Abstract: Accurate ab initio molecular dynamics (AIMD) simulations of complex, fluxional chemical systems are severely limited by the high computational scaling of correlated electronic structure methods. To overcome this bottleneck, we present a robust, graph...
79. Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms ​
Author: Abu Tyeb Azad, Fahim Ahmed, Ishita Sur Apan, Ezharuddin Jubaer, Sumaiya Karim Katha, Armun Alam, Amin Ahsan Ali, Aman Chadha, Md Mofijul Islam, AKM Mahbubur Rahman
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV
arXiv:2607.21780v1 Announce Type: cross Abstract: Document packets, multiple documents concatenated into a single file, are common in government and administrative workflows, yet splitting them into their constituent documents is difficult, especially for low-resource languages. We introduce Khondo ...
80. MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery ​
Author: Grace Fan, Eden Wu, Majid Daliri, Juliana Freire
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2607.21781v1 Announce Type: cross Abstract: Join discovery is a core task in dataset search, enabling users to find columns that can be joined with a given query column. Early approaches focused on equi-joins, but data lakes and open-data repositories often contain columns whose values refer t...
81. Probing Speaker Identity Sensitivity in Audio Deepfake Detectors ​
Author: Daniyal Kabir Dar, Arun Ross
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CR, cs.LG
arXiv:2607.21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold w...
82. ToolGuardian: Declarative Security for AI Agent-Tool Interactions ​
Author: Arun Ravindran, Saurabh Deochake
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2607.21835v1 Announce Type: cross Abstract: LLM agents increasingly rely on external tools, expanding capability while creating a new security boundary: third-party tools may appear benign at the interface level while embedding unsafe behavior in implementation. Existing defenses rely on weak ...
83. SCALE: Self-Supervised Constraint-Aware Layout GEneration for Local P&R DRV Fixing at Advanced Nodes ​
Author: Chia-Tung Ho, Haoyu Yang, Guanglei Zhou, Yoshi Nishi, Yaguang Li, Walker Turner, Cunxi Yu, Yiran Chen, Brucek Khailany
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.21850v1 Announce Type: cross Abstract: As semiconductor manufacturing advances toward sub-2nm nodes, local place-and-route (P&R) design-rule violation (DRV) fixing is increasingly limited by complex rule interactions, dense multi-layer routing geometries, and foundry-specific constraints....
84. LeAct: Learning to Reason from Expert Actions ​
Author: Ziran Yang, Chengshuai Shi, Raj Ghugare, Benjamin Eysenbach, Karthik Narasimhan, Chi Jin
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.21856v1 Announce Type: cross Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines, classical planners, theorem...
85. Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA ​
Author: Joan Figuerola Hurtado
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.21861v1 Announce Type: cross Abstract: We study baking documents directly into the weights of a 4-bit Gemma-4-e4b model via LoRA, so a system can answer questions about a corpus closed-book: no retrieval and no context-window budget. Across roughly 100 training runs from single documents ...
86. ISPCloak: Weaponizing ISP for Optimization-Free Physical Camouflage against Deepfake Detectors ​
Author: Jiale Zhao, Jiajun Wan, Lei Tang, Ye Qin, Kebing Jin, Jinghui Qin
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CR
arXiv:2607.21897v1 Announce Type: cross Abstract: The rapid advancement of generative models has spurred the critical need to evaluate the worst-case robustness of deepfake detectors. In this paper, we reveal a fundamental blind spot in current forensic paradigms: while existing detectors excel at c...
87. Interventional Score Geometry for Causal Inference ​
Author: Mojtaba Eslami
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, econ.EM
arXiv:2607.21914v1 Announce Type: cross Abstract: Let $p(x)$ be the joint density of variables $X$, and let $\psi(x)=\nabla_x\log p(x)$ be its score field. Geometry constructed from $p$ and $\psi$ alone cannot identify causal direction: structural models with the same observational distribution have...
88. Generalized Neural Operator for Parametric and Boundary-Value Problems ​
Author: Ruoyan Li, Yizhou Sun, Wei Wang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CE, cs.AI
arXiv:2607.21932v1 Announce Type: cross Abstract: Developing foundational neural simulators for Partial Differential Equations (PDEs) requires robust generalization across diverse physical parameters and boundary conditions. However, current deep learning approaches largely face a structural trade-o...
89. MA-DAR: Manifold-Aligned Dynamic Adaptive Routing for Continual Temporal Knowledge Graph Reasoning ​
Author: Xiangjun Shi, Chong Mu, Jinchuan Zhang, Lizong Zhang, Yuefeng He, Shang Liu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21949v1 Announce Type: cross Abstract: Continual temporal knowledge graph (TKG) reasoning aims to continuously incorporate newly emerging facts while preserving previously acquired knowledge. Replay-based continual learning has achieved promising performance by revisiting historical repre...
90. ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset ​
Author: Shashank Rao Marpally, Allan Wang, Atharva Ghotavadekar, Renato Alexandre Ribeiro, Nhat Le, Pilar Bachiller-Burgos, Pranav Goyal, Subham Agrawal, Yasuhiro Nitta, Howard Ziyu Han, Daeun Song, Masaki Kuribayashi, Kohei Uehara, Xiyue Wang, Yangzhe Kong, Duc M. Nguyen, Amirreza Payandeh, Gerardo P'erez-Gonz'alez, Alejandro Torrej'on-Harto, Jeeho Ahn, Tisha Jain, Andrew Stratton, Elvin Yang, Jorge de Heuvel, Nico Ostermann-Myrau, Sai Anudeep Sajja, Mithilya Raj, Daisuke Sato, Gaston Rouquette, Nikolas Martelaro, Maki Sugimoto, Hironobu Takagi, Chieko Asakawa, Maren Bennewitz, Aaron Steinfeld, Xuesu Xiao, Christoforos Mavrogiannis, Harold Soh
Published: 7/27/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.21964v1 Announce Type: cross Abstract: Understanding how robots and humans move in shared spaces is essential for designing effective social robot navigation policies and predicting human behavior. However, existing datasets often lack the diversity needed to capture differences in cultur...
91. TextSLIP: Text Self-Supervised CLIP for Medical Report Generation ​
Author: Haoyu Jiang, Ziping Cong
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.21970v1 Announce Type: cross Abstract: Automating radiology report generation is important for improving reporting consistency and clinical workflows . While Contrastive Language--Image Pretraining (CLIP) has advanced medical vision language modeling, existing CLIP-style approaches may st...
92. Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning ​
Author: Shujin Wu, Cheng Qian, Xiusi Chen, Heng Ji
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.21971v1 Announce Type: cross Abstract: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We hypothesize that the success of such evolution frameworks hinges on meta-skills, such as self-reflect...
93. J-CoT: Chain-of-Thought in J-Space ​
Author: Junde Wu, Jiayuan Zhu, Fengling Liu, Minhao Hu, Jiazhen Pan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.21981v1 Announce Type: cross Abstract: Chain-of-thought prompting improves language-model reasoning by carrying intermediate states across successive computation steps. However, relying on natural language as the only recurrent interface is overly restrictive, since many transient computa...
94. Unified Static-Dynamic Pruning for Efficient LLM Inference ​
Author: Jinhyeok Kim, Yejoon Lee, Jaeyoung Do
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.AR, cs.LG
arXiv:2607.21985v1 Announce Type: cross Abstract: The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dominate inference cost. Weight pruning offers a prom...
95. From Perturbation Correction to Geometry-Aware Sampling: Sharpness-Guided Equilibrium Sampling for Balanced Flat Minima in Long-Tailed Learning ​
Author: Jiaxin Deng, Junbiao Pang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.21999v1 Announce Type: cross Abstract: Long-tailed learning couples two sources of poor generalization: head classes dominate training exposure, while under-represented classes often converge to sharper regions of the loss landscape. Conventional re-sampling addresses the former without c...
96. Practical Graph Optimisation and AI-Driven Models for Active Directory Security Hardening ​
Author: Huy Q. Ngo
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2607.22009v1 Announce Type: cross Abstract: Microsoft's Active Directory (AD) is a directory service that enables the IT admin to manage security permissions and control access within a Windows domain network. As a core management system in many of organisation, AD has become a primary target ...
97. Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning ​
Author: Hao Yang, Jin Wang, Xuejie Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.22013v1 Announce Type: cross Abstract: Multimodal chain-of-thought (CoT) reasoning integrates visual and textual cues through step-by-step inference. In small models with limited token budgets, modality-interaction fusion often suppresses tiny cross-modal differences. In particular, multi...
98. EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection ​
Author: Hao Yang, Jin Wang, Xuejie Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.22016v1 Announce Type: cross Abstract: MEMEs are widely used on the internet and often carry strong elements of sarcasm or irony. Understanding their hidden meanings typically requires a joint interpretation of text and vision. Existing methods focus on the dual-stream vision-language mod...
99. Agent Security Needs Redefinition through a Holistic Framework ​
Author: Vincent Siu, Jingxuan He, Kyle Montgomery, Zhun Wang, Chenguang Wang, Dawn Song
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2607.22024v1 Announce Type: cross Abstract: Agent security is widely treated as a question about action content. Defenses ask whether an instruction looks malicious. Benchmarks ask whether an agent performs a harmful sounding action. \textbf{We argue that agent security is fundamentally a cont...
100. Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference Accelerators ​
Author: Afzal Ahmad, Gaoyu Mao, Shoubo Hu, Hui-Ling Zhen, Mingxuan Yuan, Xinyu Chen, Wei Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AR, cs.AI
arXiv:2607.22038v1 Announce Type: cross Abstract: Multi-task inference models share a single backbone across diverse tasks, yet execute identical computation regardless of which task is active - wasting energy and cycles on task-irrelevant operations. We observe that the task command, typically avai...
101. CEL: Comprehensive Counterfactual Explanations Library and Benchmark ​
Author: Oleksii Furman, {\L}ukasz Lenkiewicz, Marcel Musia{\l}ek, Maciej Zi\k{e}ba
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22045v1 Announce Type: cross Abstract: Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome. While early methods primarily focused on min...
102. Multiplicity of Stable Attractors in Disordered Neural Models ​
Author: Raffaele Marino, Roberto Livi, Antonio Politi
Published: 7/27/2026, 4:00:00 AM
Categories: cond-mat.dis-nn, cond-mat.stat-mech, cs.AI, nlin.CD
arXiv:2607.22047v1 Announce Type: cross Abstract: We show how large-deviation statistics allows one to obtain reliable estimates of the multiplicity of stable fixed-points in a model of neural ordinary differential equations previously employed in computational tasks. The result is obtained by devel...
103. Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination ​
Author: Isak Hwang, Yoon Pyo Lee
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.22067v1 Announce Type: cross Abstract: The integration of large language models (LLMs) into the nuclear power industry requires outputs grounded in domain-specific knowledge. This study evaluates a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) on its capacity to apply...
104. FSE: Continual Learning for Named Entity Recognition by Fast-Slow Experts ​
Author: Yunan Zhang, Yang Fan, Heng Li, Xiangping Wu, Qingcai Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.22075v1 Announce Type: cross Abstract: Continual Learning for Named Entity Recognition (CLNER) enable models to incrementally learn new entity types without forgetting previously acquired ones. However, existing methods suffer from catastrophic forgetting and insufficient exploitation of ...
105. A Leakage-Free Stacked Ensemble Method for Multiclass Classification ​
Author: S. P. Sharmila, Aruna Tiwari
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22081v1 Announce Type: cross Abstract: Multiclass classification is a fundamental problem across a wide range of domains. It is still challenging due to possession of high inter-class similarity, class imbalance datasets, and variability in data distributions. Rule-based classifiers such ...
106. MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond ​
Author: Lorenzo Concina, Seraphina Fong, Marco Matassoni, Alessio Brutti
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, eess.AS
arXiv:2607.22100v1 Announce Type: cross Abstract: Lightweight projectors are an established way to connect pre-trained speech encoders with large language models (LLMs), mapping acoustic features into token-level embeddings for tasks like ASR and spoken question answering. Existing systems, however,...
107. Benchmarking Text-to-SQL under Role-Based Access Control ​
Author: Yang Fei, Yangfan Jiang, Yin Yang, Xiaokui Xiao
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2607.22115v1 Announce Type: cross Abstract: Given a database S and a natural language question Q, text-to-SQL systems aim to generate an SQL query that correctly answers Q when executed against S. Currently, popular text-to-SQL benchmarks mostly assume unrestricted access to S; in practice, ho...
108. One Hand Watches The Other: Dynamic Multi-Agent Cooperation for Sample-Efficient Bimanual Manipulation in Dynamic Environments ​
Author: Jan Ole von Hartz, Abhinav Valada, Joschka Boedecker
Published: 7/27/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2607.22119v1 Announce Type: cross Abstract: Multi-stream robot manipulation policies achieve unparalleled sample efficiency and generalization by modeling actions relative to environmental reference frames. However, existing approaches typically assume these frames to be strictly exogenous. Th...
109. CARDIAG: A Dense Segment Classification Benchmark of Deep Learning Architectures for Coronary Angiography ​
Author: Dominik Bernard Lau, Hubert Malinowski, Jerzy Szyjut, Adam Brzeski, Tomasz Dziubich, Rados{\l}aw Targo'nski, Tomasz Figatowski, Natalia Zieli'nska
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.22139v1 Announce Type: cross Abstract: Accurate pixel-level classification of coronary angiograms is critical for cardiovascular disease assessment, yet the field lacks standardized evaluation protocols. In this work we demonstrate a new benchmark for the assessment of deep learning model...
110. TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex ​
Author: Yuliang Yan, Shuo Yan, Haochun Tang, Yiqin Sun, Enyan Dai
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22143v1 Announce Type: cross Abstract: Molecular glue degraders have emerged as a promising strategy for targeted protein degradation by inducing ternary complex formation between an E3 ubiquitin ligase and a target protein. Despite their therapeutic potential, computational design of mol...
111. dRAE: Representation Autoencoder with Hyper-Spherical Codes ​
Author: Tianren Ma, Lin Long, Chuyan Chen, Mu Zhang, Junbo Zhao, Tong Zhang, Qixiang Ye
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.22148v1 Announce Type: cross Abstract: In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semanti...
112. DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents ​
Author: Junming Chen, Junyang Jiang, Xu Chen, Zibo Liang, Kai Zheng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CL, cs.LG
arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation and production operations: live-environment fidelity (multi-turn read-write interaction with a runnin...
113. Learning Spatiotemporal Decision Priors for Efficient Path Planning under Partial Observability ​
Author: Yi Liu, Hongda Zhang, Leyao Zou, Chunlei Meng, Ziqing Zhou, Yuning Chen, Zhuo Zou, Lida Xu, Zhongxue Gan, Chun Ouyang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.22166v1 Announce Type: cross Abstract: Path planning under partial observability remains challenging because an agent must make long-horizon navigation decisions from only locally bounded observations. Nevertheless, historical trajectories contain reusable experience-guided directional pr...
114. From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models ​
Author: Shixin Fang (Fudan University), Jiachen Wo (Fudan University), Wenjuan Qin (Fudan University), Sihang Jiang (Fudan University), Yanghua Xiao (Fudan University)
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.22182v1 Announce Type: cross Abstract: Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation limits cross-study comparison, obscures capabilities tasks recruit, and ...
115. Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs ​
Author: Yuheng Zong, Minghua Wang, Xin Zhao, Zhi-Hui Zhan, Antonio Plaza, Jon Atli Benediktsson
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.22205v1 Announce Type: cross Abstract: Remote sensing multimodal large language models (RS-MLLMs) have improved general aerial-image understanding. However, Earth observation applications require fine-grained scenario specialization, constrained by scarce high-quality scenario data and in...
116. Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity ​
Author: Pengzhao Lyu, Yeun Joon Kim, Hanlin Xiao, Yingyue Luna Luan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.22218v1 Announce Type: cross Abstract: Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from human judgments. Ac...
117. TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-Resolution ​
Author: Sicheng Gao, Zhuyun Zhou, Yixuan Liu, Tong Shen, Zongwei Wu, Radu Timofte
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.22231v1 Announce Type: cross Abstract: Video super-resolution (VSR) using large-scale Diffusion Transformer (DiT) priors achieves exceptional perceptual quality but is often impractical due to the quadratic computational cost of processing dense spatio-temporal token sequences. Existing e...
118. Optimization of time-consuming experimental conditions using pseudo-experimental data guided by adaptive polynomial regression ​
Author: Hirotaka Sugawara, Yujin Taguchi, Kei Minagawa, Yusuke Hiki, Takashi Morikura, Akira Funahashi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22238v1 Announce Type: cross Abstract: Bayesian optimization (BO) is an optimization method that sequentially proposes the next candidate explainable variables for optimizing target variables by balancing exploration and exploitation. BO is often used under a limited evaluation budget, su...
119. IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning ​
Author: Wei Zhang, Xinwu Liu, Yihang Cheng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22251v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning method for large language models, but its performance depends strongly on how a fixed rank budget is distributed across Transformer modules. Existing adaptive-rank methods us...
120. Explicit Iteration Complexity of Exact Data-Driven Inverse Optimization for Integer Linear Programs ​
Author: Akira Kitaoka
Published: 7/27/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.LG, stat.ML
arXiv:2607.22263v1 Announce Type: cross Abstract: A data-driven inverse optimization problem (DDIOP) is the problem of estimating the objective-function parameters (weights) that explain observed optimal-solution data, and it arises in many applications, including integer linear programming (ILP). I...
121. Neptuna: A Comprehensive Machine Learning Framework for Benchmarking Complex Multiphase Flows ​
Author: Harish Ramachandran, Bj"orn Kimpel, Thomas Paula, Josef Winter, Steffen Schmidt, Nikolaus Adams
Published: 7/27/2026, 4:00:00 AM
Categories: physics.flu-dyn, cs.AI
arXiv:2607.22280v1 Announce Type: cross Abstract: Compressible multiphase flows involving shocks and material interfaces arise in applications such as bubble collapse and droplet breakup, where strong nonlinear interactions produce complex interface deformation, mixing, and multiscale dynamics. Deve...
122. Towards Trustworthy and Cost-Efficient Data Integration: From Na\"ive RAG to Agentic RAG ​
Author: Chuangtao Ma, Arijit Khan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2607.22319v1 Announce Type: cross Abstract: Large language models (LLMs) and AI agents have demonstrated strong potential for data integration in zero-shot and few-shot settings. However, they continue to face significant accuracy and cost challenges in enterprise environments due to a persist...
123. Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization ​
Author: Hao Wang, Kun Yuan, Wenlin Zhong, Minglei Zhang, Han Xiao, Ming Sun, Honggang Qi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.22334v1 Announce Type: cross Abstract: Open-weight language models from different families exhibit complementary capabilities, motivating their consolidation into a compact student through on-policy distillation (OPD). However, full-vocabulary OPD typically assumes a shared tokenizer, whi...
124. Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education ​
Author: Stephan Vonschallen, Karim Kaufmann, Dominique Oberle, Friederike Eyssel, Theresa Schmiedel
Published: 7/27/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.22345v1 Announce Type: cross Abstract: Generative social robots (GSRs) powered by large language models offer new possibilities for personalized tutoring in higher education, but also introduce risks related to misinformation, missing transparency, or reinforcing incorrect student respons...
125. Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment Interactions ​
Author: Jorge Bacca, Kebin Contreras, Luis Toscano-Palomino, Mauro Dalla Mura
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, eess.IV
arXiv:2607.22352v1 Announce Type: cross Abstract: We introduce time-reversed imaging, a new paradigm that infers what just happened in a scene from fading multimodal traces. Instead of extrapolating or interpolating video frames, our goal is to infer past human-environment interactions from residual...
126. SiPhy: Single-Image Physical Property Reasoning ​
Author: Hoang Le, Joonwoo Kwon, Elkhan Ismayilzada, Yufei Zhang, Zijun Cui
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.22355v1 Announce Type: cross Abstract: Inferring physical properties such as mass, stiffness, and elasticity from a single image is essential for simulation and embodied AI, yet most existing approaches rely on multi-view reconstruction or physics-based supervision. We introduce SiPhy, a ...
127. Indexing: the Beginning and the End ​
Author: Alexander Kozachinskiy, Vicente Opazo, Felipe Urrutia
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22361v1 Announce Type: cross Abstract: We study information bottlenecks in modern deep-learning architectures -- RNNs, softmax transformers, linear-attention transformers and state-space models -- through the lens of the indexing primitive. In this primitive, the input consists of $n$ bit...
128. Interior interpretability with attention rollout: contraction and propagation profiles in Transformers ​
Author: Umberto Biccari, Qian Huang, Enrique Zuazua
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22367v1 Announce Type: cross Abstract: Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers. We introduce \emph{interior interpretab...
129. HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding ​
Author: Chao Fang, Jun Yin, Man Shi, Marian Verhelst
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.LG
arXiv:2607.22389v1 Announce Type: cross Abstract: With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. To tackle this challenge, we propose HiKV, a novel algorithm-hardware co-design that ex...
130. A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation ​
Author: Fin Gentzen, Marla Grunewald, Iulisloi Zacarias, Mounir Bensalem, Admela Jukan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.NI, cs.AI
arXiv:2607.22400v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents, transitioning from static conversational interfaces to dynamic systems capable of complex reasoning, tool execution, and decision-making. However, the operational reliabilit...
131. PRIMS: Physics-guided Representation for Fluid Identification in Multimodal Sensing ​
Author: Hai-Long Nguyen, Trung Thanh Nguyen, Lars Holm, Dennis Alveringh, Duc Viet Le
Published: 7/27/2026, 4:00:00 AM
Categories: physics.flu-dyn, cs.AI
arXiv:2607.22422v1 Announce Type: cross Abstract: Accurate on-device fluid identification is essential for microfluidic applications, yet maintaining reliability under varying flow, pressure, and temperature remains a key challenge. Existing learning-based methods often treat sensor signals as domai...
132. Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability ​
Author: Ahmed M. Abuzuraiq, Philippe Pasquier
Published: 7/27/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.LG, cs.MM
arXiv:2607.22428v1 Announce Type: cross Abstract: Explainable AI (XAI) in creative practice can be less about technocentric explanation and more about enabling artists to inspect modify and debug models as part of making Yet largescale texttoimage diffusion systems are typically presented as opaque ...
133. Robot Learning to Communicate through Projected Visual Abstractions ​
Author: Danyang Yan, Boyuan Wang, Jiaxun Liu, Boyuan Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.22434v1 Announce Type: cross Abstract: Humans routinely communicate through abstractions of their bodies, including shadows, silhouettes, and reflections. Yet robots remain largely confined to expressing themselves through their physical morphology. Enabling robots to communicate through ...
134. Hyperball May Not Be a Free Lunch ​
Author: Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22444v1 Announce Type: cross Abstract: For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing updates. However, the source of their advantage remains unclear. Start...
135. Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining ​
Author: V'ictor Rinc'on Yepes
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22458v1 Announce Type: cross Abstract: Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio models (AST, CLAP, BEATs-bio and BirdNET) with a downstream task none of them saw during training: recovering phylogenetic distance fro...
136. Beyond Perspectives: A Trio-Ethnography of Interpretation Evolution in LLM-Supported Programming Education ​
Author: Jennie Ren, Jordan H. McDowell, Kyrie Zhixuan Zhou
Published: 7/27/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2607.22463v1 Announce Type: cross Abstract: Generative AI is reshaping programming education, yet educators often infer students' AI-supported learning from classroom observations alone. This experience report presents a trio-ethnography involving two computing educators with different teachin...
137. Learning to Prepare Molecular Ground States with Transformer Models ​
Author: Alex Koziell-Pipe, Jasmine Brewer, Jem Guhit, Marwa H. Farag, Kripa Panchagnula, Gabriel Laude, Fabian Finger, Carlo Gaggioli, Ludmila Szulakowska, Oliver J. Backhouse, Christos Papalitsas, Jason G. Mustakis, Thomas Soini, David Munoz Ramo, Stephen Clark, Elica Kyoseva, Enrico Rinaldi
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.AI
arXiv:2607.22468v1 Announce Type: cross Abstract: Quantum state preparation is a key component of many quantum algorithms. Performing this step efficiently is essential for realizing practical quantum advantage in quantum chemistry applications. Iterative algorithms like ADAPT-VQE can produce shallo...
138. MineValiCoder: Reliable Code Generation with Test Case Quality Mining and Bipartite Graph-Based Mutual Validation ​
Author: Zhen Zhao, Qihang Yang, Feifei Dai, Xiangfang Li, Bo Li
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.22471v1 Announce Type: cross Abstract: Large Language Model (LLM)-based Test-Driven Development (TDD) has advanced automated code generation. However, existing approaches depend heavily on human-crafted test cases and cannot operate effectively when only natural-language requirements are ...
139. \k{appa}-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating ​
Author: Jianghui Wang, Silong Yong, Francesco Orabona, Marco Canini, Katia P. Sycara, Yaqi Xie
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22489v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) has become a widely adopted technique for efficient neural network fine-tuning, decomposing model updates into low-rank matrices. However, LoRA remains computationally costly because it updates all matrices uniformly, regar...
140. CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference ​
Author: Jiyuan Tan, Vasilis Syrgkanis
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, econ.EM
arXiv:2607.22511v1 Announce Type: cross Abstract: Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large language model (LLM) reviewer. However, such reviewers...
141. Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science ​
Author: Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL
arXiv:2607.22513v1 Announce Type: cross Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationa...
142. Quantum Spectral Model: Data Reuploading with Input-Conditioned Frequency Support ​
Author: Peiyong Wang, Udaya Parampalli, Casey R. Myers
Published: 7/27/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG
arXiv:2607.22516v1 Announce Type: cross Abstract: A central design principle in modern machine learning and artificial intelligence is to align a model's inductive bias with the structure of its input data. For matrix-valued inputs, relevant matrix-level relationships can be characterised through sp...
143. SM4RT: Learning Structured Motion Geometry for 4D Reconstruction ​
Author: Shing Ho J. Lin, Wenzhao Zheng, Dong Zhuo, Yuqi Wu, Jie Zhou, Jiwen Lu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO
arXiv:2607.22534v1 Announce Type: cross Abstract: Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding remains a fundamental challenge. Most existing motion perception methods (e.g., sparse tracking, dens...
144. Integrating Reasoning Systems for Trustworthy AI, Proceedings of the 4th Workshop on Logic and Practice of Programming (LPOP) ​
Author: Anil Nerode, Yanhong A. Liu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LO, cs.PL
arXiv:2410.19738v2 Announce Type: replace Abstract: This proceedings contains abstracts and position papers for the work to be presented at the fourth Logic and Practice of Programming (LPOP) Workshop. The workshop is to be held in Dallas, Texas, USA, and as a hybrid event, on October 13, 2024, in c...
145. Analyzing the Ethical Logic of Eight Large Language Models ​
Author: W. Russell Neuman, Chad Coleman, Manan Shah
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2501.08951v2 Announce Type: replace Abstract: This study examines the expressed ethical logic of eight prominent large language models from OpenAI, Meta, Perplexity, Anthropic, Google, Mistral, DeepSeek, and xAI. Each model answered direct questions about its ethical principles and responded t...
146. From Mind to Machine: The Rise of Manus AI as a Fully Autonomous Digital Agent ​
Author: Minjie Shen, Yanshu Li, Lulu Chen, Zhichao Fan, Yanhang Li, Qikai Yang, Haochen Yang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2505.02024v4 Announce Type: replace Abstract: Manus AI is a general-purpose AI agent introduced in early 2025, marking a significant advancement in autonomous artificial intelligence. Developed by the Chinese startup Monica.im, Manus is designed to bridge the gap between "mind" and "hand" - co...
147. Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning ​
Author: Renos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille, Wei Xia, Stefano Soatto
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2511.02130v2 Announce Type: replace Abstract: We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapter on reasoning models, demonst...
148. DeepFeature: LLM-Empowered Context-aware Feature Generation for Wearable Biosignals ​
Author: Kaiwei Liu, Yuting He, Bufang Yang, Mu Yuan, Chun Man Victor Wong, Ho Pong Andrew Sze, Zhenyu Yan, Hongkai Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2512.08379v2 Announce Type: replace Abstract: Biosignals collected from wearable devices are widely utilized in healthcare applications. Machine learning models used in these applications often rely on features extracted from biosignals due to their effectiveness, lower data dimensionality, an...
149. Beyond Text-to-SQL: Can LLMs Really Debug Enterprise ETL SQL? ​
Author: Jing Ye, Yiwen Duan, Yonghong Yu, Victor Ma, Yang Gao, Xing Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.18119v2 Announce Type: replace Abstract: SQL is central to enterprise data engineering, yet generating fully correct SQL code in a single attempt remains difficult, even for experienced developers and advanced text-to-SQL LLMs, often requiring multiple debugging iterations. We introduce O...
150. Statistical Early Stopping for Reasoning Models ​
Author: Yangxinyu Xie, Tao Wang, Soham Mallick, Yan Sun, Georgy Noarov, Mengxin Yu, Tanwi Mallick, Edgar Dobriban
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ML
arXiv:2602.13935v3 Announce Type: replace Abstract: While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given ill-posed or ambiguous queries. We introduce statistically principl...
151. Exploring Robust Multi-Agent Workflows for Environmental Data Management ​
Author: Boyuan Guan, Jason Liu, Yanzhao Wu, Kiavash Bahreini
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.01647v2 Announce Type: replace Abstract: Embedding LLM-driven agents into environmental FAIR data management is compelling - they can externalize operational knowledge and scale curation across heterogeneous data and evolving conventions. However, replacing deterministic components with p...
152. Same World, Differently Given: History-Dependent Perceptual Reorganization in Artificial Agents ​
Author: Hongju Pae
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.04637v2 Announce Type: replace Abstract: What kind of internal organization would allow an artificial agent not only to adapt its behavior, but to sustain a history-sensitive perspective on its world? I present a minimal architecture in which a slow perspective latent $g$ feeds back into ...
153. Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control ​
Author: Mahiro Nakao, Kazuhiro Takemoto
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.RO
arXiv:2604.26577v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants, yet their safety in this context remains poorly characterized. We introduce a dataset of 270 harmful instructions spannin...
154. Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs ​
Author: Naomi Esposito, Anthony Tricarico, Luisa Porzio, Ali Aghazadeh Ardebili, Massimo Stella
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.HC, cs.LG, cs.SI
arXiv:2604.27618v2 Announce Type: replace Abstract: Understanding the impact of large language models (LLMs) on mathematics education requires data on LLMs' mathematical performance and biases. To this end, we introduce Math Education Digital Shadows (MEDS), a dataset mapping how LLMs reason about m...
155. AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair ​
Author: Yuelin Hu, Zhenbo Yu, Zhengxue Cheng, Wei Liu, Li Song
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.SE
arXiv:2605.04624v2 Announce Type: replace Abstract: Agent-repair leaderboards reorder under evaluator reconfiguration, and a measurable share of the reordering is produced by methods that consult evaluator-derived signal during internal selection of candidate repairs. We document this failure mode o...
156. A Statistical Multi-Objective Framework for Assessing Sensitivity of Radiomic AI Models to Acquisition Parameters ​
Author: D. Gil, I. Sanchez, C. Sanchez
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.14667v2 Announce Type: replace Abstract: A main barrier for the deployment of AI radiomic systems in clinical routine is their drop in performance under heterogeneous multicentre acquisition protocols. This work presents a performance-oriented framework for quantifying scan parameter sens...
157. Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models ​
Author: Junyao Yang, Chen Qian, Kun Wang, Linfeng Zhang, Quanshi Zhang, Yong Liu, Dongrui Liu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2605.17770v4 Announce Type: replace Abstract: The advancement of Large Reasoning Models (LRMs) has catalyzed a paradigm shift from reactive fast thinking'' text generation to systematic, step-by-step slow thinking'' reasoning, unlocking state-of-the-art performance in complex mathematical ...
158. Universal Quantum Transformer ​
Author: Sungyong Chung, Alireza Talebpour
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, quant-ph
arXiv:2606.00045v2 Announce Type: replace Abstract: Classical continuous-space neural networks fundamentally struggle to lock into exact formal rules, whether mathematical, such as modular arithmetic and non-Abelian group algebra, or linguistic, such as systematic compositional generalization. To ap...
159. Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems ​
Author: Bingnan Xiao, Chenhao Yang, Wei Ni, Xin Wang, Tony Q. S. Quek
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.24416v2 Announce Type: replace Abstract: Network operators' changing policies, service requirements, and stringent real-time constraints render existing methods designed with fixed objectives and constraints ineffective. This paper presents Agentic long-term performance optimization (Agen...
160. AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets ​
Author: Ming Chen, Pranav Pai
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, cs.MA
arXiv:2607.15781v2 Announce Type: replace Abstract: Geospatial datasets support applications from urban planning to climate modeling, yet consistent assessment of FAIR compliance is difficult. Existing evaluators use different rubrics and evidence sources and may fail on JavaScript-rendered pages or...
161. When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering ​
Author: Ziteng Hu, Jiachi Chen, Wenhao Lv, Huan Zhang, Yingjie Xia
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.17063v2 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has led practitioners to increasingly rely on them for answering questions about hardware description languages (HDLs). Because HDL is ultimately synthesized into physical hardware, an imprecise...
162. Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering ​
Author: Chendi Wang, Liam Cunningham, Tom Yishay, Jieying Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.20487v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to answer questions about political information, including in election-adjacent information settings where factual errors and ideological distortions are high-stakes. We present a reproducible meas...
163. DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making ​
Author: Raffi Khatchadourian
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2607.20491v2 Announce Type: replace Abstract: A financial AI agent can repeat a decision while changing the tools, order, or recorded arguments and results used to reach it. Outcome-only evaluation misses this variation, even when it matters for replay and change control. DFAH-Bench operationa...
164. KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback ​
Author: Yan Zhu, Y. Chen, Rebecca Faust
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.20556v2 Announce Type: replace Abstract: In large-scale text analysis tasks, pre-trained language models are often used to embed text corpora for downstream analysis. However, such models may struggle to capture domain-specific semantics and adapting them typically requires large amounts ...
165. OPOD: On-Policy Omni Distillation ​
Author: Tong Zhao, Yuyang Hu, Yutao Zhu, Reed Li, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.20918v2 Announce Type: replace Abstract: Omni-modal models can handle text, images, and audio in one system, but improving all of these abilities together remains difficult. Training a single model on pooled multimodal data often fails to match models specialized for individual modalities...
166. AREX: Towards a Recursively Self-Improving Agent for Deep Research ​
Author: Shuqi Lu, Chaofan Li, Kun Luo, Zhang Zhang, Hui Wang, Hongwang Xiao, Lei Xiong, Jiahao Wang, Sen Wang, Xiyan Jiang, Wanli Li, Yuyang Hu, Hongjin Qian, Bingyu Yan, Jianlyu Chen, Ziyi Xia, Yingxia Shao, Kang Liu, Zhicheng Dou, Di He, Chaozhuo Li, Qiwei Ye, Zhongyuan Wang, Zheng Liu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.21461v2 Announce Type: replace Abstract: Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification a...
167. Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks ​
Author: Mack Nixon, Liam Wright, Yevgeniya Kovalchuk, Alison Fang-Wei Wu, Martin Danka, Andy Boyd, David Bann
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.21482v2 Announce Type: replace Abstract: Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governance requirements that typical...
168. OpenForgeRL: Train Harness-native Agents in Any Environment ​
Author: Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou, Qianhui Wu, Hao Cheng, Wenlin Yao, Nikhil Singh, Zhou Yu, Jianfeng Gao
Published: 7/27/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2607.21557v2 Announce Type: replace Abstract: Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-e...
169. Participatory Budgeting with Project Groups ​
Author: Pallavi Jain, Krzysztof Sornat, Nimrod Talmon, Meirav Zehavi
Published: 7/27/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.DS, cs.MA
arXiv:2012.05213v2 Announce Type: replace-cross Abstract: We study a generalization of the standard approval-based model of participatory budgeting (PB), in which voters are providing approval ballots over a set of predefined projects and---in addition to a global budget limit, there are several gro...
170. Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning ​
Author: Seyeon Kim, Joonhun Lee, Namhoon Cho, Sungjun Han, Wooseop Hwang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.PR, stat.ML
arXiv:2408.02295v4 Announce Type: replace-cross Abstract: Conventional uncertainty-aware temporal difference (TD) learning often models TD errors as zero-mean Gaussian. This assumption can miss the heavy-tailed and heteroscedastic residuals induced by bootstrapping and exploration. We introduce a st...
171. Carpe Diem: Critical Learning Period-Aware Contract-Based Incentives for Federated Learning ​
Author: Thanh Linh Nguyen, Dinh Thai Hoang, Diep N. Nguyen, Quoc-Viet Pham
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.GT
arXiv:2503.07869v4 Announce Type: replace-cross Abstract: Critical learning periods (CLPs) in federated learning (FL) refer to early stages during which low-quality contributions (e.g., sparse training data availability) can permanently impair the performance of the global model. However, existing i...
172. When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas ​
Author: Steffen Backmann, David Guzman Piedrahita, Terry Jingchen Zhang, Emanuel Tewolde, Rada Mihalcea, Bernhard Sch"olkopf, Zhijing Jin
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2505.19212v2 Announce Type: replace-cross Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a critical concern. While prior work has examined LLMs' moral judgment and strategic beha...
173. Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning ​
Author: Wu Fei, Shuxian Liang, Yibo Yang, Yang Lin, Jing Tang, Lei Chen, Xiansheng Hua, Hao Kong
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2507.01551v3 Announce Type: replace-cross Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~(LLMs). However, introducing additional process reward models incurs substantial computational overh...
174. A Robust Pipeline for Differentially Private Federated Learning on Imbalanced Clinical Data using SMOTETomek and FedProx ​
Author: Rodrigo Tertulino
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, cs.SE
arXiv:2508.10017v2 Announce Type: replace-cross Abstract: Federated Learning (FL) presents a groundbreaking approach for collaborative health research, allowing model training on decentralized data while safeguarding patient privacy. FL offers formal security guarantees when combined with Differenti...
175. MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph ​
Author: Duzhen Zhang, Zixiao Wang, Zhong-Zhi Li, Yahan Yu, Shuncheng Jia, Jiahua Dong, Haotian Xu, Xing Wu, Yingying Zhang, Tielin Zhang, Jie Yang, Xiuying Chen, Le Song
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2508.12393v3 Announce Type: replace-cross Abstract: The rapid expansion of medical literature challenges the scalable structuring of domain knowledge. Knowledge Graphs (KGs) offer a solution, yet current construction methods lack generalizability and ignore the temporal dynamics of evolving kn...
176. An Ontology-Based Approach to Optimizing Geometry Problem Sets for Skill Development ​
Author: Michael Bouzinier, Sergey Trifonov, Matthew Chen, Tarun Venkatesh, Lielle Rifkin
Published: 7/27/2026, 4:00:00 AM
Categories: math.HO, cs.AI
arXiv:2509.02758v3 Announce Type: replace-cross Abstract: Euclidean geometry has historically played a central role in cultivating logical reasoning and abstract thinking within mathematics education, but has experienced waning emphasis in recent curricula. The resurgence of interest, driven by adva...
177. From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem ​
Author: James Jewitt, Hao Li, Bram Adams, Gopi Krishnan Rajbahadur, Ahmed E. Hassan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2509.09873v2 Announce Type: replace-cross Abstract: Hidden license conflicts in the open-source AI ecosystem pose serious legal and ethical risks, exposing organizations to potential litigation and users to undisclosed risk. However, the field lacks a data-driven understanding of how frequentl...
178. A Comparative Benchmark of Federated Learning Strategies for Mortality Prediction on Heterogeneous and Imbalanced Clinical Data ​
Author: Rodrigo Tertulino
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY
arXiv:2509.10517v3 Announce Type: replace-cross Abstract: Machine learning can predict in-hospital mortality, but data privacy and the statistical heterogeneity of clinical data hamper its use. Federated Learning (FL) is privacy-preserving, yet its behavior under non-IID and imbalanced conditions ne...
179. SurvDiff: A Diffusion Model for Generating Synthetic Data in Survival Analysis ​
Author: Marie Brockschmidt, Maresa Schr"oder, Stefan Feuerriegel
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2509.22352v3 Announce Type: replace-cross Abstract: Survival analysis is a cornerstone of clinical research by modeling time-to-event outcomes such as metastasis, disease relapse, or patient death. Unlike standard tabular data, survival data often come with incomplete event information due to ...
180. Vector-Valued Reproducing Kernel Banach Spaces for Neural Networks and Operators ​
Author: Sven Dummer, Tjeerd Jan Heeringa, Jos'e A. Iglesias
Published: 7/27/2026, 4:00:00 AM
Categories: math.FA, cs.AI, cs.LG, stat.ML
arXiv:2509.26371v3 Announce Type: replace-cross Abstract: Recently, there has been growing interest in characterizing the function spaces underlying neural networks. While shallow and deep scalar-valued neural networks have been linked to scalar-valued reproducing kernel Banach spaces (RKBS), $\math...
181. Wasserstein Gradient Flows for Scalable and Regularized Barycenter Computation ​
Author: Eduardo Fernandes Montesuma, Yassir Bendou, Mike Gartrell
Published: 7/27/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2510.04602v4 Announce Type: replace-cross Abstract: Wasserstein barycenters provide a principled approach for aggregating probability measures, while preserving the geometry of their ambient space. Existing discrete methods are not because as they assume access to the complete set of samples f...
182. InteractComp: Evaluating Search Agents With Ambiguous Queries ​
Author: Mingyi Deng, Lijun Huang, Yani Fan, Fanqi Kong, Jiayi Zhang, Fashen Ren, Jinyi Bai, Fuzhen Yang, Dayi Miao, Zhaoyang Yu, Yifan Wu, Yanfei Zhang, Fengwei Teng, Yingjia Wan, Song Hu, Yude Li, Xin Jin, Conghao Hu, Haoyu Li, Qirui Fu, Tai Zhong, Xinyu Wang, Xiangru Tang, Nan Tang, Chenglin Wu, Yuyu Luo
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.24668v2 Announce Type: replace-cross Abstract: Language agents have demonstrated remarkable potential in web search and information retrieval. However, many search-agent benchmarks assume that user queries are complete and unambiguous. This assumption leaves under-tested a practical failu...
183. A Theoretical Framework for Environmental Similarity and Vessel Mobility as Coupled Predictors of Marine Invasive Species Pathways ​
Author: Gabriel Spadon, Vaishnav Vaidheeswaran, Claudio DiBacco
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CE, cs.AI
arXiv:2511.03499v3 Announce Type: replace-cross Abstract: Marine invasive species spread through global shipping and generate substantial ecological and economic impacts. Traditional risk assessments require detailed records of ballast water and traffic patterns, which are often incomplete, limiting...
184. CrypTorch: PyTorch-based Auto-tuning Compiler for Machine Learning with Multi-party Computation ​
Author: Jinyu Liu, Gang Tan, Kiwan Maeng
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.PL
arXiv:2511.19711v2 Announce Type: replace-cross Abstract: MPC-based ML uses multi-party computation (MPC) to run machine learning (ML) workloads across multiple parties without each having to share their private data or model parameters. However, existing frameworks frequently degrade accuracy and p...
185. Security Without Detection: Economic Denial as a Primitive for Edge and IoT Defense ​
Author: Samaresh Kumar Singh, Joyjit Roy, Sriharsha Anand Pushkala
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.DC, cs.LG
arXiv:2512.23849v2 Announce Type: replace-cross Abstract: Sophisticated attackers can evade detection-based security by using encryption, stealth tactics, and low-rate attack patterns. This challenge is particularly acute in Internet of Things (IoT) and edge environments, where limited resources mak...
186. CoCo-Fed: A Unified Framework for Memory- and Communication-Efficient Federated Learning at the Wireless Edge ​
Author: Zhiheng Guo, Zhaoyang Liu, Zihan Cen, Chenyuan Feng, Xinghua Sun, Xiang Chen, Tony Q. S. Quek, Xijun Wang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, math.IT
arXiv:2601.00549v3 Announce Type: replace-cross Abstract: The deployment of large-scale neural networks within the Open Radio Access Network (O-RAN) architecture is pivotal for enabling native edge intelligence. However, this paradigm faces two critical bottlenecks: the prohibitive memory footprint ...
187. Atlas 2 -- Foundation models for clinical deployment ​
Author: Maximilian Alber, Timo Milbich, Alexandra Carpen-Amarie, Stephan Tietz, Jonas Dippel, Lukas Muttenthaler, Beatriz Perez Cancer, Alessandro Benetti, Panos Korfiatis, Elias Eulig, J'er^ome L"uscher, Jiasen Wu, Sayed Abid Hashimi, Gabriel Dernbach, Simon Schallenberg, Neelay Shah, Moritz Kr"ugener, Aniruddh Jammoria, Jake Matras, Patrick Duffy, Matt Redlon, Philipp Jurmeister, David Horst, Lukas Ruff, Klaus-Robert M"uller, Frederick Klauschen, Andrew Norgan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2601.05148v2 Announce Type: replace-cross Abstract: Pathology foundation models substantially advanced the possibilities in computational pathology --- yet tradeoffs in terms of performance, robustness, and computational requirements remained, which limited their clinical deployment. In this r...
188. GPU-Accelerated ANNS: Quantized for Speed, Built for Change ​
Author: Hunter McCoy, Zikun Wang, Prashant Pandey
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2601.07048v4 Announce Type: replace-cross Abstract: Approximate nearest neighbor search (ANNS) is a core problem in machine learning and information retrieval applications. GPUs offer a promising path to high-performance ANNS: they provide massive parallelism for distance computations, are rea...
189. SwiftMem: Fast Agentic Memory via Query-aware Indexing ​
Author: Anxin Tian, Yiming Li, Xing Li, Hui-Ling Zhen, Lei Chen, Xianzhi Yu, Zhenhua Dong, Mingxuan Yuan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.08160v2 Announce Type: replace-cross Abstract: Agentic memory systems have become critical for enabling LLM agents to maintain long-term context and retrieve relevant information efficiently. However, existing memory frameworks often perform query-agnostic retrieval over the full memory e...
190. Decentralized Multi-Agent Swarms for Autonomous Grid Security in Industrial IoT: A Consensus-based Approach ​
Author: Samaresh Kumar Singh, Joyjit Roy, Chirag Agrawal
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.ET
arXiv:2601.17303v2 Announce Type: replace-cross Abstract: As Industrial Internet of Things (IIoT) environments scale to tens of thousands of connected devices, centralized security architectures introduce latency bottlenecks that sophisticated attackers can exploit to compromise an entire manufactur...
191. Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning ​
Author: Muhammad Ahmed Atif, Nehal Naeem Haji, Mohammad Shahid Shaikh, Muhammad Ebad Atif
Published: 7/27/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.LG
arXiv:2601.17454v2 Announce Type: replace-cross Abstract: Centralized value learning underlies a broad class of multi-agent reinforcement learning methods, but its claimed advantage is typically evaluated in settings that confound coordination structure with function approximation and partial observ...
192. Automatic Stability and Recovery for Neural Network Training ​
Author: Barak Or
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.17483v2 Announce Type: replace-cross Abstract: Training modern neural networks is increasingly fragile, with rare but severe destabilizing updates often causing irreversible divergence or silent performance degradation. Existing optimization methods primarily rely on preventive mechanisms...
193. The pretraining domain outweighs the training objective in setting the privacy-utility trade-off of differentially private medical image analysis ​
Author: Soroosh Tayebi Arasteh, Mina Farajiamiri, Mahshad Lotfinia, Behrus Hinrichs-Puladi, Jonas Bienzeisler, Mohamed Alhaskir, Mirabela Rusu, Christiane Kuhl, Sven Nebelung, Daniel Truhn
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2601.19618v2 Announce Type: replace-cross Abstract: Differential privacy protects the patients whose images train medical imaging models, but it lowers diagnostic accuracy, and the initialization is the strongest known remedy. Practice increasingly favors large generic self-supervised encoders...
194. Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems ​
Author: Samaresh Kumar Singh, Joyjit Roy
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.SE
arXiv:2602.04120v4 Announce Type: replace-cross Abstract: Though Explainable AI (XAI) has made significant advancements, its inclusion in edge and IoT systems is typically ad-hoc and inefficient. Most current methods are "coupled" in such a way that they generate explanations simultaneously with mod...
195. Predictive Query Language: A Domain-Specific Language for Predictive Modeling on Relational Databases ​
Author: Vid Kocijan, Jinu Sunil, Jan Eric Lenssen, Viman Deb, Xinwei Xe, Federico Reyes Gomez, Matthias Fey, Jure Leskovec
Published: 7/27/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.LG
arXiv:2602.09572v3 Announce Type: replace-cross Abstract: The purpose of predictive modeling on relational data is to predict future or missing values in a relational database, for example, future purchases of a user, risk of readmission of the patient, or the likelihood that a financial transaction...
196. Administrative Law's Fourth Settlement: AI and the Scrutable State ​
Author: Nicholas Caputo
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2602.09678v3 Announce Type: replace-cross Abstract: Since 1887, administrative law has confronted a problem of institutional cognition. Expert agencies are needed to govern technologically complex systems, but expertise makes agency decisions difficult for courts, Congress, and the public to u...
197. WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics ​
Author: Sneha Maurya, Spandana Govindgari, Girish Kumar, Akhara AI
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2604.00024v2 Announce Type: replace-cross Abstract: Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Women's Health Benchmark (WHBench), a targeted evaluation suite of 47 expert-crafted scenarios ac...
198. Energy-based Tissue Manifolds for Longitudinal Multiparametric MRI Analysis ​
Author: Kartikay Tehlan, Lukas F"orner, Sina Wendrich, Nico Schmutzenhofer, Michael Fr"uhwald, Matthias Wagner, Nassir Navab, Thomas Wendler
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.07180v3 Announce Type: replace-cross Abstract: We propose a geometric framework for longitudinal multi-parametric MRI analysis based on patient-specific energy modelling in sequence space. Rather than operating on images with spatial networks, each voxel is represented by its multi-sequen...
199. Hidden Failure Modes of Gradient Modification under Adam in Continual Learning, and Adaptive Decoupled Moment Routing as a Repair ​
Author: Yuelin Hu, Zhenbo Yu, Zhengxue Cheng, Wei Liu, Li Song
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.22407v2 Announce Type: replace-cross Abstract: Many continual-learning methods modify gradients upstream (e.g., projection, penalty rescaling, replay mixing) while treating Adam as a neutral backend. We show this composition has a hidden failure mode. In a high-overlap, non-adaptive 8-dom...
200. Simpson's Paradox in Behavioral Curves: How Aggregation Distorts Parametric Models of User Dynamics ​
Author: Chao Zhou
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR
arXiv:2605.11017v2 Announce Type: replace-cross Abstract: Behavioral curve modeling -- fitting parametric functions to engagement-versus-exposure data -- is standard practice in recommendation, advertising, and clinical dosing. We show that aggregation introduces a systematic distortion: Simpson's p...
201. DriftXpress: Faster Drifting Models via Projected RKHS Fields ​
Author: Ali Falahati, Elliot Creager, Gautam Kamath, Shubhankar Mohapatra
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.12183v2 Announce Type: replace-cross Abstract: Drifting Models have emerged as a new paradigm for one-step generative modeling, achieving strong image quality without iterative inference. The premise is to replace the iterative denoising process in diffusion models with a single evaluatio...
202. Short-Term-to-Long-Term Memory Transfer for Knowledge Graphs under Partial Observability ​
Author: Taewoon Kim, Vincent Fran\c{c}ois-Lavet, Michael Cochez
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.22142v3 Announce Type: replace-cross Abstract: Reinforcement learning under partial observability requires deciding what information to retain, yet most memory-based approaches do not explicitly model short-term-to-long-term transfer of symbolic observations. We study this transfer proces...
203. ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents ​
Author: Jianan Ma, Xiaohu Du, Ruixiao Lin, Yaoxiang Bian, Jialuo Chen, Yunhao Feng, Xiaofang Yang, Shiwen Cui, Changhua Meng, Xinhao Deng, Jingyi Wang, Zhen Wang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE
arXiv:2605.22321v2 Announce Type: replace-cross Abstract: As autonomous agents (e.g., OpenClaw) increasingly operate with deep system-level privileges to execute complex tasks, they introduce severe, unmitigated security risks. Existing LLM safety testing methods are largely built around prompt-leve...
204. Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO ​
Author: Yiran Xu, Yiming Ren, Zicheng Lin, Chufan Shi, Yukang Chen, Dingdong Wang, Tianhe Wu, Junjie Wang, Yujiu Yang, Yu Qiao, Ruihang Chu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.30789v3 Announce Type: replace-cross Abstract: We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies primarily increase diversity by injecting more token-level randomn...
205. NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation ​
Author: NVIDIA, :, Aarti Basant, Amlan Kar, Despoina Paschalidou, Fangyin Wei, Francesco Ferroni, Guillermo Garcia Cobo, Haithem Turki, Huan Ling, Jaewoo Seo, James Lucas, Jay Zhangjie Wu, Jialiang Wang, Jonathan Lorraine, Jun Gao, Kai He, Katarina Tothova, Kevin Xie, Micha{\l} Tyszkiewicz, Qi Wu, Riccardo de Lutio, Ruilong Li, Sanja Fidler, Seung Wook Kim, Tianchang Shen, Tianshi Cao, Tobias Pfaff, William Lew, Xindi Wu, Xuanchi Ren, Yifan Lu, Yuxuan Zhang, Zan Gojcic, Zian Wang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO
arXiv:2606.03159v2 Announce Type: replace-cross Abstract: As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-loop simulation, the driving policy model actively interacts with the environment, where its a...
206. Pretraining Recurrent Networks without Recurrence ​
Author: Akarsh Kumar, Phillip Isola
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.06479v2 Announce Type: replace-cross Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations. Standard backpropagation through time (BPTT) addresses this problem poorly: it is sequential in time, limiting parallelism, and suffers...
207. BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression ​
Author: Shaohao Rui, Xiaofeng Mao, Zhanyu Zhang, Peijia Lin, Yansong Zhu, Yibo Zhang, Haibin Wan, Zhangrui Zhao, Weijie Ma
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2606.10135v4 Announce Type: replace-cross Abstract: Interactive video world models commonly convert bidirectional video generators into causal autoregressive systems through control fine-tuning, autoregressive training, causal initialization, and few-step distillation. This pipeline is costly,...
208. RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation ​
Author: Renzhi Wu, Zikun Cui, Junjie Yang, Tai Guo, Hong Li, Xian Chen, Li Yu, Ke Pan, Sri Reddy, Mahesh Srinivasan, Nipun Mathur, Haomin Yu, Hong Yan
Published: 7/27/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2606.18379v4 Announce Type: replace-cross Abstract: Graph-based retrieval at billion-node scale requires jointly solving three tightly coupled problems -- graph construction, representation learning, and real-time serving -- yet existing work addresses each in isolation. We present RankGraph-2...
209. An Empirical Study of OpenPangu Quantization on Ascend NPUs ​
Author: Tong Shi, Jiacheng Wang, Hui Xie, Ying Li, Aishan Liu, Jinyang Guo, Xianglong Liu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.21257v3 Announce Type: replace-cross Abstract: OpenPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive post-training quantization on Ascend NPUs has not been systematically characterized. This paper conducts a...
210. Gemma 4 Technical Report ​
Author: Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor C\u{a}rbune, Michelle Casbon, Mayank Chaturvedi, Aditya Chawla, Victor Cotruta, Alice Coucke, Phil Culliton, Robert Dadashi, Lucas Dixon, Mohamed Elhawaty, Utku Evci, Cl'ement Farabet, Johan Ferret, Filippo Galgani, Sertan Girgin, Jean-Bastien Grill, Maarten Grootendorst, Jiaxian Guo, Cassidy Hardin, Yanzhang He, Steven M. Hernandez, Omri Homburger, L'eonard Hussenot, Juyeong Ji, Armand Joulin, Aishwarya Kamath, Parnian Kassraie, Olivier Lacombe, Preethi Lahoti, Ga"el Liu, Gus Martins, Luciano Martins, Tatiana Matejovicova, Ramona Merhej, Nikola Momchev, Sneha Mondal, Ryan Mullins, Sindhu Raghuram Panyam, Shreya Pathak, Sarah Perrin, Andr'e Susano Pinto, Etienne Pot, Ang'eline Pouget, Alexandre Ram'e, Sabela Ramos, Douglas Reid, David Rim, Morgane Rivi`ere, Karsten Roth, Louis Rouillard, Omar Sanseviero, Pier Giuseppe Sessa, Shane Settle, Danila Sinopalnikov, Sara Smoot, Piotr Stanczyk, Andreas Steiner, Lawrence Stewart, Ilya Tolstikhin, Michael Tschannen, Anton Tsitsulin, Nino Vieillard, Renjie Wu, Pingmei Xu, Haichuan Yang, Edouard Yvinec, Biao Zhang, Li Zhang, Joe Zou, Nicolas Aagnes, Abdelrahman Abdelhamed, Jakub Adamek, Shivani Agrawal, Shubham Agrawal, Ibrahim Alabdulmohsin, Jean Baptiste Alayrac, Uri Alon, Chandramouli Amarnath, Ankesh Anand, Chrysovalantis Anastasiou, Setareh Ariafar, Fran\c{c}ois-Xavier Aubet, Kyriakos Axiotis, Federico Barbero, Joelle Barral, Alexei Bendebury, Urs Bergmann, Stanley Bileschi, Kat Black, Mathieu Blondel, Sebastian Borgeaud, Arthur Bra\v{z}inskas, Ryan Burnell, Robert Busa-Fekete, Mu Cai, Daniele Calandriello, Glenn Cameron, Charlotte Caucheteux, Rahma Chaabouni, Garima Chadha, Jetha Chan, Blake Jianhang Chen, Jesse Chen, Lin Chen, Xu Chen, Derek Cheng, Tzu-hsiang Chien, Nikolai Chinaev, Yi Chou, Zhaohui Chu, Benjamin Coleman, Pooja Consul, Sam Conway-Rahman, Scott Crowell, Dylan Cutler, Vivek Dani, Samira Daruki, Anil Das, Daniel Deutsch, Nishanth Dikkala, Li Ding, Qiuhan Ding, Shenil Dodhia, Konstantin Donhauser, Tulsee Doshi, Anca Dragan, Alex Druinsky, Sahil Dua, Zoltan Egyed, Danielle Eisenbud, Daniel Eppens, Cindy Fan, Bahare Fatemi, Yassir Fathullah, Vlad Feinberg, Milen Ferev, Sebastian Flennerhag, Takumi Fujimoto, Jo~ao Gabriel Oliveira, Isaac Galatzer-Levy, Jo~ao Gante, Simon Geisler, Soham Ghosal, Antonious M. Girgis, Tamara von Glehn, Alec Go, Alhaad Gokhale, Alex Grills, Yiming Gu, Mayank Gupta, Pramod Gupta, Guru Guruganesh, Raia Hadsell, Hamza Harkous, Jitendra Harlalka, Demis Hassabis, Anja Hauth, Joe Heyward, Arian Hosseini, Chih-Yang Hsia, I-Hung Hsu, Xiaopeng Huang, Yangsibo Huang, Kevin Hui, Adrian Hutter, Te I, Fotis Iliopoulos, Advait Jain, Ganesh Jawahar, Ziwei Ji, Qilin Jin, Melvin Johnson, Kandarp Joshi, Arun Kandoor, Wang-Cheng Kang, Koray Kavukcuoglu, Mehran Kazemi, Kathleen Kenealy, Amr Khalifa, Phoebe Kirk, Ivan Korotkov, Suraj Kothawade, Vitaly Kovalev, Neel Kovelamudi, Adam Kraft, Ravin Kumar, Vivek Kumar, Harish Kuppam, Justin Lannin, Chen-Yu Lee, Seungji Lee, Dmitry Lepikhin, Alon Levkovitch, Dongdong Li, Qiujia Li, Valentin Li'evin, Ethan Lin, Ziqian Lin, Casper Liu, Tianlin Liu, Tianqi Liu, Xin Liu, Ivan Lobov, Mayank Lunayach, Min Ma, Gagan Madan, Andrii Maksai, Eric Malmi, Michal Matuszak, Daniel McDuff, Gaurav Menghani, Maciej Miku{\l}a, Daniil Mirylenka, Karolis Misiunas, Vedant Misra, Andreea Mitran, Kareem Mohamed, Maksim Mukha, Eric Noland, James O'Donnell, Brendan O'Donoghue, Kate Olszewska, Bernett Orlando, Wanqiong Pan, Rina Panigrahy, Unnati Parekh, Nicolas Perez-Nieves, Chunjong Park, Eric Paskie, Liqian Peng, Bryce Petrini, Slav Petrov, Jonas Pfeiffer, Bilal Piot, Martyna Plomecka, Siim Poder, Octavio Ponce, Arijit Pramanik, David Racz, Anish Rajan, Michelle Ramanovich, Anand Rao, Marvin Ritter, Vitor Rodrigues, Evan Rosen, Miko{\l}aj Rybi'nski, Noveen Sachdeva, Micha"el E. Sander, Rohit Sathyanarayana, Sagar Savla, Samuel Schmidgall, Tal Schuster, George Scrivener, Benoit Seguin, Andrew Sellergren, Aliaksei Severyn, Izhak Shafran, Dhruv Shah, Bobak Shahriari, Yuan Shangguan, Ashish Shenoy, Pradeep Shenoy, Rakesh Shivanna, Pauline Sho, Lucas Spangher, Wojciech Stokowiec, Tim Strother, Yao Su, Yinghao Sun, Mukund Sundararajan, Andrea Tacchetti, Mor Hazan Taege, Pouya Tafti, Jean Tarbouriech, Chetan Tekur, Shantanu Thakoor, Rahul Thapa, Madeleine Traverse, Lenart Treven, Tao Tu, Chien Te Tung, \c{C}a\u{g}lar "Unl"u, Petar Veli\v{c}kovi'c, Malini Pooni Venkat, Sagar Gubbi Venkatesh, Vidya Venkiteswaran, Francesco Visin, Alex Vitvitskyi, Kiran Vodrahalli, Weiyi Wang, Xin Wang, Tris Warkentin, Jan Wassenberg, John Wieting, Cindy Wu, Lechao Xiao, Hao Xu, Yuhui Xu, Fuzhao Xue, Arun Yadav, Jun Yan, Antoine Yang, Lin Yang, Ming-Hsuan Yang, Ziyu Ying, Jae Hyeon Yoo, Morteza Zadimoghaddam, Sajjad Zafar, Fred Zhang, Jiageng Zhang, Jianyi Zhang, Xiaofan Zhang, Chao Zhao, David Zhou, Chen Zou
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.02770v2 Announce Type: replace-cross Abstract: We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architec...
211. REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming ​
Author: Nicolas Koller, Andreas U. Schmidt
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CR, cs.PL
arXiv:2607.07738v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly applied to reverse-engineering tasks, and recent threat-intelligence reporting shows them operating inside live offensive-security workflows. Claims about their capability, however, outpace our ab...
212. The Computational Basis of Confidence in Large Language Models ​
Author: Dharshan Kumaran, Viorica Patraucean, Maks Ovsjanikov, Petar Veli\v{c}kovi'c, Nathaniel Daw
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.12447v2 Announce Type: replace-cross Abstract: Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it i...
213. NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs ​
Author: Jiarong Zhao, Zhikai Lei, Zhiheng Xi, Rui Zheng, Hang Yan, Jie Zhou, Qin Chen, Liang He
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2607.14186v5 Announce Type: replace-cross Abstract: Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, eac...
214. VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory ​
Author: Xiaoran Xu, Yupeng Wu, Tianyu Xue, Yifan Xu, Xuanran Dong, Xiaoshan Yang, Changsheng Xu
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.14514v2 Announce Type: replace-cross Abstract: Training-free ObjectNav agents increasingly use vision-language models (VLMs), yet typically discard acquired scene knowledge after each request. We study cross-episode ObjectNav, where each request is an independently initialized, single-goa...
215. It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability ​
Author: Carson Rodrigues
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.16292v3 Announce Type: replace-cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge. We ask whether their predicted responses, obtained with no scanner, are a useful feature lens for a downstream ...
216. DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines ​
Author: Runming He, Zhen Hao Wong, Hao Liang, Zimo Meng, Chengyu Shen, Xiaochen Ma, Wentao Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2607.16617v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this disconnect ...
217. Reliability Scales Inversely: Bigger Language Models Compound Mistakes Faster ​
Author: Kushal Chakrabarti
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.18292v2 Announce Type: replace-cross Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account -- more data, retrieval, or scale -- misses an auto-regressive risk residual that increases with scale...
218. Now We Know? A Systematic Comparison of TerraMind and THOR ​
Author: Frederick Schindlegger, Kenzo Bounegta, Eva Gmelich Meijling, Johannes Jakubik, Arnt-B{\o}rre Salberg, Theodor Forgaard, Nicolas Longepe, Valerio Marsocci
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2607.18504v2 Announce Type: replace-cross Abstract: Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much of the gap is architecture, how much is decoder capacity, and how much is a use-case-specifi...
219. Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents ​
Author: Guanxiong Chen, Qianjun Xia, Jiawei Peng, Heng Zhang, Bole Ma, Justin Qian, Ziyi Jiao, Bingyang Zhou, Luoxin Ye, Kaifeng Zhang, Kunyi Wang, Weijia Zeng, Yunuo Chen, Pengzhi Yang, Ziqiu Zeng, Siyuan Luo, Huamin Wang, Chao Liu, Alan Yuille, Fan Shi, Changxi Zheng, Yunzhu Li, Chenfanfu Jiang, Peter Yichen Chen
Published: 7/27/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2607.19190v3 Announce Type: replace-cross Abstract: Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters...
220. Post-Training in Time Series Foundation Models: A Unifying Framework ​
Author: Shifeng Xie, Ambroise Odonnat, Zehao Xiao, Lei Zan, Malik Tiomoko, Lujia Pan, Themis Palpanas, Boris N. Oreshkin, Chenghao Liu, Keli Zhang
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.20002v2 Announce Type: replace-cross Abstract: Time series foundation models (TSFMs) have emerged as general-purpose models for time series analysis, but pretraining alone is often insufficient for reliable downstream deployment. Bridging this gap requires further intervention to handle d...
221. AI-Driven Surrogate Models for Predicting Electrode-Scale Discharge Behavior in Lithium-Ion Batteries ​
Author: Mengda Xing (CRIL, UA), Jean-Marie Lagniez (CRIL, UA), Alejandro Franco (LRCS)
Published: 7/27/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.20577v2 Announce Type: replace-cross Abstract: Physics-based simulations are essential for understanding the electrode-scale discharge behavior of lithium-ion batteries (LIBs) but suffer from prohibitive computational costs. To address this, we introduce a novel deep learning surrogate pi...
222. The Geometry of Personality: Activation Steering with Jungian Cognitive Functions ​
Author: Liu Zai (University of Glasgow), Yumeng Wang (Leiden University), Junchen Fu (University of Glasgow), Joemon M. Jose (University of Glasgow)
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.20803v2 Announce Type: replace-cross Abstract: Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as the Big Five. We investigate whether personality can instead be represented and controlled ...
223. One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies ​
Author: Minh Ngoc Ta, My Anh Tran Nguyen, Duong D. Nguyen, Yuxia Wang, Preslav Nakov
Published: 7/27/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.21143v2 Announce Type: replace-cross Abstract: Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to ask, when to stop, and when to answer. We introduce RegretBench, a multi-turn benchmark that ...