Skip to content

arXiv cs.AI - 2026-07-28 ​

587 items collected.


1. Concept-based Visual Counterfactual Explanations with Diffusion Models ​

Author: Yassine Oueslati, Daniil Kirilenko, Martin Gjoreski, Marc Langheinrich
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2607.22544v1 Announce Type: new Abstract: Visual counterfactual explanations aim to answer "what minimal change to this image would flip the model's prediction?", and are increasingly important as vision models are deployed in safety-critical domains (e.g., medicine). Existing diffusion-based ...

📖 Read original article


2. SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series ​

Author: Giovanni B. Esposito, Francesco Antici, Daniele Cesarini, Andrea Bartolini
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.PF

arXiv:2607.22548v1 Announce Type: new Abstract: Data centers and their compute nodes require accurate and flexible digital twins capable of modeling the complex interplay of workloads, environmental parameters, and physical metrics. Current machine learning approaches for HPC and its telemetry typic...

📖 Read original article


3. QFoldAgent: An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction ​

Author: Winson Chen, Yuqi Zhang, Sixu Chen, Nuo Xu, Qiang Guan, Caiwen Ding
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22549v1 Announce Type: new Abstract: Hybrid quantum-classical protein structure prediction depends strongly on Hamiltonian penalty weights, yet existing lattice-based workflows typically fix these coefficients by hand and evaluate only very short fragments in simulation. We present QFoldA...

📖 Read original article


4. Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy ​

Author: Kazem Faghih, Yize Cheng, Shoumik Saha, Mobina Pournemat, Armin Gerami, Soheil Feizi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.22554v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change ...

📖 Read original article


5. DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs ​

Author: Mahmood Bayeshi, Veysel Kocaman, Muhammed Ali Naqvi, Yigit Gul, David Talby
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22555v1 Announce Type: new Abstract: Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with explanations. Frontier LLMs are strong generalists, but single-shot prompting often yields brittle diagno...

📖 Read original article


6. MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models ​

Author: Dong Li, Yanchi Liu, Xujiang Zhao, Wei Cheng, Zhengzhang Chen, Xintao Wu, Zhong Chen, Chen Zhao, Haifeng Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22556v1 Announce Type: new Abstract: Continual learning (CL) is essential for small language models (SLMs) to adapt to evolving real-world needs in resource-constrained deployments. However, directly updating their limited parameter space causes catastrophic forgetting. While memory-based...

📖 Read original article


7. Codifying the Judge: Scalable Evaluation via Program Distillation ​

Author: Tzu-Heng Huang, Shengqi Qiu, Frederic Sala
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.22561v1 Announce Type: new Abstract: LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternati...

📖 Read original article


8. SF-AMS: Strategic Forgetting for Structured Memory in LLM Agent ​

Author: Ning Yang, Siqi Li, Miaoxin Shen, Yuan Zhou, Meng Zhang, Tong Li, Haijun Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22562v1 Announce Type: new Abstract: Managing long-context dependencies remains a primary bottleneck in LLM agents, as redundant and irrelevant information can degrade multi-step reasoning. Strategic Forgetting for Agent Memory Systems (SF-AMS) is proposed as a framework for maintaining c...

📖 Read original article


9. Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents ​

Author: Sagar Chethan Kumar, Rohith Kanathur, Dhaval Patel, Kaoutar El Maghraoui
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22563v1 Announce Type: new Abstract: Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks such as AssetOpsBench rely on manually authored scenarios and cover a li...

📖 Read original article


10. Loss-Aware Feature-Map Pruning in Convolutional Neural Networks Using Multi-Armed Bandits ​

Author: Salem Ameen, Sunil Vadera
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22564v1 Announce Type: new Abstract: Convolutional neural networks often contain redundant feature maps that increase storage and inference cost. This paper presents a loss-aware feature-map pruning framework using multi-armed bandits. Feature-map pruning is structured because it removes ...

📖 Read original article


11. DSTFView: Multi-View Cloud-Edge Workload Forecasting with Dual-Input Spatio-Temporal-Frequency Modeling ​

Author: Qingzhong Li, Hui Ma, Yajun Zhang, Qingchang Ma, Zhou Long
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.22565v1 Announce Type: new Abstract: With the widespread deployment of edge-side AI inference, edge platforms are increasingly required to support latency-sensitive, highly concurrent, and reliability-critical applications. However, existing methods often struggle to balance multidimensio...

📖 Read original article


12. MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models ​

Author: Zeyu Zhang, Ziqing Wang, Kaize Ding
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22566v1 Announce Type: new Abstract: MedLoCoMo is a Medical Long-Context Memory benchmark for patient-specific clinical reasoning over multi-admission medical dialogue. Existing medical QA benchmarks largely test short context knowledge or single document grounding, leaving open whether L...

📖 Read original article


13. Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting ​

Author: Ruiyi Tao, Xiaolong Tu, Haoxin Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.NI

arXiv:2607.22568v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed on mobile and embedded devices to improve privacy and reduce network latency. Yet on-device inference faces a fundamental constraint: high energy consumption on battery-powered, resource-limited ha...

📖 Read original article


14. Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines ​

Author: Yifei Ge, Weisong Sun, Jinkun Xiao, Yuchen Chen, Yebo Feng, Peizhuo Lv, Xia Feng, Chunrong Fang, Zhihong Zhao, Zhenyu Chen, Yang Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2607.22569v1 Announce Type: new Abstract: Coding agents are increasingly integrated into system operations, where their tool use can directly modify project artifacts, execution environments, and the underlying system. For example, if a coding agent inserts a hook into a system startup or conf...

📖 Read original article


15. Reference Feature Atlases for Mechanistic Auditing of Language Models ​

Author: Rui Wu, Tong Che
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22570v1 Announce Type: new Abstract: Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained once on a reference panel and reused for new targets, which attach by f...

📖 Read original article


16. SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs ​

Author: Prateek Chaturvedi, Yuqicheng Zhu, Hongkuan Zhou, Dongzhuoran Zhou, Yunjie He, Steffen Staab, Fei Du, Jie Tang, Evgeny Kharlamov
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22571v1 Announce Type: new Abstract: Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) enables natural language interaction with structured enterprise knowledge, yet existing agentic approaches that perform well on public benchmarks often fail to generalize to real-world enter...

📖 Read original article


17. Schema-Aware Localisation (SAL): Live Schema Grounding and Hallucination Validation for Oracle NL2SQL ​

Author: Sanjay Mishra, Divya Chukkapalli, Ganesh R. Naik
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22572v1 Announce Type: new Abstract: Large language models can generate fluent SQL from natural language, but on real enterprise Oracle databases they frequently fail at execution time: columns and aliases are hallucinated and dialect-specific syntax is missed, leading to ORA-00904 invali...

📖 Read original article


18. PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability ​

Author: Wen-Kao Li, Ze-Feng Gao, Zhong-Yi Lu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci, physics.comp-ph

arXiv:2607.22573v1 Announce Type: new Abstract: Imaginary phonon modes remain a practical bottleneck in computational materials screening because otherwise plausible structures can be locally dynamically unstable under a chosen workflow. Here we present PhononBench-MP40, a spectrum-resolved benchmar...

📖 Read original article


19. Too much evidence, too little time: From text to actionable recommendations through multi-objective evidence reasoning ​

Author: Adela Bara, Simona-Vasilica Oprea
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2607.22574v1 Announce Type: new Abstract: Evidence-based clinical decision making requires specialists to identify, evaluate and synthesize relevant scientific literature. However, PubMed searches for complex clinical cases often return hundreds of publications that cannot be reviewed manually...

📖 Read original article


20. Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models ​

Author: Mathis Pink, Vy Ai Vo, Qinyuan Wu, Jianing Mu, Javier Turek, Uri Hasson, Kenneth A. Norman, Sebastian Michelmann, Alexander Huth, Mariya Toneva
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22575v1 Announce Type: new Abstract: Human episodic memory supports the retrieval of experiences that unfold over extended timescales, yet the computational mechanisms underlying this ability remain debated due to the limited mechanistic accessibility in long-term memory experiments in hu...

📖 Read original article


21. cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs ​

Author: Xin Yang, Yemin Wang, Mingda Liu, Letian Li, Shuaishuai Cao, Zhengxiao He, Ryan Dong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22577v1 Announce Type: new Abstract: Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size-a critical bottleneck...

📖 Read original article


22. HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization ​

Author: Size Li, Zhiqing Tang, Hongrui Liang, Jianxiong Guo, Jiong Lou, Tian Wang, Weijia Jia
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22578v1 Announce Type: new Abstract: The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concurrency, multi-tenant agentic workflows. However, existing solutions typically prioritize intra-workflow optimizat...

📖 Read original article


23. Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization ​

Author: Muhammad Junaid Ali, Smail Niar, El-Ghazali Talbi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.22583v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However, deploying them in embedded and edge computing environments remains challenging because of strict latency, memory,...

📖 Read original article


24. Source-Aware Reranking for Retrieval-Augmented Generation: A Reliability Prior Approach ​

Author: Yuktha Tata Koganti, Hugo Garrido-Lestache Belinchon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.22584v1 Announce Type: new Abstract: Standard Retrieval-Augmented Generation pipelines rank retrieved documents by semantic similarity alone, without accounting for source provenance or credibility. This work evaluates a simple and interpretable modification to RAG retrieval ranking that ...

📖 Read original article


25. The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation ​

Author: Naman Vats, Oleg Golev
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22585v1 Announce Type: new Abstract: Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools, manages context, and decides when to stop) is often under-specified. Model-to-model comparison is v...

📖 Read original article


26. MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models ​

Author: Jinsong Shu, Chenyang Wu, Zhongle Xie, Baokun Wang, Lidan Shou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.22586v1 Announce Type: new Abstract: Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows linearly with context length and becomes a major bottleneck due to the large number of visual tokens. Recent prefill...

📖 Read original article


27. TriSP: Tri-Signal Structured Pruning for Large Language Models ​

Author: Manel Kara laoua, Soumia Bouyahiaoui, Aicha Boutorh
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22587v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance across diverse tasks but their deployment is constrained by the memory and compute cost of their parameters. Structured pruning addresses this by removing entire structures such as attention heads...

📖 Read original article


28. ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation ​

Author: Samyak Jhaveri, Erel Kaplan, Tom Yotam, Le Chen, Tomer Bitan, Niranjan Hasabnis, Gal Oren
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.DC

arXiv:2607.22588v1 Announce Type: new Abstract: Modern compute-intensive software must migrate across a changing ecosystem of accelerators, programming APIs, compiler stacks, and portability layers, including CUDA, OpenMP, OpenCL, and OpenMP target offload. Large language models and autonomous codin...

📖 Read original article


29. Lexical discovery in unknown environments orchestrated by Large Language Models ​

Author: Rafael Sendra-Arranz, I~naki Dellibarda Varela, Eduardo Rocon, 'Alvaro Guti'errez, Manuel Cebrian
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.22591v1 Announce Type: new Abstract: Populations of autonomous agents deployed in unknown environments (e.g. planetary or deep-sea exploration) must develop shared vocabularies to refer to entities that have no name in any human language. We propose the Neuro-Symbolic Lexical Discovery (N...

📖 Read original article


30. Structure Over Scale: Schema-Constrained Causal Graphs for RAG ​

Author: Marc Saouda (Boston Consulting Group), Rajprakash Bale (Boston Consulting Group), Eren Aldis (Boston Consulting Group), Cloves Almeida (Boston Consulting Group)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR

arXiv:2607.22592v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (GraphRAG) grounds answers in structured knowledge, but current systems extract entities and relationships exhaustively, producing graphs whose size and construction cost scale with corpus length rather than w...

📖 Read original article


31. xMIx: High-Performance Serving-Time Platform for Mechanistic Interpretability Apps ​

Author: Michael Blum, Mark Silberstein, Yaniv David
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22595v1 Announce Type: new Abstract: Mechanistic interpretability (MI) has emerged as a powerful approach for analyzing and intervening in inference computations, with a growing number of applications such as jailbreak attempt detection, truthfulness evaluation, and hallucination detectio...

📖 Read original article


32. An Agentic Orchestration of Atomistic Simulations ​

Author: Rahul Somasundaram, Adela Habib, Khanh Dang, Sachin Shivakumar, Ryley G. Hill, Golo Wimmer, Avanish Mishra, Aleksandra Pachalieva, Arthur Lui, Hari Viswanathan, Michael Grosskopf, Saryu Fensin, Russell Bent, Nathan DeBardeleben, Earl Lawrence
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci

arXiv:2607.22596v1 Announce Type: new Abstract: Atomistic simulations are central to materials design, but their execution involves complex, multi-step workflows that require significant human expertise. Here, we present an agent-based system embedded within the URSA (Universal Research and Scientif...

📖 Read original article


33. HyCE-RAG: Hypergraph Chain-of-Evidence Retrieval-Augmented Generation for Explainable Multi-hop Question Answering ​

Author: Hong-Yu An, Yun-Jian Zhang, Chen-Wei Liang, Tian-Yi Zhang, Jian Ding, Yi-Lun Wu, Ao-Bo Li, Wei-Cong Su, Saifullah, Mujiangshan Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22597v1 Announce Type: new Abstract: Multi-hop question answering requires systems to retrieve evidence from multiple documents and connect scattered facts into a coherent reasoning process. Standard retrieval-augmented generation (RAG) mainly relies on semantic similarity between a query...

📖 Read original article


34. Differencing the Diffusion Trajectory toward Uncertain Components for Time Series Forecasting ​

Author: Chen Su, Yuanhe Tian, Yan Song
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22599v1 Announce Type: new Abstract: Diffusion models have become a widely used framework for probabilistic time series forecasting, modeling the distribution of future values given an observed history. In time series forecasting, however, the future continues the observed history, creati...

📖 Read original article


35. Chart Deception in Vision-Language Models: From Vulnerability to Mitigation ​

Author: Ridwan Mahbub, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Mizanur Rahman, Mir Tafseer Nayeem, Enamul Hoque
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22600v1 Announce Type: new Abstract: Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axes, distorted aspect ratios, inappropriate encodings, and misleading color mappings-can systematicall...

📖 Read original article


36. DeepLook: Deeper Thinking with Lookahead ​

Author: Tingxin Yang, Zefeng Wang, Mengyue Wang, Xingcheng Zhou, Yunpu Ma
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22602v1 Announce Type: new Abstract: Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on difficult reasoning tasks than parameter scaling alone. However, existing approaches remain inefficient in how comp...

📖 Read original article


37. Group Preference Collapse in Personalized Multimodal Large Language Models ​

Author: Fan Lyu, Wenqi Zhang, Joost van de Weijer
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22603v1 Announce Type: new Abstract: Personalized multimodal large language models (MLLMs) aim to generate user-specific responses, but existing methods mainly rely on profile-level information and overlook diverse user preferences. We identify group preference collapse, where multi-user ...

📖 Read original article


38. Evaluating LLMs as Interpretable Controllers for Dynamical Systems ​

Author: Aleksander {\O}stensen, Alberto Mino Calero, Anastasios M. Lekkas, Adil Rasheed
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22609v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for decision-making and reasoning tasks, yet their potential as controllers for physical systems remains largely unexplored. This work investigates whether LLMs can function as interpretable controller...

📖 Read original article


39. Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations ​

Author: Jessica Tang, Shraddha Barke, Sharad Agarwal
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.22610v1 Announce Type: new Abstract: When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those dependencies propagate across prior turns? Existing context attribution methods process the full context in a si...

📖 Read original article


40. Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure ​

Author: Arun Malik, Deepal Jayasinghe, Bradley Klemick, Prachi Shah, Nitish Talasu, Vineet Tushar Trivedi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.ET, cs.MA

arXiv:2607.22611v1 Announce Type: new Abstract: The deployment of autonomous AI agents in production infrastructure introduces fundamental security challenges that traditional role-based access control (RBAC) models cannot address. Unlike deterministic automation, AI agents exhibit stochastic behavi...

📖 Read original article


41. DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training ​

Author: Hanlin Du, Zhiyuan Yan, Haiquan Chen, Jiarui Fang, Yungang Bao, Sa wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22614v1 Announce Type: new Abstract: RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU reallocatio...

📖 Read original article


42. Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning ​

Author: Aamir Hamid, Bharg Barot, Satvik Racharla, Tim Finin, Primal Pappachan, Roberto Yus
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22621v1 Announce Type: new Abstract: While large language models (LLMs) enable strong question answering (QA), budgeted deployment is complicated by nondeterminism and heterogeneous resource profiles (cost, latency, and energy). We present OPTI-Q, a database-inspired, cost-based optimizer...

📖 Read original article


43. CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process ​

Author: Minghao Yang, Yanjun Xu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22624v1 Announce Type: new Abstract: Recently, there have been several works in the Text-to-SQL domain that utilize Small Language Models (SLMs) for training. These approaches achieve performance close to that of large models in generating SQL, using only the computational power of a sing...

📖 Read original article


44. TokenMem: Faithful Knowledge Injection for Frozen LLMs ​

Author: Chengzhang Yu, Chenyang Zheng, Zening Lu, Yingru He, Yutong Huang, Yiming Zhang, Yue Xu, Zhanpeng Jin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.22625v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge, but suffers from knowledge conflicts: when retrieved information contradicts parametric memory, the shared self-attention pathway produces unpredictable...

📖 Read original article


45. Masked Distillation: Internalizing the Chain-of-Thought in Language Models ​

Author: Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.22629v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate traces dominate latency, memory usage, and serving cost, even though the final answer correctness i...

📖 Read original article


46. VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing ​

Author: Yexiang Liu, Wen Zhong, Sijie Zhu, Xin Gu, Fan Chen, Junxian Duan, Jie Cao, Longyin Wen, Zhenfang Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22632v1 Announce Type: new Abstract: The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. However, vlog assessment is highly subjective and remains challenging due to a lack of standardized crit...

📖 Read original article


47. Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering ​

Author: Guixin Su, Qiankun Pi, Mayi Xu, Wenli Li, Ming Zhong, Yuanyuan Zhu, Jiawei Jiang, Tieyun Qian
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22633v1 Announce Type: new Abstract: Table Question Answering (TableQA) aims to reason over tables to answer user queries. Existing research treats all questions uniformly and evaluates solely through overall accuracy, obscuring a critical reality that LLMs excel at simple lookups yet str...

📖 Read original article


48. PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding ​

Author: Zheng Wang, Zhifan Ye, Qi Cheng, Yonggan Fu, Ziyan Wang, Feng Zhu, Haozhe Zhao, Jan Kautz, Pavlo Molchanov, Humphrey Shi, Minjia Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22634v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs, generating tokens in parallel. This makes them effective draft models for speculative decoding (SD), producing an entire block of draft tokens ...

📖 Read original article


49. CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants ​

Author: Xuzhao Geng, Haozhao Wang, Xuelian Li, Zhenyu Yang, Haonan Lu, Rui Zhang, Ruixuan Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22635v1 Announce Type: new Abstract: Target-oriented dialogue systems have demonstrated strong capabilities in completing user goals through interactive conversations. However, existing studies are primarily designed for single, explicit goal completion, while phone call assistants face a...

📖 Read original article


50. Answering Path Queries under Linear and Guarded Existential Rules ​

Author: Jean-Fran\c{c}ois Baget (LIRMM, Inria, University of Montpellier, CNRS, France), Meghyn Bienvenu (Univ. Bordeaux, CNRS, Bordeaux INP, LaBRI, France), Marie-Laure Mugnier (LIRMM, Inria, University of Montpellier, CNRS, France), Micha"el Thomazo (Inria, DIENS, ENS, PSL University, CNRS, France)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22636v1 Announce Type: new Abstract: Ontology-mediated query answering is concerned with the problem of answering queries over knowledge bases consisting of a database instance and an ontology. While most work in the area focuses on conjunctive queries (CQs), navigational queries have gai...

📖 Read original article


51. Fast Cross-Scenario Adaptation of CSI Models via Channel Conditional Parameter Generation ​

Author: Xudong Zou, Siyu Wu, Zunlei Feng, Jie Song, Yuanyu Wan, Mingli Song, Jiacong Hu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.IT, math.IT

arXiv:2607.22637v1 Announce Type: new Abstract: Deep learning has shown strong potential for massive multiple-input multiple-output (Massive MIMO) physical-layer tasks, including channel state information (CSI) feedback and channel estimation. However, environmental heterogeneity can severely degrad...

📖 Read original article


52. TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs ​

Author: Sai Shruthi Sistla, Ashutosh Hathidara, Christopher Toukmaji, Mayank Shrivastava, Karthikeyan Asokkumar
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22639v1 Announce Type: new Abstract: Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via constrained beam search. Toolsense shows that this regime has two critical drawbacks: it destroys para...

📖 Read original article


53. CRAFT: Learn the Schema, Execute the Plan ​

Author: Aakash Kolekar, Sahika Genc, Shahriar Shariat, Bunyamin Sisman, Tibor Mezi, Barbara Poblete, Shree Vandana Kachroo, Calvin Chi, Parth Parmar, Ari Singer, Prayaas Jain, Cindy Barker, Benoit Dumoulin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA, cs.SE

arXiv:2607.22642v1 Announce Type: new Abstract: Enterprise coding agents translate natural-language analytical requests into executable code over proprietary APIs, schemas, and metric definitions. Yet the prevailing deployment pattern injecting exhaustive schema and tool documentation into each prom...

📖 Read original article


54. Reason Before You Retrieve: Agentic Planning for Multi-modal RAG ​

Author: Tianyu Yang, Shir Simon, Zhenzhen Li, Minhao Cheng, Xiangliang Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2607.22643v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal input over a flat evidence space. This design often struggles with two key...

📖 Read original article


55. DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification ​

Author: Mohammed Yousif, Prabhjot Singh, Arjun Pankajakshan, Madhu Reddiboina
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2607.22644v1 Announce Type: new Abstract: Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or type. This leads to inefficient use of compute and human resources: simple documents are over-processed...

📖 Read original article


56. Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models ​

Author: Yijia Dai, Zhaolin Gao, Yahya Sattar, Jennifer J. Sun, Sarah Dean
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.22646v1 Announce Type: new Abstract: Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but the algorithm underlying this capability remains undetermined: prior work has proposed several cand...

📖 Read original article


57. PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving ​

Author: Meghana Maghyastha, Robert Underwood, Randal Burns, Bogdan Nicolae
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.DC

arXiv:2607.22648v1 Announce Type: new Abstract: Inspired by the design of client caching in Content Delivery Networks (CDNs), PTStore distributes and replicates popular tensors that form reusable KV cache prefixes, which are the main technique used by state of art approaches to accelerate inferences...

📖 Read original article


58. STAIF: A Stage-wise Optimization for Complex Instruction Following ​

Author: Jian Hong, Chen Cheng, Quan Liu, Yuhao Chen, Enhong Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.22649v1 Announce Type: new Abstract: Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment methods, such as DPO, optimize holistic reward signals that often underemphasize strict satisfaction ...

📖 Read original article


59. ARdena: Scenario-driven control of real-time LLM agents ​

Author: Luka Borozan, Domagoj Matijevi'c
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22651v1 Announce Type: new Abstract: Large language models (LLMs) have enabled increasingly capable conversational agents, but reliably controlling their behavior in real-time interactive environments remains a significant challenge. Existing approaches often rely on model fine-tuning or ...

📖 Read original article


60. KG2Code: Bridging Knowledge Graphs and Large Language Models via Executable Code for Question Answering ​

Author: Yike Wu, Nan Hu, Guilin Qi, Guohui Xiao, Chen Jiang, Xinchun Zou, Yuchen Lu, Songlin Zhai, Yongrui Chen, Yuyang Zhang, Xiaoguang Li, Lifeng Shang, Jiaoyan Chen, Jeff Z. Pan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22652v1 Announce Type: new Abstract: Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream knowledge-intensive tasks, particularly knowledge graph question answering (KGQA). Existing approaches p...

📖 Read original article


61. Do Language Models Converge to Themselves? Recursive Self-Refinement as Textual Relaxation ​

Author: Xuening Wu, Qianya Xu, Yanlan Kang, Zeping Chen, Yubin Liu, Shenqin Yin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22653v1 Announce Type: new Abstract: Large language models are increasingly used in recursive refinement workflows, where an initial draft is repeatedly revised by the same model. Despite their growing use, the long-term dynamics of such workflows remain poorly understood. Does repeated r...

📖 Read original article


62. MINT-V2X: A Mobility-Integrated Network Trajectory Dataset for Predictive Resource Management ​

Author: Abdullah Anjum, Abdolazim Rezaei, Mehdi Sookhak
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22654v1 Announce Type: new Abstract: Vehicle-to-Everything (V2X) communication systems are based on datasets that not only contain vehicle trajectory data but also wireless network parameters with a realistic level of fidelity, enabling the creation of prediction and optimization models. ...

📖 Read original article


63. EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation ​

Author: Jie Zhao, Jie Feng, Can Rong, Zhihan Hou, Peng Lu, Yong Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22655v1 Announce Type: new Abstract: Estimating origin-destination (OD) flows under disruptive events is important for disaster response and urban resilience. Existing deep OD models trained on routine mobility often degrade when extreme events abruptly alter regional functions and popula...

📖 Read original article


64. StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech ​

Author: Yuzhe Wang (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Thomas Thebaud (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Jennifer Hu (Department of Cognitive Science, Johns Hopkins University, Baltimore, USA), Jes'us Villalba-Lopez (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Venkatesh Ravichandran (Amazon AGI, USA), Georgi Tinchev (Amazon Research, UK), Najim Dehak (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Laureano Moro-Vel'azquez (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.SD, eess.AS

arXiv:2607.22658v1 Announce Type: new Abstract: Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceBench, a benchmark for measuring interpersonal stance in conversational s...

📖 Read original article


65. TRE: Training-Free Hallucination Detection for Diffusion Language Models ​

Author: Pengcheng Weng, Yanyu Qian, Yue Tan, Yixin Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22661v1 Announce Type: new Abstract: Diffusion large language models (D-LLMs) have recently gained increasing attention, yet their reliability is significantly hindered by the hallucination problem. Existing hallucination detection approaches for D-LLMs mainly follow a training-based para...

📖 Read original article


66. CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data ​

Author: Peiguang Li, Yongwei Zhou, Juncheng Diao, Yuchun Fan, Jian Yang, Jianxiao Yang, Zhongda Su, Shuguang Jiao, Xiao Wei, Zhiye Zou, Gan Dong, Zhizhao Zeng, Rongxiang Weng, Jingang Wang, Xunliang Cai
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22662v1 Announce Type: new Abstract: Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance. However, these pipelines typically rely on singular optimization objective...

📖 Read original article


67. Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models ​

Author: Xingyu Mou, Zijin Huang, Tianze Zhang, Yuxin Ma, Lanning Wei, Zengfeng Huang, Da Zheng, Lun Du
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22663v1 Announce Type: new Abstract: Block diffusion has emerged as the dominant paradigm for scaling discrete diffusion language models (dLLMs), because decoding text in fixed-size blocks preserves parallel generation within each block while keeping the quadratic attention cost tractable...

📖 Read original article


68. Obliviate: Efficient Unlearning in Recommender Systems ​

Author: Tushar Prakash, Brijraj Singh, Niranjan Pedanekar, Narayan Chaturvedi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22665v1 Announce Type: new Abstract: Machine unlearning is becoming increasingly critical in the context of data privacy regulations, particularly for recommendation systems that are directly trained on user interaction data. The goal of this work is to remove requested interaction data a...

📖 Read original article


69. Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance ​

Author: Andrei Starodubov, Yaqub Aris Prabowo, Andreas Hadjipieris, Roberto Galeazzi, Ioannis Kyriakides
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.IT, cs.LG, cs.RO, cs.SY, eess.SP, eess.SY, math.IT

arXiv:2607.22667v1 Announce Type: new Abstract: This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-vessel tracking in heterogeneous maritime sensor networks. The proposed approach is motivated by information-theoretic sensor management: instea...

📖 Read original article


70. AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models ​

Author: Rohan Naphade, Minzhou Pan, Bo Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22671v1 Announce Type: new Abstract: Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation, their risk taxonomies become incomprehensive and their attack prompts become ineffective. We present...

📖 Read original article


71. How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift ​

Author: James Elcock, William F. Shen, Xinchi Qiu, Nicholas D. Lane
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.22676v1 Announce Type: new Abstract: Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment dom...

📖 Read original article


72. DOSA: A Tree-Guided, Self-Regressive Framework for Long Document Structure Analysis ​

Author: Bohou Li, Benjamin Sowell, Mehul Shah, Mark Lindblad, Henry Lindeman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22679v1 Announce Type: new Abstract: In visually-rich documents, information is encoded not only in individual page objects such as tables, headers, and text blocks, but also in the structural relations among them, making document structure analysis fundamental to information retrieval an...

📖 Read original article


73. A Vocabulary for Multi-Agent Automated Research Systems ​

Author: Bardiya Akhbari
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA

arXiv:2607.22682v1 Announce Type: new Abstract: We introduce a vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare. The vocabulary specifies 1) who the agents are, 2) what operations are available in the system, 3) who m...

📖 Read original article


74. Imprompt: A Language Framework for Prompt Programming ​

Author: Chentian Wu, Shengyuan Yang, Adithya Murali
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.PL

arXiv:2607.22683v1 Announce Type: new Abstract: With the unprecedented success of Language Models (LMs), the science of Prompt Engineering has evolved the powerful idea of Prompt Programming, where prompts are treated as a programmable control surface for describing complex tasks and leveraging LM c...

📖 Read original article


75. Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents ​

Author: Zhengyu Chen, Teng Xiao, Huaisheng Zhu, Yige Yuan, Luan Zhang, Jingang Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.22688v1 Announce Type: new Abstract: Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, evaluated, and learned from. Existing pipelines typically train models under...

📖 Read original article


76. Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents ​

Author: Zedong Yu, Qianxing Li, Zhi Gao, Liuyu Xiang, Chenrui Shi, Yang Liu, Huiming Wu, Yujie Wei, Yuhao Fei, Yubo Fu, Zhaofeng He
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22689v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as clicking, typing, and scrolling on desktops and mobile devices. However, c...

📖 Read original article


77. LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory ​

Author: Jing Yu, Yibo Zhao, Jiaming Zhang, Xiang Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22690v2 Announce Type: new Abstract: Long-term memory enables LLM agents to leverage past interactions, but dialogue histories quickly exceed the context window, forcing agents to retrieve relevant subsets at query time. Because useful evidence is sparse and scattered across verbose conve...

📖 Read original article


78. HiLLTS: Zero-Shot Hierarchical LLM-Guided Traffic Signal Control for Sustainable Transportation ​

Author: Yue Ding, Tendai Mukande, Mingming Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22691v1 Announce Type: new Abstract: Urban traffic congestion significantly increases fuel consumption, greenhouse gas emissions, and commuter delays, resulting in substantial economic losses and environmental harm in modern cities. Traditional traffic signal control strategies such as fi...

📖 Read original article


79. Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture ​

Author: Anabela C. Areias, Catarina Botelho, Ant'onio Farinhas, Areti Vassilopoulos, Dora Janela, Xin Tong, Nuno M. Guerreiro, Maya D'Eon, Fab'iola Costa, Ricardo Rei
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existing safety approaches primarily detect risk but rarely shape how models respond as conversational ris...

📖 Read original article


80. Bayesian Repetition Penalty: A Principled Adjacent-Conditional Framework for Reversing Attention Collapse in Autoregressive Language Models ​

Author: Wenjie Fan, Bin Ma, Dong Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22694v1 Announce Type: new Abstract: Attention collapse in autoregressive language models -- manifested as repetitive token loops where the model becomes trapped in self-reinforcing attractors -- is a persistent pathology that existing decoding-time heuristics fail to address at its root ...

📖 Read original article


81. PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window ​

Author: Ryan Thornton, Mir Mehedi Ahsan Pritom, Maanak Gupta
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2607.22695v1 Announce Type: new Abstract: Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widespread deployment. While this automation provides clear utility, completing these tasks often requires the insertion o...

📖 Read original article


82. Test-Time Coverage: Test-Conditioned Data Curation for Deployment-Aware Learning ​

Author: Nadine Chang, Maying Shen, Shizhe Diao, Jialiang Wang, Jingde Chen, Thomas Breuel, Pavlo Molchanov, Rafid Mahmood, Jose M. Alvarez
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2607.22697v1 Announce Type: new Abstract: Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution. However, standard data curation methods score training-side criteria rather than directly optimizing deployment...

📖 Read original article


83. Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures ​

Author: Supantho Rakshit, Adele Goldberg, Henry Conklin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22699v1 Announce Type: new Abstract: As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilingually, t...

📖 Read original article


84. RoleMix: Unifying Sequential and Non-Sequential Features via Semantic Tokenization for Post-Click Conversion Rate Prediction ​

Author: Wenan Wang, Qin Zhao, Zhixiang Lu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22700v1 Announce Type: new Abstract: Post-click conversion rate (PCVR) prediction is central to industrial recommendation, but remains challenged by the structural mismatch between sparse, unordered multi-field features and long, domain-specific behavior histories. Existing models often p...

📖 Read original article


85. MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation ​

Author: Hyewon Lee, Minkyung Song, Junghyun Oh, Seunghoon Han, Sungsu Lim
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.DL, cs.IR

arXiv:2607.22706v1 Announce Type: new Abstract: This paper presents the MPR-CiteG framework, which achieved second place in the ScienceON AI Challenge by addressing two fundamental challenges in generative AI: inefficient retrieval and the absence of source verification. We propose a dual-component ...

📖 Read original article


86. SEGRA: Structured Experience-Guided Graph Reasoning Agent for Gremlin Based Question Answering ​

Author: Saiyue Lyu, Mariam Dundua, Vishaal Kapoor, Sarthak Ahuja, Neda Kordjazi, Evren Yortucboylu, Harsh Amin, Rebecca Steinert
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22713v1 Announce Type: new Abstract: Enterprise IT support knowledge graphs capture rich relationships among cases, users, devices, symptoms, taxonomic categories, root causes, and historical resolutions. Yet querying them in Gremlin requires knowledge of graph schemas, traversal semantic...

📖 Read original article


87. Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning ​

Author: Mohit Jiwatode, Ronja Fuchs, Robin Schm"ocker, Bodo Rosenhahn, Alexander Dockhorn
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22732v1 Announce Type: new Abstract: LLM-based game agents often perform poorly on more complex tasks. This work examines whether these failures are linked to limited spatial reasoning and evaluates whether causal prompt augmentation and multi-step planning can improve win-rates while man...

📖 Read original article


88. Commitment To Cooperation With Self-Negotiated Contracts ​

Author: Tim Wyse, Kaitlin Bustos, Yulia Volkova, Max Kleiman-Weiner
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22750v1 Announce Type: new Abstract: As AI agents operate with increasing autonomy in a multi-agent world, they will need to learn to cooperate with other agents and with humans to generate mutual benefits. However, cooperation is a challenge because the costs of cooperation are often inc...

📖 Read original article


89. Disentangling Multi-View Scanning in Mamba for Network Traffic Anomaly Detection ​

Author: Xinglin Lian, Chengtai Cao, Ting Zhong, Fan Zhou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22829v1 Announce Type: new Abstract: Network Traffic Anomaly Detection (NTAD) is a critical task in cybersecurity, yet timely and accurate anomaly detection remains challenging. Mamba has emerged as a particularly promising backbone for NTAD due to its linear-time complexity for long-sequ...

📖 Read original article


90. Coordinated Networking for On-Device Agent-Augmented Real-Time Communication ​

Author: Goodsol Lee, Juheon Yi, Jinglu Wang, Haowen Xu, Saewoong Bahk, Yan Lu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22854v1 Announce Type: new Abstract: AI agents are enabling a new paradigm of agent-augmented real-time communication (RTC), where humans focus on high-level collaboration, while agents autonomously retrieve, analyze, and generate information in real time to support their interactions. Th...

📖 Read original article


91. What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents ​

Author: Shawn Ray
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.LG

arXiv:2607.22868v1 Announce Type: new Abstract: Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whether intervention changes future behavior. We separate three questions. First, relative to fixed ora...

📖 Read original article


92. Physical AI Governance: From Theory to Practice Across Life Cycle ​

Author: Wang Yang, Shaobo Wang, Hongxuan Liu, Xiaoran Cai, Yunyu He, Jingzong Zhou, Mengzhong Ma, Yi Yu, Rohit Sharma, Jingjing Fu, Peng Qi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.RO

arXiv:2607.22877v1 Announce Type: new Abstract: With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact with, and act in the physical world. Unlike traditional AI, Physical AI operates under real-time safety...

📖 Read original article


93. How Well Can AI Generate Backlogs from App Mockups? ​

Author: Andrea Lezcano Airaldi, Lourdes Romera, Walid Maalej
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2607.22902v1 Announce Type: new Abstract: Creating sprint backlogs requires considerable effort, as items such as epics, user stories, and tasks can be missed or inconsistently specified. We propose a multimodal approach to support backlog generation from visual app mockups, an artifact availa...

📖 Read original article


94. Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Coding Agent Teams ​

Author: Shouren Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA

arXiv:2607.22917v1 Announce Type: new Abstract: Large Language Model (LLM) agents have significantly improved coding and programming workflows. Claude Code, in particular, is one of the most powerful LLM coding agents and is capable of conducting complex coding tasks. However, several drawbacks can ...

📖 Read original article


95. SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI ​

Author: Mahdi Eslamimehr
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2607.22926v1 Announce Type: new Abstract: High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-first, authorization-separated architecture in which credible catastrophic-enablement risk constrains admissibility...

📖 Read original article


96. Design Theater: A Benchmark for Generative UI ​

Author: Kashif Imteyaz, Kaif Imteyaz, Nakul Rajpal, Kaif Shaikh, Michael Muller, Saiph Savage
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22928v1 Announce Type: new Abstract: Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces. Alongside the interface, these tools generate user-facing design rationales that explain their layout, accessibility, and design cho...

📖 Read original article


97. Let AI Agents Translate Networks, Not Reason About Them ​

Author: Hongyu H`e, Maria Apostolaki
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.NI, cs.SC

arXiv:2607.22947v1 Announce Type: new Abstract: A formal model enables verifying reachability, localizing an outage, or anticipating the blast radius of a change. Yet, virtually no production network has one, since writing a model by hand demands rare expertise and is hard to keep current as the net...

📖 Read original article


98. Share No More Than the Request Requires: Federated Disclosure for Perspective-Aware AI ​

Author: Sourena Khanzadeh, Daniel Platnick, Marjan Alirezaie, Hossein Rahnama
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.CY, cs.SI

arXiv:2607.22953v1 Announce Type: new Abstract: Modern AI systems bring societal risks such as mass surveillance, extreme concentrations of power, and loss of user autonomy---calling into question a model where third-parties collect and control massive amounts of user data. Users require a sovereign...

📖 Read original article


99. ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control ​

Author: Yan Zhang, Shibo Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.22962v1 Announce Type: new Abstract: LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated fact written at one step therefore persists as a false premise for every subsequent step, a failure...

📖 Read original article


100. Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming ​

Author: Zirong Chen, Meiyi Ma
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LO, cs.SC

arXiv:2607.23019v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate steps are not guaranteed to be logically sound. We present Reason Popper-ly, a neurosymbolic framework that uses ...

📖 Read original article


101. Stress-testing large language model agents in a robotic chemistry laboratory ​

Author: Lulu Guo, Yingkai Sun, Xiaobo Li, Luyao Ge, Ziming Wang, Haitao Zheng, Jingyu Li, Huijuan Zhang, Bingxu Chen, Daobin Liu, Yuebo Liu, Jie Li, Xiaohui Li, Linjiang Chen, Yi Luo, Jun Jiang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2607.23045v1 Announce Type: new Abstract: AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. Here, we use a robotic chemistry laboratory as a physical-world testbed to make scientific agency meas...

📖 Read original article


102. SymStep: Symbolic Step Verification for Logical Reasoning ​

Author: Aida Usmanova, Rui Gao, Dilshod Azizov, Ricardo Usbeck, Zangir Iklassov
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23055v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps. We introduce SymStep: an LLM makes one atomic claim at a time (DEDUCE: Alice, pet, Cat), then a li...

📖 Read original article


103. Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning ​

Author: Narthana Sivalingam, Santhirarajah Sivasthigan, Buddhi Wijenayake, Roshan Godaliyadda, Vijitha Herath, Parakrama Ekanayake
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23077v1 Announce Type: new Abstract: Modeling multi-entity temporal data requires capturing dependencies across entities, time, and their interactions. Transformer-based approaches perform well but often rely on deep stacks of layers to learn these heterogeneous dependencies implicitly, i...

📖 Read original article


104. Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization ​

Author: Dongjie Chen, Ping Zhao, Bohua Zhan, Yulong Wang, Shushu Chen, Liangjun Feng, Hao Zhou, Min Shen, Linmu Wang, Weijia Sheng, Xiangyu Wei, Weijie Ding, Jianhui Huang, Yaoqing Gao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.PL

arXiv:2607.23089v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled automated kernel generation and optimization, but most existing approaches rely on surface signals such as compilation feedback and profiling metrics. These signals reveal that a kernel is sl...

📖 Read original article


105. SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows ​

Author: Summer Sun (Shaqiu Community)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23123v1 Announce Type: new Abstract: Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced within a constrained workflow as the unit of evaluation. We introduce SQBench, a benchmark for evalu...

📖 Read original article


106. AgentOmnia: Scaling Agentic Models for Full-Scenario Applications ​

Author: Hao Jiang, Gangtao Xin, Yingdi Huang, Guojie Zhu, Jiangshan Zhang, Xinyuan Lin, Yunkun Xu, Chengyu Shen, Wenlong Fei, Jiawei Li, Yujie Fu, Sichen Kang, Tingyu Xie, Yedi Hu, Jingren Zhang, Hongcheng Gao, Jianshu Zeng, Chong Chen, Chang Guo, Chao Feng, Feng Wang, Fulin Lin, Jinchao Ma, Lang Mei, Li Huang, Liyan Liu, Qing He, Shuting Tao, Siyu Mo, Xiangnan Chen, Xiaohan Yu, Xiaoyang Li, Yanheng Hou, Yanyu Wu, Zhihan Yang, Wentao Zhang, Yang Gao, Zhao Cao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.23124v1 Announce Type: new Abstract: Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-scenario agentic scaling and present AgentOmnia, a framework coordinating ...

📖 Read original article


107. CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion ​

Author: Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.MM

arXiv:2607.23159v1 Announce Type: new Abstract: Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although most are discarded. Training-free caching makes each rollout 2-3x faster at near-lossless quality. Composition is sa...

📖 Read original article


108. An Ontology for Machine Learning Interatomic Potentials ​

Author: Daniel Hern'andez, Jong Hyun Jung, Yuji Ikeda, Yongliang Ou, Pranav Kumar, Tom Sch"achtel, Wenchuan Liu, Xin Li, Xi Zhang, Xiang Xu, Lifang Zhu, Fritz K"ormann, Steffen Staab, Blazej Grabowski
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.DB, cs.DL

arXiv:2607.23219v1 Announce Type: new Abstract: Machine learning interatomic potentials (MLIPs) approximate quantum-mechanical energies and forces---conventionally computed by density functional theory (DFT) or wave-function methods---at a fraction of the cost. The field encompasses a growing ecosys...

📖 Read original article


109. Characterisation of Density-based FM generation methods in the context of Information Fusion ​

Author: Yanhao Huang, Christian Wagner
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23243v1 Announce Type: new Abstract: Fuzzy Integral (FI) based aggregation provides a powerful mechanism for nuanced aggregation, for example, in ensemble approaches or decision-level fusion more generally. The main challenge of this approach is the appropriate parametrization of the Fuzz...

📖 Read original article


110. CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics ​

Author: Xinhong Xu, Yimeng Zhang, Yuanlong Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.23258v1 Announce Type: new Abstract: Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central question remains unresolved: \textbf{whether a model pretrained on one collection of recordings can generalize to new ...

📖 Read original article


111. SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents ​

Author: Yang Wan, Zhenhao Zhang, Jierui Wang, Linchao Zhu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23263v1 Announce Type: new Abstract: Deciding whether a trajectory actually fulfills its instruction governs how we measure computer-use agents on long-horizon graphical-user-interface tasks and how we train them with reinforcement learning. This judgment has long relied on rule-based eva...

📖 Read original article


112. TopoFE: topology-aware LLM-guided Automated Feature Engineering ​

Author: Sha Li, Naren Ramakrishnan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.23286v1 Announce Type: new Abstract: Automatic feature engineering (AutoFE) for tabular learning can be naturally formulated as a program synthesis problem, where the objective is to discover predictive feature transformations from an exponentially large search space. Recent advances in l...

📖 Read original article


113. RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning ​

Author: Xi Chen, Hongru Zhou, Shiyu Feng, Hanyu Zhou, Huahui Yi, Rongsheng Wang, Tiancheng He, Kun Wang, Pingping Liu, Qiankun Li, Sicheng Lin, Huiying Ou, Xiaohong Zheng, Tianying Zang, Zhuohang Wu, Leheng Jiang, Kexin Cao, Wenhan Zhang, ChengYi Li, Zhiyang Wang, Songlin Li, Benyou Wang, Ningbei Yin, Shaoting Zhang, Weili Fu, Jian Li, Kang Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23290v1 Announce Type: new Abstract: Rare diseases collectively affect an estimated 3.5% to 5.9% of the population, yet more than 70% of patients are misdiagnosed and many endure years of evaluation before a diagnosis is reached, because early presentations are nonspecific and relevant ex...

📖 Read original article


114. Ordered Network Analysis of Epistemic Emotions during Collaborative Problem Solving ​

Author: Sifatul Anindho, Videep Venkatesha, Jaclyn Ocumpaugh, Nathaniel Blanchard
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2607.23317v1 Announce Type: new Abstract: Investigating how affective states such as confusion and frustration persist and transition during co-situated collaborative problem solving (CPS) is important for understanding the dynamics of epistemic emotions. However, the accurate identification o...

📖 Read original article


115. ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications ​

Author: Toby Liang, Gopal Sarda, Sagar Davasam, Vikas Yadav
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23326v1 Announce Type: new Abstract: The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world settings presents unique challenges due to complex system constraints and unexpected user behaviors. Amo...

📖 Read original article


116. Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation ​

Author: Paul Simpson, John Kozak, Lisa Doake
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional rules of the form "A is required UNLESS B applies, UNLESS C overrides B". The failure reproduces at...

📖 Read original article


117. Key-Interval A*: Accelerating Grid Pathfinding via Structural Abstraction ​

Author: Taiquan Sui
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23393v1 Announce Type: new Abstract: Existing exact methods for 4-connected grid pathfinding reduce online search, but often either retain fine-grained search states or require substantial preprocessing. This paper presents Key-Interval A* (KIA*), an optimal pathfinding algorithm that use...

📖 Read original article


118. Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning ​

Author: Adhyyan Narang, Artin Tajdini, Claire Zhang, Jamie Morgenstern
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23394v1 Announce Type: new Abstract: Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly. Standard defenses, such as data filtering, ...

📖 Read original article


119. NeurGO: Learning to Generate Elite Candidates for Meta-Black-Box Expensive Optimization ​

Author: Jintao He, Huixiang Zhen, Wenyin Gong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23408v1 Announce Type: new Abstract: Expensive black-box optimization is ubiquitous in science and engineering, where function evaluations are costly and the evaluation budget is limited. Traditional evolutionary algorithms and Meta-BlackBox Optimization (MetaBBO) approaches typically con...

📖 Read original article


120. Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels ​

Author: Haining Zheng, Qian Dong, Rodolfo K. Depena, Jonathan D. Bhatia, Feng Xiao, Peng Xu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.MA

arXiv:2607.23438v1 Announce Type: new Abstract: As AI systems increasingly exhibit agentic behavior, discussions of autonomy often conflate what systems are technically capable of doing with what they should be permitted to do in practice. This paper introduces a governance framework that explicitly...

📖 Read original article


121. Do LLMs Know Their Vulnerable Scenarios? ​

Author: Ziheng Peng, Huiqi Deng, Haoran Jing, Xuankun Rong, Jiahui Han, Xiting Wang, Na Zou, Xia Hu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2607.23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed att...

📖 Read original article


122. Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability Diagnosis ​

Author: Xinhao Yao, Yuanzhuo Liu, Changhao Wang, Yunfei Yu, Haoran Tan, Yuyao Zhang, Ruifeng Ren, Minlong Peng, Yong Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23524v1 Announce Type: new Abstract: Deep search is becoming a core capability of modern agent systems, yet it is typically evaluated solely based on end-to-end answer accuracy. This coupled evaluation paradigm entangles retrieval quality, long-context comprehension, evidence verification...

📖 Read original article


123. ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness ​

Author: Qiao Yan, Yihan Wang, Zhenghao Xing, Jiaqi Xu, Pheng-Ann Heng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23537v1 Announce Type: new Abstract: Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality. As a result, it remains unclear how vision-language mod...

📖 Read original article


124. Verification-Notebook Learning for Source-Aware Multimodal Misinformation Detection ​

Author: Junyuan Tan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23581v1 Announce Type: new Abstract: Multimodal misinformation verification is challenging because misleading signals may come from different parts of a post and require different forms of evidence. LVLMs are well suited to this task, but their verification performance often depends on th...

📖 Read original article


125. Are You Still the Agent I Authorized? Earned Authority under a Fixed Ceiling for Evolving Agents ​

Author: Zhaoxi Zhang, Xiaomei Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2607.23586v1 Announce Type: new Abstract: Long-lived AI agents increasingly evolve after deployment by retaining experience, acquiring skills and tools, revising workflows, delegating work, and moving across task phases. This improves adaptation but creates a distinct authorization problem. To...

📖 Read original article


126. Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning ​

Author: Wenxuan Zhang, Yuhui Wang, Donggang Jia, Xiaoqian Shen, Jian Ding, Ivan Viola, J"urgen Schmidhuber, Mohamed Elhoseiny
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23605v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-making across turns. Although end-to-end training in agentic environments can improve such multi-turn decision-mak...

📖 Read original article


127. SpecAHD: Localize to Specialize for Automated Heuristic Design in Large-Scale Routing Problems ​

Author: Kezhao Lai, Yutao Lai, Hai-Lin Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23676v1 Announce Type: new Abstract: LLM-based automated heuristic design (AHD) typically scores executable programs on complete instances or within fixed solver components. In large-scale routing problems, localized reconstruction reduces the size of each optimization task, but repair re...

📖 Read original article


128. Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems ​

Author: Mingzhou Fan, Siyuan Xu, Mingxuan Yuan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23678v1 Announce Type: new Abstract: Large language models (LLMs) enable autonomous agents for reasoning, planning, and tool use. Recent systems increasingly organize these agents as graphs of specialized, interconnected nodes. Although graph-based orchestration supports flexible decompos...

📖 Read original article


129. Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV ​

Author: Zefeng Cai, Zerui Cai
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23693v1 Announce Type: new Abstract: Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries and drops the rest. Eviction and episodic-memory schemes therefore rest on a premise rarely tested directly, that a retained event is sti...

📖 Read original article


130. Offline-to-Online Creative Optimization with Generative Models and Adaptive Testing ​

Author: Kevin Lee, Benjamin Letham, Zhiyuan Jerry Lin, Elodie Samson, Eric Onofrey, Poppy Zhang, Shawndra Hill, Eytan Bakshy
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.23696v1 Announce Type: new Abstract: Ad creative optimization is increasingly constrained by evaluation rather than generation. Generative models can produce many plausible creatives, but reliable evaluation requires online experiments, in which only a limited slate can be tested. We stud...

📖 Read original article


131. Offline-Online Curriculum RL for Multimodal Reasoning ​

Author: Wendi Deng, Hang Du, Guoshun Nan, Haokun Tian, Jiaqi Yu, Xinlei Cao, Jaile Li, Jingfeng Chen, Ling Deng, Ting Li, Hao Yang, Jun Liu, Xudong Jiang, Sicong Leng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23700v1 Announce Type: new Abstract: Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers. This behavior undermines interpretability and reliability, suggesting reliance on spurious short...

📖 Read original article


132. E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios ​

Author: Weihuang Zheng, Tianyuan Zou, Eileen Ye, Alphet Liu, Youyong Kong, Ya-Qin Zhang, Duran Zheng, Maxm Pan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23722v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, and committing state changes. We refer to this capability as multi-step t...

📖 Read original article


133. Training Language Models to Cooperate with Inference-Time Controllers ​

Author: Moumita Choudhury, Vanshaj Khattar, Jing Liu, Toshiaki Koike-Akino, Ankush Chakrabarty, Shlomo Zilberstein, Ye Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23771v1 Announce Type: new Abstract: Large language model (LLM) performance increasingly depends not only on the base model, but also on the inference-time controller used to organize reasoning. Existing post-training methods, however, typically optimize for a single fixed interaction pat...

📖 Read original article


134. From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement ​

Author: Qinsi Wang, Jing Shi, Huazheng Wang, Kun Wan, Yiran Wu, Bo Liu, Qingyun Wu, Hai Helen Li, Yiran Chen, Handong Zhao, Wentian Zhao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23802v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics ...

📖 Read original article


135. ACM: Agentic Context Management for Long Horizon Tasks ​

Author: Xiaochuan Li, Ryan Ming, Meng Chu, Shuai Shao, Rong Jin, Chenyan Xiong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23809v1 Announce Type: new Abstract: Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment. Existing context compression methods inevitably incur information loss and are triggered by rigid heuristic rules, leav...

📖 Read original article


136. Do Visual Features Improve Other-Initiated Repair Detection? A Dyadic Multimodal Approach ​

Author: Anh Ngo, Nicolas Rollet, Catherine Pelachaud, Chlo'e Clavel
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23845v1 Announce Type: new Abstract: Other-initiated Self-repair, or in short Other-initiated Repair (OIR), is an essential mechanism in conversational interaction, whereby a recipient signals a problem in speaking, hearing, or understanding, prompting the previous speaker to resolve it. ...

📖 Read original article


Author: Haijiang Yan, Jian-Qiao Zhu, Liqiang Huang, Ming Meng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23854v1 Announce Type: new Abstract: Humans often find good solutions to combinatorial optimization problems that are computationally hard even for advanced computer algorithms. In the Euclidean traveling salesman problems (TSP), people rapidly produce tours that are near-optimal, despite...

📖 Read original article


138. Cost-Aware Recovery-Pathway Identification and Bayesian Optimization for Autonomous Materials Discovery ​

Author: Debajyoti Ray, Niranjan Srinivas
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2607.23896v1 Announce Type: new Abstract: Autonomous laboratories automate experimental execution, but a campaign must also decide which recovery pathway merits optimization. We formulate this as a sequential decision problem with a discrete pathway-identification stage and a continuous within...

📖 Read original article


139. GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models ​

Author: Jun Ling, Tao Huang, Junzhuo Liu, Bowen Tang, Peng Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23913v1 Announce Type: new Abstract: Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substantially increase downstream language-model inference cost. Existing token-reduction methods assess tok...

📖 Read original article


140. Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory ​

Author: Saurabh Ranjan, Konstantina Sokratous, Brian Odegaard
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY, q-bio.NC

arXiv:2607.23927v1 Announce Type: new Abstract: A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabula...

📖 Read original article


141. MemTX: Transactional Belief Commit for Stateful Agent Memory ​

Author: Xiaoyang Li, Yiqi Wang, Haohui Lu, Zhi Chen, Mo Li, Pingan Song, Mingkai Zheng, Taotao Cai
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23929v2 Announce Type: new Abstract: LLM agents increasingly coordinate through persistent shared memory: one agent's write becomes another agent's premise, and eventually a tool call with real side effects. Current agent memory systems treat every accepted write as immediately actionable...

📖 Read original article


142. From Cognitive Architectures to Language Agents: A Mechanism-Level Review of Lineage, Convergence, and Migration Gaps ​

Author: Haodi Fan, Zucong Lan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23942v1 Announce Type: new Abstract: Memory, planning, reflection, and tool use are often compared as feature labels, obscuring the control semantics that determine how an agent actually runs. This review connects ten historical cognitive architectures, eight language-agent runtime famili...

📖 Read original article


143. DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models ​

Author: Hao Yang, Jin Wang, Xuejie Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23944v1 Announce Type: new Abstract: Human visual reasoning typically follows a coarse-to-fine attention process, starting from global scene understanding and gradually focusing on question-relevant regions. However, multimodal large language models may deviate from this pattern due to at...

📖 Read original article


144. EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff ​

Author: Xiao Ma, Zhiquan Hu, Yi Wei, Chenchen Zhao, Yijun Chen, Jicheng Zhao, Yuming Li, Chuang Dai
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23955v2 Announce Type: new Abstract: Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups provide no comparative signal and may hide useful search behavior. We present EviBack, an evidence- constrained ...

📖 Read original article


145. Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries ​

Author: Taeyoung Kim
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.23967v1 Announce Type: new Abstract: Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time relaxation mechanism for grokking in linear models trained with full-batch heavy-ball optimization and weight de...

📖 Read original article


146. Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks ​

Author: Stefan G. Creadore
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, q-bio.QM

arXiv:2607.23975v1 Announce Type: new Abstract: Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does not establish scientific validity. We developed Plato-Bio, a biology-routed extension of the open Plato/Denario a...

📖 Read original article


147. Exploring Budgeted Image Classification with Content-Sensitive Resource Allocation ​

Author: Athanasios G. Papadopoulos
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23997v1 Announce Type: new Abstract: The ever-growing adoption of Artificial Intelligence (AI) creates the need to deploy Deep Neural Networks in a variety of computational environments. We consider dynamic environments, where computational requirements are subject to change, and we pose ...

📖 Read original article


148. Self-Supervised Consistency Enhanced Disentangled Learning for Neural Decoding Generalization in Brain-Machine Interface ​

Author: Jiyu Wei, Di Hong, Zhanjie Zhang, Dazhong Rong, Qinming He, Yueming Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.RO, eess.SP

arXiv:2607.24023v1 Announce Type: new Abstract: Brain-Machine Interfaces (BMIs) provide a direct communication pathway between the brain and external devices, enabling humans to control assistive and robotic technologies, with potential applications in rehabilitation, human motor augmentation, and h...

📖 Read original article


149. A Cyclic Adaptation-Generalization Framework with Uncertainty-Guided Self-Paced Learning for Long-Term Brain-Machine Interfaces ​

Author: Jiyu Wei, Di Hong, Zhanjie Zhang, Dazhong Rong, Qinming He, Yueming Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.RO, eess.SP

arXiv:2607.24031v1 Announce Type: new Abstract: Brain-Machine Interfaces (BMIs), which link the brain to external devices, hold great potential in rehabilitation, human performance augmentation, and human-centered robotics. However, invasive BMIs face a critical challenge for long-term deployment du...

📖 Read original article


150. The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research ​

Author: Carlo Iacono (Charles Sturt University, Australia)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked purposes. First, it audits a maximum-variation purposive corpus of 40 empirical records appearing b...

📖 Read original article


151. Quantum-Inspired Evolutionary Neighborhood Search for Arrival-Departure Track Utilization Adjustment under Short-Term Disturbances ​

Author: Xiaobin Li, Wuming Lei, Yanbin Gao, Weiguang Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24049v1 Announce Type: new Abstract: Short-term disturbances at major passenger railway stations alter train arrival and departure times as well as the release sequence of station resources. Effective recovery therefore requires coordinated adjustment of arrival-departure track allocation...

📖 Read original article


152. Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation ​

Author: Jingkun Luo, Da-Tian Peng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24054v1 Announce Type: new Abstract: A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes intended reasoning from answer acquisition. Outcome evidence and exposure detection do not establish ...

📖 Read original article


153. The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards ​

Author: Keyu Li, Jin Gao, Dequan Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24063v1 Announce Type: new Abstract: On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is toward how much compute that factuality costs. Static leaderboards score factuality in isolation and tr...

📖 Read original article


154. MiSS: A Logic-Driven Explanation of Minimal Sufficient Coalitions for Point Cloud Classifiers ​

Author: Mengda Xing (UA, CRIL), Jean-Marie Lagniez (UA, CRIL)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24074v1 Announce Type: new Abstract: We present MiSS, a black-box, query-based framework for explaining 3D point cloud classifiers through perturbation-relative sufficiency reasoning. MiSS treats a superpoint partition as an interpretable abstraction layer and asks whether the original pr...

📖 Read original article


155. Towards High-Level Semantic Intelligence ​

Author: Xiujie Song, Gefei Yang, Yining You, Jiahui Gan, Qi Jia, Shota Watanabe, Tianxi Wan, Mengyue Wu, Kai Yu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24082v1 Announce Type: new Abstract: Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of AI reveals a clear trajectory from simple to complex semantic processing. While early AI systems...

📖 Read original article


156. MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents ​

Author: Yiwen Ma, Songjun Tu, Qichao Zhang, Dong Li, Linjing Li, Dongbin Zhao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24097v1 Announce Type: new Abstract: Memory-augmented LLM agents typically answer queries by retrieving relevant memories and feeding them directly to an answer model. This retrieval-as-evidence paradigm assumes retrieved memories are already suitable for reasoning, leaving the answer mod...

📖 Read original article


157. Scaling GUI Agents with Visual State Transitions ​

Author: Xiangyan Liu, Kaixin Li, Haonan Wang, Biao Wu, Meng Fang, Longxu Dou, Chao Du, Michael Qizhe Shieh, Tianyu Pang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24112v1 Announce Type: new Abstract: We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal model on visual state transitions by jointly optimizing inverse dynamics (predicting actions from s...

📖 Read original article


158. Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems ​

Author: Ali Zahid Raja
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.24117v1 Announce Type: new Abstract: Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source...

📖 Read original article


159. A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference ​

Author: Zhuoran Song, Haozhe Jiang, Chunyu Qi, Minnan Pei, Gang Li, Xiaoyao Liang, Haibing Guan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24148v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators, such as Dadu-Corki, improve efficiency but treat VLA models as full-pre...

📖 Read original article


160. Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness ​

Author: Yang Li, Hai Liu, Dian Shao, Yu Wang, Xiyu Chen, Sergey Volkov, Bozhi Wang, Ziyu Sun, Sihang Liu, Ye Luo, Xiaowei Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.24162v1 Announce Type: new Abstract: Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation budgets. Existing approaches - heuristic search, black-box optimization...

📖 Read original article


161. Falsifiable Commitment Planning for Self-Correcting Web Agents ​

Author: Guangyi Liu, Huan Zhao, Quanming Yao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24167v1 Announce Type: new Abstract: Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, reused skill, or plan assumption no longer supports the user instruction. Existing agents can plan, reflect, or reu...

📖 Read original article


162. Myopia Prevention and Control 3.0: Artificial Intelligence--Driven Risk Stratification, Proactive Monitoring, and Personalized Intervention ​

Author: Tieniu Wang, Cangzhu Huang, Qianhui Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24187v1 Announce Type: new Abstract: The convergence of artificial intelligence (AI), digital sensing, and ubiquitous computing has created an unprecedented opportunity to transform myopia prevention from a reactive, population-based model into a proactive, precision-driven one. Despite e...

📖 Read original article


163. Integrating Factual and Normative Industrial Knowledge via Constraint-Aware Graph Attention for Process Plan Recommendation ​

Author: Yuntong Chen, Yingqi Li, Yingying Xiao, Ziang Wang, Zewei Liu, Jiahao Liu, Xitian Tian, Lijiang Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2607.24213v1 Announce Type: new Abstract: Integrating heterogeneous industrial knowledge, including factual relations and decision constraints, remains a core challenge in industrial information systems. Machining process planning exemplifies this problem because engineers must select operatio...

📖 Read original article


164. Epistemic Norms for AI Safety and Alignment Research ​

Author: Keivan Navaie
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24243v1 Announce Type: new Abstract: Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence...

📖 Read original article


165. Generative Artificial Intelligence (GenAI) to convert images of queuing networks into verifiable simulation models: an open-weight LLM workflow approach ​

Author: Thomas Monks, Alison Harper, Amy Heather, Navonil Mustafee
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24259v1 Announce Type: new Abstract: Recent work has explored the use of Large Language Models (LLMs) to automate simulation model building, typically by generating executable code directly from natural language descriptions. However, this raises challenges for verification and reproducib...

📖 Read original article


Author: Junlin Liu, Jiangwang Chen, Zixin Song, Shuaiyu Zhou, Chunji Lv, Hank Wu, Kailin Jiang, Jinyang Wu, Bohan Yu, Chenxi Zhou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24280v1 Announce Type: new Abstract: Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distilla...

📖 Read original article


167. Unequal Trips, Unequal Places: Diagnosing and Mitigating Delay Inequity in Autonomous Vehicle Fleet Coordination ​

Author: Nicole Hu, Mingtao Zhang, Haoyang LI, Chen Jason Zhang, Li Qing
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24336v1 Announce Type: new Abstract: City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is distributed across trips and regions. We conduct a distributional audit on three real-city road-network and taxi...

📖 Read original article


168. Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families ​

Author: Dushyant Sharma
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.24339v1 Announce Type: new Abstract: Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, perseveration when stuck. These are failures of propensity, not capability; they concern what a model does under sustained...

📖 Read original article


169. Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Age ​

Author: Weijie Xia, Stefanie Horian, Hanyue Huang, Queena K. Qian, Jie Yang, Pedro P. Vergara Barrios
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24341v2 Announce Type: new Abstract: Recent studies use Large language models (LLMs) to simulate human opinions and decisions by prompting models with demographic, attitudinal, or persona-based descriptions. Yet such simulations rarely model the practical, cognitive, or social frictions t...

📖 Read original article


170. Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization ​

Author: Haoyue Liu, Xiaoyu Ma, Ye Chen, Yuexian Zou, Xiaoying Tang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24354v1 Announce Type: new Abstract: Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding promising results. However, on multimodal tasks, the effectiveness of APO is fundamentally bottlenec...

📖 Read original article


171. Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers ​

Author: Jinliang Deng, Yiming Niu, Yibo Pan, Zhiqi Shao, Qin Luo, Yongxin Tong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24419v1 Announce Type: new Abstract: Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect failures and iteratively revise classifier designs. Recent LLM-based agents have demonstrated the potential for au...

📖 Read original article


172. From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis ​

Author: Liwei Dong, Jiahao Zhao, Nan Xu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24459v2 Announce Type: new Abstract: Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems. We study scientific-computing experience consolidation: converting verified runtime ...

📖 Read original article


173. Making Mathematical Knowledge Explainable, Accessible and Interoperable Through Large Language Model Integration ​

Author: Jan Range, Bj"orn Schembera, Dominik G"oddeke
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.DL

arXiv:2607.24512v1 Announce Type: new Abstract: Mathematical models are central to formalizing research problems, yet their documentation often falls short of FAIR principles. Knowledge bases such as the Mathematical Model Database (MathModDB) address this gap by providing curated, semantically rich...

📖 Read original article


174. Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis ​

Author: Tianqiao Zhao, Meng Yue, Jianhui Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY

arXiv:2607.24539v1 Announce Type: new Abstract: Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framework in order to condu...

📖 Read original article


Author: G{'e}nesis Montenegro (WIMMICS), Mokhtar Boumedyen Billami (WIMMICS), Catherine Faron (WIMMICS), Fabien Gandon (WIMMICS), Pierre Monnin (WIMMICS)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24551v1 Announce Type: new Abstract: Maintenance regulations are complex legal texts that are difficult to exploit when addressing a specific case and challenging to integrate into operational systems. This paper presents a two-stage LLM-assisted workflow for French maintenance regulation...

📖 Read original article


176. Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models ​

Author: Murilo Salem, Lu'isa B"ohm, Daniel Pontes, Anderson Ferrugem
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24562v1 Announce Type: new Abstract: Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style. Conformal risk control (CRC) gives rigorous marginal risk guarantees for selective prediction with abstention, but marginal guarantees d...

📖 Read original article


177. TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs ​

Author: Federico Valletta, Giacomo Longo, Enrico Russo, Alessio Merlo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2607.24563v1 Announce Type: new Abstract: Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outputs remain fallible and are often stored without the evidence, provenance, and validation history needed to decid...

📖 Read original article


178. DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing ​

Author: Tobias J. Bauer, Christian Riess, Daniel Loebenberger, Christian Bergler
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.IR

arXiv:2607.24567v1 Announce Type: new Abstract: Semantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensional data spaces have gained extensive consideration in recent years. Deep learning-based methods offer better seman...

📖 Read original article


179. LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports ​

Author: Jonas Schr"oder, Jonas Schweisthal, Oliver M"uller, Markus Weinmann, Stefan Feuerriegel
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24573v1 Announce Type: new Abstract: Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and th...

📖 Read original article


180. SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents ​

Author: Hang Ni, Weijia Zhang, Fan Liu, Mengqian Lu, Hao Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24588v1 Announce Type: new Abstract: Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout th...

📖 Read original article


181. Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future Directions ​

Author: Zhimin Zhang, Chengzhen Ma, Jia Chai, Rongxin Zhan, Huansheng Ning, Lingfeng Mao, Dan Zhang, Suiping Jiang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24589v1 Announce Type: new Abstract: The development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared achievements. The rise of Artificial Intelligence (AI) has significantly accelerated the global processes of digiti...

📖 Read original article


182. Efficiency Matters in Autonomous Research ​

Author: Haiqian Yang, Yuan Cao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.24647v1 Announce Type: new Abstract: AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome. In this paper, we argue that the efficiency of t...

📖 Read original article


183. Reason-Mediated Behavioral Models for Auditing LLM Social Simulators ​

Author: Atharva Pandey, Gautam Jajoo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24649v1 Announce Type: new Abstract: Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can match th...

📖 Read original article


184. Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating ​

Author: Maruthi Vemula, Neeraj Praneeth Gajula
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24667v1 Announce Type: new Abstract: A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choic...

📖 Read original article


185. ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams ​

Author: Ali Ansari, Yasmin Mohammadi, Farnoush Nili, Parsa Esmaeilkhani, Longin Jan Latecki, Eduard Dragut
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.DB

arXiv:2607.24707v1 Announce Type: new Abstract: Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ERUnderstand, the firs...

📖 Read original article


186. Creative Integration: A Decidable Criterion of Creativity ​

Author: Yoshinori Nomura
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.13977v1 Announce Type: cross Abstract: "Integrative" solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes the world cheaper to describe -- from a tidy re-description. Building on the lineage that treats creativity and ...

📖 Read original article


187. Evaluating Large Language Models for Symbolic Security Protocol Analysis ​

Author: Paolo Modesti, Syed Ahmed, Ioannis Sfyrakis, Derek Enodolomwanyi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.20712v1 Announce Type: cross Abstract: Security protocol verification relies on formal tools such as ProVerif and OFMC. This study evaluates whether Large Language Models (LLMs) can perform comparable analysis. We test GPT and DeepSeek in chat and reasoning modes over three runs on 130 ob...

📖 Read original article


188. Comparing Optimization Models for Radiotherapy Scheduling ​

Author: C. C. Rambaldi Migliore, D. Stanicel, N. Musliu, G. Iacca, M. Roveri
Published: 7/28/2026, 4:00:00 AM
Categories: math.OC, cs.AI

arXiv:2607.22539v1 Announce Type: cross Abstract: The Radiotherapy Scheduling Problem (RTSP) involves determining an optimal schedule for patients undergoing radiation treatments, a task that has a massive impact on clinical outcomes given the central role of radiotherapy in cancer care. The daily b...

📖 Read original article


189. Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B ​

Author: Tejasvi C. Addagada
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CR

arXiv:2607.22545v1 Announce Type: cross Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail addresses in a singl...

📖 Read original article


190. Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review ​

Author: Haowen Li, Yoichi Ishibashi, Masafumi Oyamada
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.22553v1 Announce Type: cross Abstract: Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and revi...

📖 Read original article


191. A Formal Kinetic Theory for Zeroth-Order Newton Dynamics:Stein-Corrected Hessian Estimation and Curvature--Variance Trade-offs ​

Author: Shihao Ji, Mingyu Li, Zihui Song
Published: 7/28/2026, 4:00:00 AM
Categories: math.OC, cs.AI

arXiv:2607.22567v1 Announce Type: cross Abstract: Zeroth-order Newton-type methods are useful when gradients and Hessians are unavailable, but they behave quite differently from first-order gradient-free methods. We develop a kinetic framework for algorithms that estimate both gradient and Hessian f...

📖 Read original article


192. A didactical-driven teacher assistant for a dimensional modeling course ​

Author: Laurent Brisson (IMT Atlantique - DSD), Maria Segarra (IMT Atlantique - INFO, Lab-STICC_MOTEL), Gr'egory Smits (IMT Atlantique - INFO, Lab-STICC_MOTEL)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2607.22598v1 Announce Type: cross Abstract: Educational chatbots powered by large language models (LLMs) show promising effects on learning outcomes, yet most systems delegate pedagogical decisions such as content selection and didactic structuring implicitly to the LLM, making tutoring strate...

📖 Read original article


193. Quotient Tree Arithmetic: Deferred-Division Computation with Bounded Symbolic Depth and Cross-Subtree Cancellation ​

Author: Gregory Magarshak
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SC, cs.AI, cs.LG, cs.MS

arXiv:2607.22612v1 Announce Type: cross Abstract: We introduce Quotient Tree Arithmetic (QTA), a computational substrate in which values are represented as deferred quotient pairs (N, D) whose ratio is evaluated lazily at a designated materialization boundary. The framework applies to any domain: IE...

📖 Read original article


194. Revitalizing Public Urban Places through Cultural and Political Memory: A Technological Approach with LLMs and Augmented Reality ​

Author: Lara Vartziotis, Tina Vartziotis, Valentin Keckeisen, Frank Beutenmueller, Martin Obstbaum, Sotirios Kotsopoulos, Kostas Moraitis
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL

arXiv:2607.22613v1 Announce Type: cross Abstract: This paper explores the intersection of memory, place, and identity, examining how new technologies, particularly Apple Vision Pro, can illuminate this nexus. Leveraging digital twins and virtual reality, it investigates how memory is woven into land...

📖 Read original article


195. Masked Autoencoders Learn Perception-Relevant Representations from Resting State Neural Data ​

Author: Aleksandr Kovalev, Antonio Lozano, Fabrizio Grani, Cristina Soto Sanchez, Leili Soo, Roc'io L'opez-Peco, Adrian Villamarin-Ortiz, Roberto Moroll'on Ruiz, Mar'ia del Mar Ayuso Arroyave, Alfonso Rodil, Eduardo Fern'andez
Published: 7/28/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.LG, cs.NE

arXiv:2607.22615v1 Announce Type: cross Abstract: Clinical neuroprosthetics face a data bottleneck: labeled perception trials are scarce while hours of spontaneous neural activity are largely underutilized. Here, we test whether self-supervised learning can use these unlabeled datasets to improve pe...

📖 Read original article


196. Learning When to Reason for Text-to-SQL via SFT and DPO ​

Author: Soohyuk Jang, Jiheum Yeom, Nohil Park, Sang Hun Kim, Yoonyoung Choi, Kiwook Bae, Sungroh Yoon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.22622v1 Announce Type: cross Abstract: Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmarks at the cost of high inference-time overhead. However, a large fraction of real-world queries are ...

📖 Read original article


197. AI-Assisted Causal Inference and Mediation Analyses of Environmental and Psychosocial Determinants of Subjective Cognitive Difficulties in the All of Us Research Program ​

Author: Cong Cao, Shuangge Ma
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG

arXiv:2607.22640v1 Announce Type: cross Abstract: Short-term environmental exposures have been linked to cognitive and behavioral outcomes, although many reported associations may reflect broader geographic and contextual differences. Using longitudinal data from the All of Us Research Program (2018...

📖 Read original article


Author: Ahmed Abolfadl, Marwa Mahmoud Abla, Mervat Abu-Elkheir, Maggie Mashaly
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.LG

arXiv:2607.22641v1 Announce Type: cross Abstract: Predicting emerging trends is vital for businesses, researchers, and policymakers; yet traditional approaches often lack scalability and adaptability. This paper presents a trend prediction framework based on Automated Machine Learning (AutoML), desi...

📖 Read original article


199. Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS ​

Author: Viktoriia Makovska, George Fletcher
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.22657v1 Announce Type: cross Abstract: Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing machine-unlearning algorithms can suppress this behavior. We introduce Level-based Evaluation of Na...

📖 Read original article


200. SetGo: Metadata Readiness for Scientific AI Datasets ​

Author: Sean R. Wilkinson, Polina Shpilker, Wesley Brewer
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DL, cs.AI

arXiv:2607.22677v1 Announce Type: cross Abstract: Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corre...

📖 Read original article


201. Towards Nexus-Score: Metadata Gaps Limit Scholarly AI Attribution ​

Author: Aadi Narayana Varma Dantuluri, Sushrut Thorat, Paras Chopra
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.IR

arXiv:2607.22684v1 Announce Type: cross Abstract: Artificial intelligence systems increasingly mediate how science is found and credited. We asked whether missing metadata prevents AI systems from crediting work. As a boundary test, an AI system citing without access to task-relevant paper lists oft...

📖 Read original article


202. MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion ​

Author: Jiacheng Liu, Jason Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22696v1 Announce Type: cross Abstract: High-resolution video diffusion models built on Diffusion Transformers (DiTs) deliver strong fidelity but quickly exhaust the memory budget of a single workstation. A 100 billion-plus parameter DiT easily requires over a terabyte of persistent state,...

📖 Read original article


203. Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model ​

Author: Xiao Wang, Hao Si, Qiang Chen, Yu-Xiang Zhang, Beihe Zhang, Jianhua Yang, Qingquan Yang, Dengdi Sun, Wanli Lyu, Guosheng Xu, Jin Tang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.22704v1 Announce Type: cross Abstract: Nuclear fusion has made significant progress in recent years and is expected to become one of the most important pathways to addressing global energy challenges. This paper focuses on observing plasma using visible-light cameras, analyzing its spatio...

📖 Read original article


204. RMS@CC-MMD 2026: Multimodal Misogyny Detection via Geometric Interaction and Multi-View Consensus ​

Author: Md. Ajwad Hossain
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2607.22709v1 Announce Type: cross Abstract: The proliferation of internet memes has introduced new complexities to automated content moderation, particularly in detecting misogyny. Memes often rely on a semantic clash between visual and textual modalities, where hateful intent is implicit and ...

📖 Read original article


205. CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents ​

Author: Mingwei Zheng, David OBrien, Siwei Cui, Pardis Pashakhanloo, Rajdeep Mukherjee, Myeongsoo Kim, Sachit Kuhar
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SE

arXiv:2607.22711v1 Announce Type: cross Abstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. However, the conventional append-only trajectory architecture found in practice tightly couples file-read ...

📖 Read original article


206. scMIR: a vision-language foundation model for single-cell light microscopy image representation ​

Author: Yifan Shang, Jiahui Tan, Xiangxiang Zeng, Renjie Zhou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, physics.optics

arXiv:2607.22712v2 Announce Type: cross Abstract: Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heterogeneity pose challenges to high-throughput automated analysis. Existing representation learning methods mostly...

📖 Read original article


207. Real-Time Semantic Segmentation with Optimized RetinaNet Architectures for Embedded Automotive Systems ​

Author: Sai Sidharth D
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22714v1 Announce Type: cross Abstract: Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded automotive platforms impose severe constraints on compute, memory, and power. This paper presents an optimized sema...

📖 Read original article


208. DAMamba-UNet3D: A Parameter-Efficient Mamba State Space U-Net with Dynamic Adaptive Scan for 3D Medical Image Segmentation ​

Author: Mohammad Arafat Hussain, Ellen Grant, Yangming Ou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22718v1 Announce Type: cross Abstract: We propose parameter-efficient SSM-based U-Net architectures for 3D medical image segmentation. Convolutional U-Nets afford O(n) local mixing per layer but lack explicit global context; transformers provide global reasoning at O(n^2) cost in sequence...

📖 Read original article


209. An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia ​

Author: Nassira Ait Mehdi, Milissa Temmam, Slimane Larabi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22721v1 Announce Type: cross Abstract: Cognitive remediation tasks often require patients to perform structured actions involving object manipulation and sequential reasoning. For patients diagnosed with schizophrenia, these tasks are crucial for addressing severe cognitive deficits. Howe...

📖 Read original article


210. A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection ​

Author: Ali Borji
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.22722v1 Announce Type: cross Abstract: Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible perturbation that causes the model to keep its original, correct prediction, even though a human would no longer...

📖 Read original article


211. Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks ​

Author: Kaibing Yang, Guangfeng Cai, Shengtian Yang, Shuo He, Yu Li, Mengyi Liu, Pengwei Chen, Jun Xu, Lei Feng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22724v1 Announce Type: cross Abstract: Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group. However, on difficult long-horizon tasks, this comparison can suffer ...

📖 Read original article


212. Structural Preservation Governs Data Augmentation in Deep Learning-Based Laser Speckle Material Classification ​

Author: Mohamed Abdallah Salem, Nourhan Zein Diab
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.22725v1 Announce Type: cross Abstract: Data augmentation is routinely used to improve generalization in image classification, but the assumptions underlying standard policies are poorly matched to coherent imaging. Laser speckle patterns are not generic textures; they arise from coherent ...

📖 Read original article


213. Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features ​

Author: Dzmitry Malyshau
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.22739v1 Announce Type: cross Abstract: We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcement learning or explicit memory. Cortex is a compact Quake policy with 10.98 million trainable parameters in a ...

📖 Read original article


214. QFedPolyp: A Communication- and Inference-Efficient Federated Learning Framework for Polyp Segmentation ​

Author: Madan Baduwal, Priyanka Paudel
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2607.22743v1 Announce Type: cross Abstract: Background and Objective: Automatic polyp segmentation supports computer-aided diagnosis and early colorectal cancer detec- tion. Centralized deep learning requires hospitals to share sensitive medical data, while federated learning preserves privacy...

📖 Read original article


215. AI-generated Images Challenge Visual Trust in High-risk Scenarios ​

Author: Yi-Zhi Wang, Yichen Xiao, Linan Yue, Weibo Gao, Yichao Du, Pengfei Fang, Shimin Di, Min-Ling Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22745v1 Announce Type: cross Abstract: Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and personal reputation. Yet existing detection benchmarks rarely examine synthetic images in public- and i...

📖 Read original article


216. Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge ​

Author: Hongruixuan Chen, He Huang, Haifeng Wang, Jian Song, Junjue Wang, Weihao Xuan, Hamish Mitchell, Jiepan Li, Wei He, Liangpei Zhang, Zijie Wang, Chen Zhong, Jiazhen Zhao, Lei Hu, Ting Hu, Hongyan Zhang, Gregory Angelides, Miriam Cha, Clifford Broni-Bediako, Junshi Xia, Taylor Perron, Naoto Yokoya
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, eess.IV

arXiv:2607.22746v1 Announce Type: cross Abstract: Rapid post-disaster response requires timely, building-level information on whether structures remain intact, are damaged, or are destroyed. Post-event optical imagery, however, may be unavailable because of cloud, smoke, or darkness. The Bright Chal...

📖 Read original article


217. Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence ​

Author: Zhaowen Fan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE

arXiv:2607.22748v1 Announce Type: cross Abstract: Modern neural networks primarily adapt through parameter modification within predefined computational structures. While recent methods introduce modularity, conditional computation, and parameter-efficient adaptation, they generally do not distinguis...

📖 Read original article


218. Post-Operative Glioma Segmentation via Loss Stabilization, Normalization and Subspace Attention ​

Author: Alexandru Cri\c{s}an, Diana Borza
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22749v1 Announce Type: cross Abstract: Tracking residual tumor after surgery is essential for catching recurrence early, but automating post-operative glioma segmentation remains a difficult task. Although transformer-based architectures, such as SwinUNETR, achieved impressive results, fe...

📖 Read original article


219. Real-time Reconstruction of Human Visual Perception from fMRI ​

Author: Rishab S. Iyer, Jiaxin Cindy Tu, Cesar Kadir Torrico Villanueva, Anish Mahishi, Ross P. Kempner, Jacob S. Prince, Ernest W. Lo, Akash Bhowmick, Hritik Arasu, Amaar Chughtai, Elizabeth A. McDevitt, Paul S. Scotti, Kenneth A. Norman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, q-bio.NC

arXiv:2607.22753v1 Announce Type: cross Abstract: Real-time closed-loop neurofeedback based on functional magnetic resonance imaging (fMRI) has led to important scientific and clinical advances. However, the sophistication of the analysis methods used in real-time fMRI lags behind the state-of-the-a...

📖 Read original article


220. Hierarchical Grading in Large Language Models ​

Author: T. Shaska
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22757v1 Announce Type: cross Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, and the training obje...

📖 Read original article


221. Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks ​

Author: Amritesh Banerjee
Published: 7/28/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2607.22758v1 Announce Type: cross Abstract: The integration of iterative LLMs within multi-agent diagnostic frameworks requires a rigorous quantitative reevaluation of underlying communication topologies. Frequently used architectural paradigms depend on scale-free or small-world networks, ass...

📖 Read original article


222. Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation ​

Author: Yunting Song, Matthew Watson, Peter Grabowski, Jun Qin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumulate hidden structural contradictions, safety risks, and systemic human ...

📖 Read original article


223. DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning ​

Author: He Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22769v1 Announce Type: cross Abstract: The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data. Existing dynamic data scheduling methods face critical limitations in industrial-scale pretraining and supervised fine...

📖 Read original article


224. Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models ​

Author: Renjie Liang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22771v1 Announce Type: cross Abstract: Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over many candidates. There are several encoders, several ways to compress their tokens, and several token...

📖 Read original article


225. LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning ​

Author: Chen Wang, Boming Kang, Qinghua Cui
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22777v1 Announce Type: cross Abstract: Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along amino-acid sequences, their training objectives do not explicitly constrain the model to learn three-dimensional ...

📖 Read original article


226. Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control ​

Author: Federico Del Pup, Elisa Tentori, Manfredo Atzori
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22779v1 Announce Type: cross Abstract: Hand gesture recognition via surface electromyography (sEMG) is fundamental to prosthetic control. In this field, deep learning approaches have become the gold standard. However, current architectures struggle to scale; model performance typically de...

📖 Read original article


227. What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation ​

Author: Minwoo Yu, Young-guk Ha
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22781v1 Announce Type: cross Abstract: High task performance does not show whether a model retains prediction-relevant structural information in its internal representation. Temporal graph models, for example, can achieve high future-link AUC while basic graph statistics remain difficult ...

📖 Read original article


228. Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs ​

Author: Ilia Sobakinskikh, Paul Alexander Bilokon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR, cs.DC, cs.PF, stat.CO

arXiv:2607.22786v1 Announce Type: cross Abstract: In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series. The financial time series are price series such as asset prices. U...

📖 Read original article


229. FMOPF: Latent Flow Matching with Constraint-Aware Interaction Priors for AC Optimal Power Flow ​

Author: Zhilin Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2607.22788v1 Announce Type: cross Abstract: AC optimal power flow determines the minimum-cost generation dispatch under nonlinear power balance constraints and is solved thousands of times daily in electricity market operations. Learning a direct mapping from load conditions to OPF solutions c...

📖 Read original article


230. Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial Training ​

Author: Ali Tabaraei, Federico Simonetta, Stavros Ntalampiras
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.SD

arXiv:2607.22794v1 Announce Type: cross Abstract: Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arising from inter-speaker variability. To address this critical issue, we present the first patient-independent mul...

📖 Read original article


231. Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis ​

Author: Yuntong Chen, Jianyu Liu, Guobin Zhao, Ziang Wang, Chao Chen, Ju Huang, Xitian Tian, Lijiang Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22797v1 Announce Type: cross Abstract: Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against physical reality before it is acted upon. Current intelligent fault diagnosers fail this standard in ...

📖 Read original article


232. LithoFormer: A Robust Framework for Stratigraphic Inference via Transformers ​

Author: Shwetha Salimath, Francesca Bugiotti, Sylvain Wlodarczyk, Sohaib Ouzineb
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22804v1 Announce Type: cross Abstract: Accurate geological characterization of subsurface reservoirs from well log data is essential to support projects such as carbon capture and storage (CCS), geothermal development, and extraction of natural resources. Existing automated techniques for...

📖 Read original article


233. OrchNAS: Orchestrated Neural Architecture Search Service for Personalised Federated Edge Intelligence ​

Author: Keya Patel, Sajib Mistry, Sheik Mohammad Mostakim Fattah, Aneesh Krishna
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22805v1 Announce Type: cross Abstract: We propose OrchNAS, an energy-aware, personalised, federated edge intelligence framework that leverages a Neural Architecture Search Service to automatically design service-adaptive models for heterogeneous edge environments. The framework orchestrat...

📖 Read original article


234. Hybrid Semantic and Spectral Ensemble for Robust Synthetic Image Source Attribution ​

Author: Md. Ajwad Hossain
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.GR, eess.IV

arXiv:2607.22808v1 Announce Type: cross Abstract: The rapid advancement of text-to-image (T2I) models has necessitated robust Synthetic Image Source Attribution (SIA) methodologies. A critical challenge in SIA is the distribution shift between pristine training images and real-world deployed images,...

📖 Read original article


235. From Hybrid Mechanistic--Data-Driven Modeling Toward Neuro-Symbolic AI: What, Why, and How ​

Author: Moein E. Samadi, Andreas Schuppert
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO, stat.ML

arXiv:2607.22811v1 Announce Type: cross Abstract: Hybrid mechanistic/data-driven models, which combine first-principles with learned components, are increasingly used in process engineering and scientific machine learning. Common hybrid modeling designs are specified primarily through their architec...

📖 Read original article


236. Agentic Autoresearch for CT Reconstruction ​

Author: Andreas Maier, Lucas Kachelriess, Siming Bayer, Yixing Huang, Yan Xia, Amber Simpson, Moritz Zaiss
Published: 7/28/2026, 4:00:00 AM
Categories: physics.med-ph, cs.AI, cs.CV

arXiv:2607.22824v1 Announce Type: cross Abstract: Comparing CT reconstruction methods fairly is labor-intensive and largely manual, and many benchmarks use idealized data. We ask whether a large language model (LLM) agent can do the labor of reconstruction research on its own, and whether a ranking ...

📖 Read original article


237. Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias ​

Author: Ofek I. Cohen, Lior Shani, Aviv Rosenberg, Ankur Samanta, Tal Wagner, Yonathan Efroni
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.22837v1 Announce Type: cross Abstract: Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns around sensitive data. However, such adaptation remains non-trivial: it often requires operationall...

📖 Read original article


238. Language-Routed RAG and Direct Option Scoring for Multilingual Financial QA: DS@GT at FinMMEval ​

Author: Justice Ayela, Kabir Sahni
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2607.22841v1 Announce Type: cross Abstract: We present DS@GT's submission to FinMMEval 2026 Task 1, a multilingual financial exam question answering benchmark spanning English, Spanish, Greek, Chinese, and Hindi. Financial certification exams such as the CFA, EFPA, and CPA demand structured do...

📖 Read original article


239. Robustifying pathology foundation models via fine-tuning ​

Author: Alexandre Filiot, Oskar Thaeter, Benoit Schmauch, Lionel Guillou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22861v1 Announce Type: cross Abstract: Pathology foundation models (FMs) produce powerful tile-level representations which remain sensitive to scanner and staining variability, undermining deployment across laboratories. We develop a novel fine-tuning recipe that improves the robustness o...

📖 Read original article


240. Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests ​

Author: Patrick Rim, Tom Long, Ekta Prashnani, Ruth Rosenholtz, Ben Boudaoud, Peter Xenopoulos, Alex Wong, Joohwan Kim, Jae-Hyun Jung
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22864v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Existing benchmarks evaluate these models as black boxes, limiting their ability to identify the underlying causes...

📖 Read original article


241. AI-interpreted Optical Scattering for Robust and Focal Depth-Aware Imaging ​

Author: Eunji Ko, Patrick Ross, Corey Hart, Wolfgang Losert
Published: 7/28/2026, 4:00:00 AM
Categories: physics.optics, cs.AI, cs.LG

arXiv:2607.22867v1 Announce Type: cross Abstract: Optical scattering has conventionally been regarded as an impediment in imaging research due to the degradation of image quality during reconstruction. Nevertheless, this study explores two cases in which optical scattering may serve a beneficial rol...

📖 Read original article


Author: Tergel Molom-Ochir, Benjamin F. Morris III, Yintao He, Archit Gajjar, Giacomo Pedretti, Hai Helen Li, Yiran Chen, Jim Ignowski, Aishwarya Natarajan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.ET

arXiv:2607.22869v1 Announce Type: cross Abstract: Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limiting edge deployment. In-memory computing (IMC) is energy-efficient on regular workloads but has been considere...

📖 Read original article


243. Spatial Prediction of Soil Microplastics and Organic Matter Using Graph Attention Networks ​

Author: Anik Dev Nath, Md Al Amin, Bikash Kumar Paul
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22875v1 Announce Type: cross Abstract: Accurate estimation of soil microplastics and organic matter is essential to assess ecosystem health and support sustainable land use. This study presents a graph-based deep learning approach using Graph Attention Networks (GATs) to model spatial dep...

📖 Read original article


244. Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study) ​

Author: Junda Zhao, Shurui Zhou, Eldan Cohen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2607.22880v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have driven growing interest in using LLMs to automate test generation. Prior work commonly evaluates generated test suites using proxy metrics such as code coverage and mutation score. However, studies...

📖 Read original article


245. Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests ​

Author: Junda Zhao, Shurui Zhou, Eldan Cohen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2607.22883v1 Announce Type: cross Abstract: While Large Language Models (LLMs) show great promise for automating unit test generation, recent studies suggest that the quality of generated tests can be negatively impacted when models are prompted with buggy code. This paper presents a new metri...

📖 Read original article


246. Controlling Embedding Spaces with Text-Conditioned Transformations ​

Author: Joseph Fioresi, Fabian Caba Heilbron, Pankaj Nathani, Mubarak Shah, Kushal Kafle
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22919v1 Announce Type: cross Abstract: Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification. These embeddings compress high-level semantics into a single vector, which comes at the cost ...

📖 Read original article


247. Not All LLM Reasoning is Visible in the Chain-of-Thought ​

Author: Vatsal Baherwani, Tom Goldstein, Ashwinee Panda
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.22925v1 Announce Type: cross Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens...

📖 Read original article


248. Invariant Discovery for Networked Systems ​

Author: Hongyu H`e, Alexander Krentsel, Sylvia Ratnasamy, Maria Apostolaki
Published: 7/28/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.LG, cs.SC

arXiv:2607.22944v1 Announce Type: cross Abstract: Invariants, the relations expected to hold among measured signals of a network, underpin applications from verification to traffic generation, telemetry imputation, and input validation, yet writing them by hand demands rare expertise in both formal ...

📖 Read original article


249. Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational Excellence ​

Author: Bin Dong, Sukhada Gholba, Brooklin Gore, Shawn Kwang, David Mitchell, Samuel Oehlert, Garrett Stewart, Brendan White, Luke Baker, Ed Balas, Britt Gathright, Chin Guok, Jon-Paul Heron, John MacAuley, Scott Richmond, Chris Robb, Chris Tracy, Kesheng Wu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2607.22948v1 Announce Type: cross Abstract: The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 initiative and to address persistent operational pain points in the Network Operations Center (NOC) workflow. ESnet...

📖 Read original article


250. Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model ​

Author: Robab Aghazadeh Chakherlou, Siddartha Khastgir, Peter Popov, Xingyu Zhao
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG

arXiv:2607.22951v1 Announce Type: cross Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified operational profile. Conventional benchmark-based evaluation, often summarized by aggregate accuracy, p...

📖 Read original article


251. HALLELUAI: A Hallucination-Aware AI System for Ultra-Realistic Image-to-Video Generation at Scale ​

Author: Aniket Sakpal, Yang Jiang, Rouzbeh Davoudi, Shayan Hassantabar, Mani Najmabadi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.22959v1 Announce Type: cross Abstract: AI-generated video is increasingly used across marketing, product storytelling, and creative workflows, yet automated; high-precision quality control remains a major constraint to scaling production. We present HALLELUAI, an end-to-end system that mo...

📖 Read original article


252. Label-free Industrial Fault Detection via Adversarial Inverse Reinforcement Learning: A System for Run-to-Failure Prognostics ​

Author: Dhiraj Neupane, Mohamed Reda Bouadjenek, Richard Dazeley, Sunil Aryal
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22987v1 Announce Type: cross Abstract: Machinery fault detection (MFD) remains heavily reliant on supervised learning, which struggles with the scarcity of fault labels in real-world settings. While reinforcement learning (RL) offers a framework to model the sequential nature of degradati...

📖 Read original article


253. An Explicit Counterexample to Stanley's Rankwise Lower-Bound Conjecture for Differential Posets ​

Author: Xinan Dai, Wenhao Deng, Yingdong Shi, Tailin Wu, Yuchen Yang
Published: 7/28/2026, 4:00:00 AM
Categories: math.CO, cs.AI

arXiv:2607.22988v2 Announce Type: cross Abstract: In Problem 6 of his 1988 paper on differential posets, Stanley asked for the least possible cardinality of a fixed rank of an $r$-differential poset and suggested that the minimum should be attained by $Y^r$, the $r$-fold Cartesian power of Young's l...

📖 Read original article


254. Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning ​

Author: Xiaokun Wang, Siyu Song, Wentao Liu, Xiaodong Zou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.22996v1 Announce Type: cross Abstract: Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the opening turn instead of guiding students through progressive inquiry, as Socratic pedagogy prescribes. We present Heu...

📖 Read original article


255. Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline ​

Author: Qing Yang, Xun Wang, Ziguan Wang, Zhenjiang Li, Hongqiang Wang, Dongdong Weng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.GR, cs.LG

arXiv:2607.22997v1 Announce Type: cross Abstract: Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next major frontier for AI, echoed by industry leaders such as Jensen Huang (``the next big thing is Phy...

📖 Read original article


256. WCM: World-Cognition Model for Generalizable Human-Robot Interaction ​

Author: Yuzhen Chen, KC Zhou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.HC, cs.LG

arXiv:2607.22999v1 Announce Type: cross Abstract: Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tasks. Current robot-control paradigms, including vision-language-action policies and world-model-based planners, ...

📖 Read original article


257. Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop ​

Author: Jeff Otterson (W. P. Carey School of Business, Arizona State University)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.23002v1 Announce Type: cross Abstract: Large language models increasingly write both code and the tests meant to check it; coverage records what ran, not what was verified. We study an adversarial test-hardening loop under a mechanical oracle: a Tester model writes tests, mutation testing...

📖 Read original article


258. Exact values and exact upper bounds for families of integers with arithmetic progression intersections (Erd\H{o}s Problem #272) ​

Author: Zhanfu Yang
Published: 7/28/2026, 4:00:00 AM
Categories: math.CO, cs.AI

arXiv:2607.23004v1 Announce Type: cross Abstract: Let $t(N)$ be the largest $t$ for which there exist distinct sets $A_1,\dots,A_t \subseteq {1,\dots,N}$ such that $A_i \cap A_j$ is a nonempty arithmetic progression for all $i \neq j$ (Erdos Problem #272). Simonovits and Sos proved $t(N)=O(N^2)$ a...

📖 Read original article


259. VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy ​

Author: Xinyan Zhong, Yuwei Shi, Yuqi Wei, Chen Shen, Tianhang Zhou, Zhenghao Wu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cond-mat.mtrl-sci, cs.AI

arXiv:2607.23006v1 Announce Type: cross Abstract: Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locating the supporting evidence within those papers. Conventional retrieval-augmented generation typically addresses ...

📖 Read original article


260. All in One: Generative Modeling as Mean-Field Game Design ​

Author: Kun Zhao, Xu Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.23026v1 Announce Type: cross Abstract: Mean-field games (MFGs) offer a unifying lens on continuous-time generative modeling: a cost tuple recovering twelve prominent models---Continuous Normalizing Flows, OT-Flow, Score-based Models, Schr"{o}dinger Bridges, and more---as special cases of...

📖 Read original article


261. Multi-Agent Privacy Game in Federated Learning: A Unified Mean-Field View ​

Author: Kun Zhao, Xu Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.23029v1 Announce Type: cross Abstract: Federated learning enables collaborative model training across distributed clients without centralising their data, yet privacy remains a persistent concern because the shared model updates can leak information about local datasets. Existing privacy-...

📖 Read original article


262. MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models ​

Author: Ashitabh Misra, Madhav Agrawal, Arham Jain, Tarek Abdelzaher
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.23047v1 Announce Type: cross Abstract: Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but existing methods solve the allocation for a single fixed memory budget. In practice the budget varies across depl...

📖 Read original article


263. Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models ​

Author: Dhruvil S, Fenil Sojitra, Ravirajsinh Chauhan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.23054v1 Announce Type: cross Abstract: Multi-head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key-value pairs through a shared low-rank bottleneck (cKV), achieving 81% KV-cache reduction during inference. Despite its adoption in massive production models, no prior work h...

📖 Read original article


264. ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation ​

Author: Ahmed Haj Ahmed, Alvin Grissom II
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23058v1 Announce Type: cross Abstract: Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded reasoning. We introduce ADAGE (Analogical Difficulty-by-design Assessment...

📖 Read original article


265. Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models ​

Author: Yusuke Sakai, Natthawut Kertkeidkachorn, Kiyoaki Shirai
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23067v1 Announce Type: cross Abstract: Contrastive decoding methods such as DoLa improve the factuality of Large Language Models (LLMs) by contrasting the output distributions of mature and premature layers. However, DoLa's dynamic layer selection relies solely on divergences in output vo...

📖 Read original article


266. Traceable LLM Reasoning for Fake-Order Fraud Detection ​

Author: Siqi You, Bingsong Xu, Zhixian Zheng, Xinjian Peng, Yang Xie, Ying Wang, Jiarong Xu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2607.23075v1 Announce Type: cross Abstract: Detecting fake-order fraud at scale remains a critical challenge for large online-to-offline (O2O) service platforms, as existing approaches often rely on expert-designed features, produce black-box decisions, and provide limited interpretability. To...

📖 Read original article


267. Scoping Review of AI, Metrology, and ESG in the Semiconductor Sector: Implications for Safe and Sustainable by Design (SSbD) ​

Author: Karen Ang, Han-Teng Liao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SI, cs.AI, cs.CY, cs.SY, eess.SY

arXiv:2607.23082v1 Announce Type: cross Abstract: The semiconductor sector faces a dual transition: scaling manufacturing execution through Artificial Intelligence (AI) while satisfying stringent sustainability mandates, such as the EU Carbon Border Adjustment Mechanism (CBAM). This paper presents a...

📖 Read original article


268. Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios ​

Author: Lixun Ma, Ruolong Ma, Bei Wang, Feng Wei, Zhenguang Liu, Lorenzo Cavallaro, Wentao Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.23088v1 Announce Type: cross Abstract: Large Language Models (LLMs) are widely used for code generation, yet their security behavior in realistic development workflows remains underexplored. Existing benchmarks often rely on explicitly specified security requirements, failing to capture r...

📖 Read original article


269. KAYROS: An Anytime and Exact Open-Source Solver for Duration-Minimization Time-Dependent Vehicle Routing. A Technical Report and a Case Study in Human-AI Engineering ​

Author: Florian Rascoussier
Published: 7/28/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.MS

arXiv:2607.23116v1 Announce Type: cross Abstract: KAYROS is an open-source solver for duration-minimization time-dependent vehicle routing problems, with or without time windows (TDVRPTW, TDVRP). In these variants, travel times change with departure time, and each route's dispatch time is a decision...

📖 Read original article


270. A scalable online machine learning approach for Stock Recommendation ​

Author: Harsh Nagarkar
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CE, cs.AI, cs.DC, cs.GT

arXiv:2607.23120v1 Announce Type: cross Abstract: Stock recommendation systems face the dual challenge of adapting to rapidly changing market conditions while maintaining low-latency predictions for end users. Traditional batch-trained models fail to capture concept drift, and monolithic architectur...

📖 Read original article


271. From Vibe to Code -- and Back: Lexical Oscillation in the Formation of Design Intent with Generative AI ​

Author: Daisaku Sato
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.23126v1 Announce Type: cross Abstract: Generative AI design tools make natural-language prompts a starting point for design, placing new articulation demands on designers. Rather than treating prompts as the transmission of pre-existing design intent, we ask how design intent is formed th...

📖 Read original article


272. False Prophets: On the Security of World Models in Agentic Systems ​

Author: Erik Imgrund, Anna Wimbauer, Klim Kireev, Konrad Rieck
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.23147v1 Announce Type: cross Abstract: Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of these tasks requires the agent to predict the results of its actions. Recent research proposes to enh...

📖 Read original article


273. In-Context Learning as Implicit Policy Gradient ​

Author: Masahiro Kaneko, Timothy Baldwin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.23153v1 Announce Type: cross Abstract: Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their corresponding evaluation scores as in-context examples. Despite these empirical findings, the theoretical found...

📖 Read original article


274. Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining ​

Author: Rares A. C. Diaconescu, Iulia Slanina, Alina Florea, Andrei B. Trache, Miruna E. Coroi, Anne Arzberger, Jie Yang, Enrico Liscio
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23175v1 Announce Type: cross Abstract: Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present the first comparative evaluation of training-free methods for aligning language generation to user-s...

📖 Read original article


275. Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design ​

Author: Shenghao Yang, Hongtao Zhang, Yuhan Yi, Zhihao Tang, Zihao Cui, Lian Wen, Han Yan, Yuan Gao, Mingbo Zhao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.23189v1 Announce Type: cross Abstract: AI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion industry. However, 3D garment generation remains in its nascent stage, where in the realm of fashion, the semanti...

📖 Read original article


276. BoneAgeTW2: Automated Skeletal Maturation Assessment via the Tanner-Whitehouse 2 Method, Deep Learning, and Clinical Report Generation with Distribution Curves ​

Author: Juan Manuel Castillo Pinto
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.23224v1 Announce Type: cross Abstract: We present BoneAgeTW2, the first fully open-source system to automate the complete Tanner-Whitehouse 2 (TW2) clinical protocol for skeletal maturity assessment end-to-end. The system employs YOLOv8 for precise detection and localization of the 20 TW2...

📖 Read original article


277. FedSLIM: Privacy-Preserving Federated MDL-Based Descriptive Pattern Mining Across Data Silos ​

Author: Samar Samir Khalil, Noha S. Tawfik, Marco Spruit
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG

arXiv:2607.23236v1 Announce Type: cross Abstract: Federated learning has achieved considerable success for predictive modelling, yet federated descriptive analytics remains largely unexplored. Existing federated pattern mining approaches are predominantly support-based and do not optimise a principl...

📖 Read original article


278. Context-Aware Concept Distillation for Trustworthy Flood Prediction ​

Author: Eli Levinkopf, Efrat Morin, Claudia V. Goldman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.23237v1 Announce Type: cross Abstract: Effective flood risk management relies on accurate forecasting, yet the "black box" nature of stateof-the-art Deep Learning models creates a barrier to trust and accountability in high-stakes public safety decisions. While existing Explainable AI (XA...

📖 Read original article


279. X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference ​

Author: Jianwen Xian, Zhiyuan Xu, Yuchen Li, Ziliang Lai, Kang He, Zhen Huang, Aichen Feng, Jinyan Chen, Yilin Zhang, Qinqin Chen, Chengru Song
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2607.23264v1 Announce Type: cross Abstract: Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote stores and overlap data movement with Tensor Core computation. Existing systems schedule when communication is ...

📖 Read original article


280. What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features ​

Author: Chen-Yi Lu, Yueh-Shao Chen, Somali Chaterji
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.23271v1 Announce Type: cross Abstract: Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., "a dog" vs. "not a dog") to nearly identical embeddings, rendering them insensitive to negation. We attribute this failure to a phenomenon we call Representation...

📖 Read original article


281. Statistically Supported LLM Ingredient and Recipe Data Collection in Computational Nutrition ​

Author: James Izzard, Hassan Eshkiki, Fabio Caraffini
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.ET

arXiv:2607.23273v1 Announce Type: cross Abstract: Computational nutrition needs precise ingredient data, but current databases are incomplete, inconsistent, and built for human reference rather than automated reasoning. LLMs could help fill these gaps, but single-pass outputs are unreliable and can ...

📖 Read original article


282. FILLER: Feature Imputation via Latent Location Exploration and Retrieval ​

Author: Santu Mondal, Chayan Maitra, Rajat K. De
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.23295v1 Announce Type: cross Abstract: In real-world machine learning applications, incomplete observations create a fundamental challenge. Researchers have come up with several ideas to address this crucial problem. However, current models still face challenges in balancing scalability a...

📖 Read original article


283. Online Fair Division with Budget Constraints ​

Author: Saar Cohen, Nicholas Teh, Paul W. Goldberg, Michael J. Wooldridge
Published: 7/28/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.LG, cs.MA, econ.TH

arXiv:2607.23310v1 Announce Type: cross Abstract: We study an online variant of discrete fair division under generalized assignment budget constraints. Goods arrive one at a time and must be assigned irrevocably to a feasible agent or to charity, which holds all unallocated goods, while fairness is ...

📖 Read original article


284. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex ​

Author: Chanwoo Park, Asuman Ozdaglar
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.23333v1 Announce Type: cross Abstract: We revisit the regret loss framework introduced in Park et al. (2025), which uses decision-theoretic regret as a direct loss function for training models to make better decisions, through the lens of probability-simplex policies. Our first result sho...

📖 Read original article


285. Patient-Agnostic Synthetic Pretraining for Efficient Patient-Specific Intraoperative 2D/3D Registration ​

Author: Minheng Chen, Youyong Kong
Published: 7/28/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV

arXiv:2607.23343v1 Announce Type: cross Abstract: Intraoperative 2D/3D registration aligns preoperative CT volumes with intraoperative X-ray or fluoroscopic images and is essential for image-guided interventions. Recent learning-based and differentiable registration methods have shown promising accu...

📖 Read original article


286. On AI Safety and Security Technical Debt in Engineering AI-Enabled Systems ​

Author: Muhammad Tukur, Hayatullahi B. Adeyemo, Tao Chen, Nour Ali, Anis Zarrad, Rick Kazman, Marco Agus, Rami Bahsoon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CR

arXiv:2607.23365v1 Announce Type: cross Abstract: Artificial intelligence (AI) systems are increasingly deployed in high-stakes domains such as healthcare, autonomous driving, finance, and education. While these systems offer powerful data-driven and adaptive capabilities, their complexity, rapid ev...

📖 Read original article


287. Fair Division with Strictly Increasing Valuations: A Tight Threshold for Two-Agent EF1 and PO ​

Author: Nicholas Teh
Published: 7/28/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, econ.TH

arXiv:2607.23367v1 Announce Type: cross Abstract: We study whether strictly positive marginal values restore the compatibility of envy-freeness up to one good (EF1) and Pareto optimality (PO) for indivisible goods. For two agents, we identify the exact threshold in the number of goods. Every instanc...

📖 Read original article


288. Explaining BiomedCLIP with Weighted Banzhaf Interactions Supported by Tree-Gram Parsing ​

Author: Jakub Rymarski (University of Warsaw, Poland), Adam Rempa{\l}a (University of Warsaw, Poland), Bart{\l}omiej Sobieski (University of Warsaw, Poland), Przemys{\l}aw Biecek (University of Warsaw, Poland)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2607.23368v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are demonstrating significant capabilities in medical tasks like radiology analysis, yet providing faithful and interpretable explanations remains a key consideration for their responsible deployment in clinical settings...

📖 Read original article


289. When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles ​

Author: Tobias Bersia, Tatiana Gaintseva
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23379v1 Announce Type: cross Abstract: Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant informati...

📖 Read original article


290. Semantic Semi-Incremental Data-Association-Free Object SLAM ​

Author: Yihao Zhang, Jungseok Hong, John J. Leonard
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2607.23384v1 Announce Type: cross Abstract: Data association between landmark measurements and landmark variables has long been a central challenge in SLAM, as estimation accuracy depends critically on associating measurements with the correct landmark variables. Recent advances in deep learni...

📖 Read original article


291. Directional Influence Function: Estimating Training Data Influence in Constrained Learning ​

Author: Xin Wang (Jeff), R. Tyrrell Rockafellar (Jeff), Xuegang (Jeff), Ban
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.23388v2 Announce Type: cross Abstract: As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints. Understanding how training samples in- fluenc...

📖 Read original article


292. Blood Pressure Estimation from PPG: A Comparative Study of Direct and ECG-Mediated Deep Learning Pipelines ​

Author: Bo Wu, Haoling Wang, Zhuodiao Kuang, Kateryna Shapovalenko
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.ET, stat.AP

arXiv:2607.23406v1 Announce Type: cross Abstract: Continuous cuffless blood pressure (BP) monitoring is essential for connected health systems and wearable devices, enabling early detection, longitudinal tracking, and personalized management of cardiovascular disease. Many prior approaches attempt t...

📖 Read original article


293. TLA$^{+}$-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA+ Specification Generation ​

Author: Arslan Bisharat, Eric Spencer, Brian Ortiz, Khushboo Bhadauria, Mujtaba Nazari, Beatriz Santos, Anisa Ramos, TaiNing Wang, George K. Thiruvathukal, Konstantin L"aufer, Mohammed Abuhamad
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.23425v1 Announce Type: cross Abstract: Large language models increasingly write TLA$^{+}$ formal specifications from natural-language descriptions, but progress is hard to measure: existing resources grade by resemblance to a reference or by whether the output parses, neither of which sho...

📖 Read original article


294. A Characterization of the Orthocomplement of the Tangent Space of Semiparametric Markov Models ​

Author: Trung Phung, Ilya Shpitser
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ME, cs.AI

arXiv:2607.23439v1 Announce Type: cross Abstract: Graphical models are ubiquitous in social and empirical science as they are intuitive and easy to use. These models belong to the broader class of Markov models, defined using solely conditional independence (CI) restrictions. In order to estimate fi...

📖 Read original article


295. Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles? ​

Author: Hai Hu, Siyuan Song, Chongtian Shao, Kejia Zhang, Tianjian Zhu, Xiaojing Zhao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23440v1 Announce Type: cross Abstract: In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-...

📖 Read original article


Author: Moniruzzaman Mahadi, Abrar Mohammed Tanzim Alam, Sayma Siddika Monalisa, Mir Mohammad Asif Abdullah, Swakkhar Shatabda, Md Adnan Arefeen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23446v1 Announce Type: cross Abstract: A small language model can receive the governing statutory provision and still answer incorrectly. We test whether fine-tuning on examples containing relevant law improves later use of retrieved law. We curate 2{,}165 bilingual QA records from six Ba...

📖 Read original article


297. Constraint-Bound Agnostic Bayesian Optimization: One Model for All Thresholds ​

Author: Jin Wang, Xi Lin, Handing Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.LG

arXiv:2607.23448v1 Announce Type: cross Abstract: Expensive constrained optimization problems in real-world industry design often involve constraint thresholds that are difficult to determine in advance. Engineers may need to adjust constraint thresholds to explore different feasibility-performance ...

📖 Read original article


298. When Every Simulation Counts: Value-Based Reinforcement Learning for Accelerated Photonics Inverse Design ​

Author: Longying Wen, Feiyang Wu, Jinglin Yu, Chongxian Yuan, Renjie Li, Zhaoyu Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: physics.optics, cs.AI, cs.LG, physics.app-ph

arXiv:2607.23469v1 Announce Type: cross Abstract: Photonic-crystal surface-emitting lasers (PCSELs) can combine high-power operation with narrow-divergence surface emission, but optimizing coupled parameters requires costly full-wave simulations. Deep Q-network (DQN) optimization can reuse simulated...

📖 Read original article


299. ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour ​

Author: Jianhang Xie, Sicheng Tan, Vishnu Naresh Boddeti, Zhichao Lu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2607.23478v1 Announce Type: cross Abstract: Fully homomorphic encryption (FHE) provides strong cryptographic guarantees for private inference, but deploying transformer models under FHE remains prohibitively expensive. A key bottleneck is that non-linear operations such as softmax, normalizati...

📖 Read original article


300. Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation ​

Author: Musa Tur Farazi, Nufayer Jahan Reza
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.23493v1 Announce Type: cross Abstract: Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in the digital age. However, classifying internet memes remains computationally challenging due to the...

📖 Read original article


301. Formalizing Flag Algebras in Lean ​

Author: Gyeongwon Jeong, Seonghun Park, Jihoon Hyun, Sang-il Oum, Hongseok Yang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.PL, math.CO

arXiv:2607.23500v1 Announce Type: cross Abstract: Razborov's flag algebra method is a powerful tool for proving asymptotic inequalities in extremal graph theory, often reducing the task to finding a finite certificate by semidefinite programming. We present a machine-checked formalization of the met...

📖 Read original article


302. Impute On-Demand: Adaptive Correlated Time Series Imputation for Changing Environments ​

Author: Zhichen Lai, Huan Li, Dalin Zhang, Dong Gong, Lina Yao, Christian S. Jensen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DB

arXiv:2607.23503v1 Announce Type: cross Abstract: Internet of Things (IoT) applications generate vast amounts of Correlated Time Series (CTS) data that often contain missing values and require imputation. Existing methods emphasize accuracy but often lack adaptability to changing IoT environments: t...

📖 Read original article


303. Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework ​

Author: Madhav S Baidya
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.23507v1 Announce Type: cross Abstract: Choosing the right text embedding model is one of the most consequential -- and most frequently under-examined -- decisions in building a retrieval or search system, yet the model that tops a leaderboard is rarely the best choice for a given deployme...

📖 Read original article


304. Do Diagrams Help Large Language Models Reason? Evidence from Syllogistic Reasoning ​

Author: Risako Ando, Koji Mineshima
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23513v1 Announce Type: cross Abstract: Diagrams are widely used to support logical reasoning, and prior studies suggest that representations such as Euler diagrams can improve human reasoning performance. Recent work has also explored their effects on large language models (LLMs). In this...

📖 Read original article


305. Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking ​

Author: Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, Yupeng Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MM

arXiv:2607.23514v1 Announce Type: cross Abstract: Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal kn...

📖 Read original article


306. Auditing Alignment Controllability in LLMs via Political Axes ​

Author: Bartol Bu'can, Nikola So\v{c}ec, Sarah Isufi, Morena Grani'c, Luka Hobor, Agneza Krajna, Mihael Kovac, Mario Brcic
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL

arXiv:2607.23519v1 Announce Type: cross Abstract: Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answe...

📖 Read original article


307. Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric ​

Author: Nikolaos Kekatos, Stylianos Basagiannis, Panagiotis Katsaros, Alexios Lekidis, Tom Nianios
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LO, cs.RO

arXiv:2607.23532v1 Announce Type: cross Abstract: Swarms of LLM-assisted autonomous robots are increasingly proposed for cooperative intelligence, surveillance, and reconnaissance (ISR) in contested environments. A growing class of their assurance failures arises not within any single platform but a...

📖 Read original article


308. Neonatal Hypoxic-ischaemic Encephalopathy Classification from the EEG and HRV Signals Using a Conformer based Masked Autoencoder ​

Author: Shuwen Yu, William P Marnane, Geraldine B. Boylan, Gordon Lightbody
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.SP

arXiv:2607.23554v1 Announce Type: cross Abstract: In this paper, we propose the MAEConformer, a novel self-supervised learning framework that combines the Conformer architecture with the Masked Autoencoder (MAE) paradigm for large-scale representation learning from unlabelled electroencephalography ...

📖 Read original article


309. GTIN: A Unified Framework for Joint Event and Time Prediction in Temporal Graphs ​

Author: Mohammad Ostadmohammadi, Sepehr Kazemi, Hamid R. Rabiee
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SI, cs.AI

arXiv:2607.23556v1 Announce Type: cross Abstract: Temporal graphs are increasingly used to model dynamic systems in diverse domains such as social networks, financial networks, and traffic networks. Predicting both what the next event will be and when it will occur in these systems is crucial for un...

📖 Read original article


310. An Unofficial FastLAS Tutorial: A Programmer's Guide ​

Author: Fabio Aurelio D'Asaro
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.LG

arXiv:2607.23557v1 Announce Type: cross Abstract: FastLAS is a scalable system for Inductive Logic Programming (ILP): you give it some background knowledge, a language bias, and a set of examples, and it searches for a set of logic program rules (a hypothesis) that explains the examples. These notes...

📖 Read original article


311. D3O: Dynamic Distribution Distillation for Ordinal Regression ​

Author: Chunlai Dong, Yaojun Hu, Yuyang Xu, Haochao Ying, Jian Wu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.23575v1 Announce Type: cross Abstract: Ordinal regression is widely used in scenarios where labels are discrete yet inherently ordered. In practice, however, ordinal labels are often obtained by discretizing underlying continuous semantics through subjective human judgment, resulting in a...

📖 Read original article


312. Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models ​

Author: Liangyu Li, Qingwen Liu, Mingqing Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2607.23602v1 Announce Type: cross Abstract: Controllers based on sampling and latent world models assign a predicted terminal cost to each candidate action sequence, choose the minimum, execute its first action block, and replan. This rule can fail even when the terminal cost perfectly and acc...

📖 Read original article


313. DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory ​

Author: Xingyang Yu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, hep-th

arXiv:2607.23614v1 Announce Type: cross Abstract: We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a boun...

📖 Read original article


314. Where Is the Cost of Third-Party API Routers in Agentic Software Development? ​

Author: Donghao Fu, Jingxin Li, Xue Jiang, Yihong Dong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2607.23624v1 Announce Type: cross Abstract: Third-party API routers have become a common layer that unifies access across increasingly diverse LLM providers. In coding-agent workflows, high-autonomy operation is widely adopted because it reduces interaction overhead. As a result, a third-party...

📖 Read original article


315. Order in Desbordante: Techniques for Efficient Implementation of Order Dependency Discovery Algorithms ​

Author: Yakov Kuzin, Dmitriy Shcheka, Michael Polyntsov, Kirill Stupakov, Mikhail Firsov, George Chernishev
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.LG, cs.PF

arXiv:2607.23632v1 Announce Type: cross Abstract: Science-intensive data profiling focuses on discovery and validation of various patterns in datasets. This study considers discovery of one such pattern - order dependency (OD). Simply put, OD states that some list of columns is ordered according to ...

📖 Read original article


316. Variational-Ising-Attention (VIA):TailoredAttentionMattersfor Science ​

Author: Rui Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.chem-ph

arXiv:2607.23634v1 Announce Type: cross Abstract: Attention enables context modeling via query-key scoring with softmax normalization. Driven by industrial long-context demands, mainstream research has converged toward sparsity and efficiency--yet softmax's independence assumption persists. For scie...

📖 Read original article


317. Extending Desbordante with Probabilistic Functional Dependency Discovery Support ​

Author: Ilia Barutkin, Maxim Fofanov, Sergey Belokonny, Vladislav Makeev, George Chernishev
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CE, cs.LG

arXiv:2607.23636v1 Announce Type: cross Abstract: Data profiling aims to extract complex patterns from data for further analysis and use that data in domains such as data cleaning, data deduplication, anomaly detection, and many more. Functional dependencies (FDs) are one of the most well-known patt...

📖 Read original article


318. CALMRec: Causally Aligned Language Memory for Long-Horizon Recommendation ​

Author: Gengyu Zhan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.23647v1 Announce Type: cross Abstract: Large language models (LLMs) can summarize heterogeneous user evidence in natural language, but current LLM recommenders often collapse enduring preferences, transient intent, and exposure-induced behavior into one profile. This makes recommendation ...

📖 Read original article


319. Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents ​

Author: Aayush Kumar, Avik Dutta, Sumit Gulwani, Gustavo Soares, Advait Sarkar, Emerson Murphy-Hill
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.SE

arXiv:2607.23670v1 Announce Type: cross Abstract: Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to develop a plan before task execution. However, it remains unclear whether the benefits of this feature...

📖 Read original article


320. An empirical investigation into the properties of standard word embeddings ​

Author: Salomon Kabongo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23675v1 Announce Type: cross Abstract: The embedding of word sequences into continuous vector spaces has been one of the most important developments in Natural Language Processing in the recent past. Such embeddings have found application in areas such as Automatic Speech Recognition, Mac...

📖 Read original article


321. The Illusion of Secure LLM Code: Closing the Security Gap via Iterative Reprompting ​

Author: Ishpuneet Singh, Shreyas Mahajan, Gurjot Singh, Maninder Singh
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.HC, cs.LG, cs.MA

arXiv:2607.23710v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly integrated into software development workflows, yet their ability to autonomously generate secure authentication code remains uncertain. This paper evaluates the security architecture of authentication sy...

📖 Read original article


322. An Exact Counterexample to Carlson's Associated-Prime Depth Conjecture from a Group of Order 128 ​

Author: Xinan Dai, Wenhao Deng, Yingdong Shi, Tailin Wu, Yuchen Yang
Published: 7/28/2026, 4:00:00 AM
Categories: math.GR, cs.AI

arXiv:2607.23732v1 Announce Type: cross Abstract: In Question~3.1 of his 1995 paper on depth and transfer, Carlson asked whether the depth of a finite-group cohomology ring is always realized by the dimension of one of its associated primes. We give a negative answer. Let [ G=\SG{128}{859},\qquad k...

📖 Read original article


323. AI Strategy: How to Choose What AI Product to Implement ​

Author: Foster Provost, Panos Ipeirotis
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG, econ.GN, q-fin.EC, stat.AP

arXiv:2607.23733v1 Announce Type: cross Abstract: Firms struggle to choose AI projects that pay off: two projects can look equally promising to smart, motivated stakeholders and yet deserve opposite decisions. At the residential real-estate brokerage Compass, one AI product (Likely-to-Sell recommend...

📖 Read original article


324. Escaping the Euclidean Void: Manifold-Informed Flow Matching for Sequential Recommendation ​

Author: Dengzhao Fang, Jingtong Gao, Yu Li, Xiangyu Zhao, Yi Chang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.23762v1 Announce Type: cross Abstract: Conventional recommenders capture users' preferences by optimizing observed user-item relations, whereas continuous generative recommendation additionally learns the trajectory of synthesizing a target item. Flow matching drives this process by gradu...

📖 Read original article


325. WISERouter: LLM Routing with Workload Budget Constraint ​

Author: Yifei Li, Zihui Gao, Laks V. S. Lakshmanan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.23765v1 Announce Type: cross Abstract: Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is prohibitive at scale. LLM routing exploits diversity in model capability and cost by assigning each query to a su...

📖 Read original article


326. Outcome-Fair Restless Multi-Armed Bandits for Stochastic Deadline Scheduling ​

Author: Shakti Sharma, Rahul Meshram
Published: 7/28/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY

arXiv:2607.23772v1 Announce Type: cross Abstract: We study a restless multi-armed bandit (RMAB) problem for a stochastic deadline scheduling application. RMAB problems are solved using the Whittle index policy. The goal in RMAB is to maximize the expected cumulative discounted reward maximization. T...

📖 Read original article


327. Scale Weight Decay and Train Better ​

Author: Anuj Apte
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.dis-nn, cs.AI, math.OC

arXiv:2607.23777v1 Announce Type: cross Abstract: The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data. This is typically done with a constant decoupled weight decay which causes the network weights to shrink steadily over the course of training....

📖 Read original article


328. A Few Words Go a Long Way: Language Guided Robot Policy Synthesis ​

Author: Daphne Chen, Archit Ritesh Jain, Eric Goossen, Emma Romig, Michael Murray, Nick Walker, Maya Cakmak
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.23784v1 Announce Type: cross Abstract: While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies that are difficult to interpret, adapt, or correct when they inevitably fail. In this work, we propose ...

📖 Read original article


329. Maximum Satisfiability of Simple Temporal Problems ​

Author: Johannes K. Fichte, Johanna Groven, Peter Jonsson, Victor Lagerkvist, Jorke M. de Vlas
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CC, cs.AI

arXiv:2607.23785v1 Announce Type: cross Abstract: The Simple Temporal Problem (STP) is a core framework for quantitative temporal constraints. As STP data can be inconsistent, we study MAXSTP: compute a maximum-cardinality consistent subset of constraints. This extension is NP-hard, and we analyze i...

📖 Read original article


330. PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis ​

Author: Chi Phan, Tianyi Zhang, Yufeng Wu, Qiaochu Xue, Jiajie Zhang, Linghan Cai, Zeyu Liu, Sudong Wang, Yueming Jin, Dan Hu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.23794v1 Announce Type: cross Abstract: Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morphology at higher magnification. However, existing pathology benchmarks and vision-language models (VLMs) ...

📖 Read original article


331. How Context Attribution Handles What the Model Already Knows ​

Author: Quoc-Huy Trinh, Lin Zhu, Sebastian Szyller
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23804v1 Announce Type: cross Abstract: Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works show the initial success in attributing the con- tributive score of the contexts. However, we observe that when ...

📖 Read original article


332. A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever ​

Author: Sietse Schelpe (Corbenic AI)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR, cs.LG, cs.PF

arXiv:2607.23806v1 Announce Type: cross Abstract: Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a ...

📖 Read original article


333. Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages ​

Author: Deovrat Mehendale, Aditya Mehndiratta, Dhruv Rathi, Kaushal Bhogale, Mitesh M. Khapra
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23808v1 Announce Type: cross Abstract: In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field...

📖 Read original article


334. Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance ​

Author: Denglin Jiang, Haoran Zhou, Anshul Wadhawan, Brendan Fahy, Vinay Ramesh, David Weisberg, Dmitriy Derkachevskiy, Helen Sheehan, Srivas Prasad, Michele Franceschini
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23813v1 Announce Type: cross Abstract: We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i) testset-full, 498 hours of full...

📖 Read original article


335. Kalypso: Relational LLM Serving ​

Author: Hojae Son, Md Ashraful Islam, Huy Gia Cao, Hui Guan, Marco Serafini
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CL

arXiv:2607.23815v1 Announce Type: cross Abstract: Large language models are increasingly used as semantic operators for filtering, extracting, ranking, joining, and transforming unstructured data. Existing semantic query processing systems invoke request-centric LLM serving systems that are unaware ...

📖 Read original article


336. TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation ​

Author: Susil Kumar Mohanty, Rohit Patel, Kosuru Yuvaraj, Jeenal Chaudhary, Disha Singhania
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG

arXiv:2607.23838v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base at query time. This makes RAG useful for private data, fast-changing information, and reducing hallucination, ...

📖 Read original article


337. Limbomorphs ​

Author: Alex Alvarez, Michael Levin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2607.23842v1 Announce Type: cross Abstract: Artificial life systems are typically defined by a set of dynamical rules over an environment, an agent, or both, from which lifelike patterns may emerge. Gifbreeder is an animated version of the interactive evolutionary computation (IEC) platform Pi...

📖 Read original article


338. A Coulomb Particle Model for Learning Kernel Attention in Transformers ​

Author: Masoud Badiei Khuzani, Sharath Honnaiah, Atiq Islam, Alex Cozzi, Abraham Bagherjeiran
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.23869v1 Announce Type: cross Abstract: Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution. We propose a particle-based method that learns this distribution by optimizing kernel-target alignm...

📖 Read original article


339. MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents ​

Author: Belal S. Alsinglawi, Weizheng Wang, Junyi Wu, Yi Jiang, Lianhai Lin, Merouane Debbah, Izzat Alsmadi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2607.23870v1 Announce Type: cross Abstract: Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing UAV and multimodal ...

📖 Read original article


340. Physics-Informed Neural Networks for Predicting Nitrous Oxide Flux ​

Author: Freddy Yu, Jashanjeet Kaur Dhaliwal, Subhadeep Chakraborty
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2607.23880v1 Announce Type: cross Abstract: Nitrous oxide (N$_2$O) is the dominant ozone-depleting substance emitted in the 21st century, and the third largest contributor to anthropogenic greenhouse gases due to its high potency and long atmospheric lifetime, with more than 70% of N$_2$O emis...

📖 Read original article


341. Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature ​

Author: Tanjin He, Aikaterini Vriza, Logan Ward, Xu Huang, Yiming Chen, Anubhav Jain, Gerbrand Ceder, Rajeev S. Assary, Ian T. Foster, Maria K. Y. Chan
Published: 7/28/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, cs.CL, cs.DL, cs.IR

arXiv:2607.23886v1 Announce Type: cross Abstract: X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described through ...

📖 Read original article


342. Visible to the Court: How AI Is (and Isn't) Litigated in U.S. Federal Court Opinions ​

Author: Julie Yu, Rock Yuren Pang, Jevan Hutson, Katharina Reinecke
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY

arXiv:2607.23888v1 Announce Type: cross Abstract: In the United States, artificial intelligence (AI) is rapidly deployed amid limited federal regulation. With courts become a recurring forum in which AI-related practices are scrutinized, it is important to empirically understand the AI litigation la...

📖 Read original article


343. Embodied GPT-5.1: Evidence of a World Model? ​

Author: Roberto Spinelli, Thiago C. Martins
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.23899v1 Announce Type: cross Abstract: This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level controller of a physical mobile robot despite having no prior embodiment, no training in simulated environments, and no exposure to sensor...

📖 Read original article


344. Understanding Tone-Dependent Inference Cost in Large Language Models ​

Author: Akhil Kumar, Om Dobariya
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23915v1 Announce Type: cross Abstract: We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed to understand the trade-offs between accuracy and inference cost on a 570 Question MMLU datase...

📖 Read original article


345. DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection ​

Author: Jyun-Ze Tang, Po-Han Huang, Ming-Ching Chang, Chih-Fan Hsu, Jeng-Lin Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.23924v1 Announce Type: cross Abstract: Vision foundation models have enabled strong training-free anomaly detection (AD). However, most existing approaches rely primarily on independent local patch features, leaving the global contextual information encoded by Vision Transformers (ViTs) u...

📖 Read original article


346. SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving ​

Author: Yihui Zhang (Beihang University), Tianyu Wo (Beihang University), Jinghao Wang (Beihang University), Xiaoyang Sun (University of Leeds), Menghao Zhang (Beihang University), Cangzhou Yuan (Beihang University), Li Li (Beihang University), Chunming Hu (Beihang University), Albert Y. Zomaya (The University of Sydney), Renyu Yang (Beihang University)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG, cs.PF

arXiv:2607.23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency. Persistent long-li...

📖 Read original article


347. Understanding Machine Unlearning Through the Lens of Mode Connectivity ​

Author: Jiali Cheng, Hadi Amiri
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.23970v1 Announce Type: cross Abstract: Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress, the loss landscape and optimization geometry of unlearning are poorly understood. In this paper, we study machi...

📖 Read original article


348. Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models ​

Author: Tapan Parikh
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23976v1 Announce Type: cross Abstract: Appending a two-word confirmation tag to a decision question -- "Is X the better choice?" versus "X is the better choice, right?" -- changes whether a language model endorses the choice. We measure this tag effect on 20 frozen, ground-truth-free deci...

📖 Read original article


349. Multimodal Semantic-Probabilistic Objectness for Open World Object Detection ​

Author: Weijun Tian, Rui Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.23981v1 Announce Type: cross Abstract: Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes. PROB improves unknown discovery by modeling class-agnostic probabi...

📖 Read original article


350. Moral Hazard in Multi-Agent Language Models ​

Author: Dane Malenfant
Published: 7/28/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2607.23982v2 Announce Type: cross Abstract: Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmstr"om's team moral-hazard model, we introduce the Dialogue Moral Hazard Game, a controlled textual game that operationalizes...

📖 Read original article


351. Adaptive Data Admission and Retention for Streaming Federated Learning ​

Author: Zhuoyi Zhao, Ben Liang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NI

arXiv:2607.23987v1 Announce Type: cross Abstract: We study streaming federated learning with limited client memory, where newly generated training data incur time-varying sampling costs and must be selectively admitted and retained over time. We consider a joint server-side admission and client-side...

📖 Read original article


352. SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding ​

Author: Seoyeon Kim, Minjae Kang, Jaehyung Kim
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23991v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirements. However, models follow these prompts only implicitly through in-context learning, which can be insufficient ...

📖 Read original article


353. Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation ​

Author: Mohan Manivannan, Dalal Alharthi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.DC

arXiv:2607.24006v1 Announce Type: cross Abstract: Cloud telemetry arrives at a scale that, paradoxically, makes intrusion understanding harder rather than easier. Attackers operate through legitimate identity, federated session tokens, and cloud native APIs indistinguishable from routine administrat...

📖 Read original article


354. Disentangling Semantic Attention from Structural Bias in the Attention Manifold ​

Author: Pengkun Jiao, Bin Zhu, Jingjing Chen, Yu-gang Jiang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.24017v1 Announce Type: cross Abstract: The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disproportionate attention toward certain semantically uninformative visual t...

📖 Read original article


355. HELIOS: An LLM-Driven Autonomous Indirect Trajectory Optimization Agent ​

Author: An-yi Huang
Published: 7/28/2026, 4:00:00 AM
Categories: astro-ph.IM, cs.AI, cs.RO

arXiv:2607.24051v1 Announce Type: cross Abstract: Low-thrust trajectory optimization is a core technology in deep-space mission design. Indirect methods based on Pontryagin's Minimum Principle (PMP) offer rigorous optimality guarantees, yet their practical application faces three bottlenecks: (1) tr...

📖 Read original article


Author: L'eo Hein, Giovanni De Nunzio, Aur'elie Pirayre, Laurent Najman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.24056v1 Announce Type: cross Abstract: Network-wide traffic volume estimation typically relies on propagating measurements from fixed sensors, making performance highly dependent on sensor density and limiting deployment in sparsely instrumented networks. We propose a link-level learning ...

📖 Read original article


357. ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning ​

Author: Wenwu Fan, Qihong Lin, Zhijie Xia, Zhuo Zheng, Sihao Wang, Qiang Chen, Liangsheng Zhu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.24062v1 Announce Type: cross Abstract: Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between training and inference. This training-inference discrepancy stems from two primary factors: an architectural separatio...

📖 Read original article


358. MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning ​

Author: Tuan-An To, Yuk-Kwan Wong, Tuan-Anh Vu, Ziqiang Zheng, Sai-Kit Yeung
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.24064v1 Announce Type: cross Abstract: Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality image-text pairs. However, the performance of VLMs often degrades in the video domain due to the essenti...

📖 Read original article


359. Towards simultaneous decoding of kinetic and kinematic movement parameters during grasp and lift task by noninvasive brain imaging ​

Author: Parth G. Dangi, Yogesh Kumaar Meena
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.ET

arXiv:2607.24081v1 Announce Type: cross Abstract: Brain-machine interfaces (BMIs) can assist individuals with limited mobility, such as stroke survivors or amputees. One of the key challenges in developing BMIs is expanding their usability and control, which can be achieved by accurately decoding mu...

📖 Read original article


360. LU-500: A Logo Benchmark for Concept Unlearning ​

Author: Keyu Li, Jin Gao, Jialing Zhang, Dequan Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.24101v1 Announce Type: cross Abstract: Concept unlearning is increasingly used to limit the reproduction of protected or unsafe visual concepts in text-to-image models. Existing evaluations, however, mostly study targets that dominate the whole image, such as styles, broad object categori...

📖 Read original article


361. A Case Study on the Acceptance of a Humanoid Robotic Head Employed in Three Public Spaces ​

Author: Marcel Heisler, Luca Randecker, Christian Becker-Asano
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.HC

arXiv:2607.24113v1 Announce Type: cross Abstract: Previous research has shown that a human-like robot's acceptance heavily depends on the setting in which it operates and its ability to perform relevant tasks. This paper, first, reports on how our robot processes natural language to generate a multi...

📖 Read original article


362. EEGForceFusion: Joint Tokenised-Continuous Representation Learning for Subject-Independent Grasp Force Decoding ​

Author: Sankalp Sunil Turankar, Yogesh Kumar Meena
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.ET

arXiv:2607.24126v1 Announce Type: cross Abstract: Brain-machine interfaces provide a link between neural activity and external devices, enabling restoration of motor function and advancing human-machine interaction using non-invasive electroencephalography (EEG). However, continuous grasp force deco...

📖 Read original article


363. Monitoring Post-Disaster Urban Recovery Using High-Resolution SAR Time Series and Unsupervised Learning: Evidence from the 2023 T\"urkiye-Syria Earthquake ​

Author: Luigi Russo, Deodato Tapete, Silvia Liberata Ullo, Paolo Gamba
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.24180v1 Announce Type: cross Abstract: Monitoring post-disaster recovery is essential for understanding how urban systems rebuild and progressively return to functionality. However, tracking reconstruction remains difficult because reliable ground-truth information is often scarce and rec...

📖 Read original article


364. Not Forgotten: Implementation and Evaluation of a Personalized Episodic Memory for the Humanoid Robot Head Kim ​

Author: Steve Aschenbrenner, Marcel Heisler, Thomas Sievers, Christian Becker-Asano
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.HC

arXiv:2607.24190v1 Announce Type: cross Abstract: Social robots that rely on large language models for conversation are unable to retain information across sessions. This absence of memory violates social expectations, potentially preventing the formation of persistent relationships. This paper pres...

📖 Read original article


365. StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting ​

Author: Heyan Chai, Xin Li, Wenjie Wang, Jianyang Qin, Chaoyang Li, Lu Wang, Hao Chen, Qing Liao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.24191v1 Announce Type: cross Abstract: Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evolution of beliefs, particularly during stance reversal...

📖 Read original article


366. Every Client Is an Environment: Federated De-confounding for Spatio-Temporal Forecasting ​

Author: Qingxiang Liu, Anqi Liang, Heng Wang, Yuxuan Liang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.24218v1 Announce Type: cross Abstract: Federated learning has emerged as a promising paradigm for spatio-temporal forecasting (STF), enabling collaborative model training without sharing raw observations. Existing federated STF methods primarily regard cross-client heterogeneity as an opt...

📖 Read original article


367. FilmBench: A Film-Grade Benchmark for Cinematic Video Generation ​

Author: Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian, Hua Li, Jingjing Fan, Mingshuang Tang, Zhong Li, Hengxia Qiang, Weibin Chen, Jinyang Zhen, Bing Zhao, Lin Qu, Jing Li, Hu Wei
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.24241v1 Announce Type: cross Abstract: Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. M...

📖 Read original article


368. ML-based Predictive Models for Power Consumption in Virtualised O-RANs ​

Author: Rishu Raj, Genevieve Akude, Urooj Tariq, Daniel Kilper
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.24256v1 Announce Type: cross Abstract: As communication networks adopt virtualized and disaggregated architectures, achieving energy efficiency has become increasingly important for both economic and environmental reasons. Traditional methods for power modeling are inadequate in these dyn...

📖 Read original article


369. Physics-Guided Generative AI for Property-Targeted 3D Porous Media Design ​

Author: Peng Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.24274v1 Announce Type: cross Abstract: Inverse design of three-dimensional porous media is central to applications in filtration, catalysis, energy storage, fuel cells, thermal management, and biomedical scaffolds, but remains challenging because many distinct pore geometries can share si...

📖 Read original article


370. A Computational Ethical Framework for Financial Digital Phenotyping for Mental Health ​

Author: Oluwadara Adedeji, Michael Mayowa Farayola, Jeff Brozena, Irina Tal, Regina Connolly, Mark Matthews
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.CY

arXiv:2607.24275v1 Announce Type: cross Abstract: Ethical governance of AI-driven systems is often expressed through high-level principles and static documentation, creating a gap between regulatory requirements and system-level verification. This challenge is particularly acute in digital phenotypi...

📖 Read original article


371. The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages ​

Author: Priyansh Srivastava
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.24276v1 Announce Type: cross Abstract: Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are trained predominantly on English-centric corpora, they introduce a systematic and often overlooked dis...

📖 Read original article


372. Teacher Knows It Best: Spontaneous Symmetry Breaking and Tipping Points in Networked Langevin Dynamics AI Sycophancy ​

Author: Sayantari Ghosh, Saumik Bhattacharya, Partha Pratim Chakrabarti
Published: 7/28/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI, math.DS

arXiv:2607.24304v1 Announce Type: cross Abstract: We formulate a statistical physics framework to model a networked stochastic dynamical system exhibiting bistability, driven by additive noise and social conformity. We apply this model to understand and mitigate AI-induced delusional spiraling-a phe...

📖 Read original article


373. Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls ​

Author: Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Khandaker Rifah Tasnia, Sifat Rahman Ahona, Juena Ahmed Noshin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.24343v1 Announce Type: cross Abstract: Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body but should not determine a recipient, account, command, or credential. Existing statistical methods ty...

📖 Read original article


374. DeepFaith: Evidence-Grounded LLMs for Faithful Incident Reporting in Multi-Stage APT Defense ​

Author: Trung V. Phan, Tri Gia Nguyen, Thomas Bauschert
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.24348v1 Announce Type: cross Abstract: Advanced Persistent Threats (APTs) are difficult to detect and interpret due to their multi-stage and stealthy nature. While recent autonomous defense systems leverage provenance graphs and learning-based models for detection and mitigation, their ou...

📖 Read original article


375. Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs ​

Author: Jianru Shen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.24371v1 Announce Type: cross Abstract: Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic coding, CPT for procedure billing, and HL7 FHIR for data exchange. While large language models demonstrat...

📖 Read original article


376. MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention ​

Author: Jianlin Yu, Jing Lin, Linghui Kong, Aiyue Chen, Weiyi Sun, Chenyu Zeng, Wangli Lan, Jinxi Li, Zhuo Zheng, Ziyang Yue, Danning Ke, Fei Yi, Tianchi Hu, Yuan Ding, Yiwu Yao, Junsong Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2607.24377v1 Announce Type: cross Abstract: The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numeric...

📖 Read original article


377. Regulating for AI Legitimacy ​

Author: Gilad Abiri
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.24391v1 Announce Type: cross Abstract: AI systems already govern. They rank speech and allocate attention, filter applicants and triage claims. The dominant frame for AI governance, alignment, asks whether such systems pursue the right objectives safely. It cannot answer a prior question:...

📖 Read original article


378. The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing ​

Author: Stefan Scholze, Johannes Partzsch, Sebastian H"oppner, Florian Kelber, Andreas Dixius, Marco Stolba, Sirine Arfa, Marc Berthel, Georg Ellguth, Jim Garside, Hector A. Gonzalez, Stephan Hartmann, Thomas Kiel-Hocker, Dongwei Hu, Matthias Jobst, Khaleelulla Khan Nazeer, Tim Langer, Chen Liu, Gengting Liu, Matthias Lohrmann, Mantas Mikaitis, Felix Neum"arker, Amirhossein Rostami, Stefan Schiefer, Tilo Schubert, Delong Shang, Bernhard Vogginger, Yexin Yan, Steve Furber, Christian Mayr
Published: 7/28/2026, 4:00:00 AM
Categories: cs.ET, cs.AI, cs.AR, cs.DC

arXiv:2607.24396v1 Announce Type: cross Abstract: In deep learning, efficiency gets more and more important to compensate for the ongoing growth in model sizes and applications. Neuromorphic hardware has long been advocated as an upcoming alternative to deep networks, taking inspiration from the bra...

📖 Read original article


379. Multivariate Time Series Forecasting with Adaptive Non-Local Observables ​

Author: Yu-Ting Lee, Huan-Hsin Tseng, Samuel Yen-Chi Chen
Published: 7/28/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG

arXiv:2607.24399v1 Announce Type: cross Abstract: Multivariate time series forecasting (MTSF) predicts future values of multiple variables from historical data. While quantum neural networks have been increasingly applied to this task, they typically rely on fixed local measurements, which restrict ...

📖 Read original article


380. DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference ​

Author: Dengke Han
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.24434v1 Announce Type: cross Abstract: Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-use...

📖 Read original article


381. LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings ​

Author: Brittany Harbison, Ashok K. Goel
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.24435v1 Announce Type: cross Abstract: Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining prevalence and agreement diagnostics wi...

📖 Read original article


382. Evaluating RAG for French immigration law: a benchmark and baseline study ​

Author: Annia Abtout, Julien Delaunay, Monika Ewa Rakoczy
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2607.24449v1 Announce Type: cross Abstract: International recruitment in France requires navigating a layered legal framework absent from existing legal AI benchmarks. We present a publicly available benchmark and first comparative evaluation for this domain, covering permit-type recommendatio...

📖 Read original article


383. ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image ​

Author: Mingzhi Xu, Yizhe Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.24453v1 Announce Type: cross Abstract: Learning from minimal human supervision is a long-standing goal in medical image analysis, where dense expert annotations are costly. We study retinal vessel segmentation in an extreme semi-supervised setting with one annotated image and a pool of un...

📖 Read original article


384. UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective ​

Author: Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling. However, adapting AR checkpoints across c...

📖 Read original article


385. DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes ​

Author: Jiahao Xie, Zhongbin Guo, Qianle Wang, Ruiqi Lu, Dongling Xiao, Wanxuan Sun, Cheng Yang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.24516v1 Announce Type: cross Abstract: While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios by intuitio...

📖 Read original article


386. Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls ​

Author: Marzieh Zare
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE

arXiv:2607.24519v1 Announce Type: cross Abstract: Pretrained EEG foundation models are increasingly proposed for clinical decoding, but their transfer across populations and robustness to negative controls remain unclear. We benchmark six models (LaBraM, EEGMamba, CBraMod, REVE, BENDR, and BIOT) on ...

📖 Read original article


387. EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings ​

Author: Xiaocheng Fang, Jieyi Cai, Guangkun Nie, Haoyu Wang, Jiarui Jin, Yujie Xiao, Bo Liu, Chenyang He, Qinghao Zhao, Gaofeng Cheng, Hongyan Li, Shenda Hong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.24553v1 Announce Type: cross Abstract: Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardiac findings. Global ECG--text alignment may entangle modality-specific factors, while long-tailed finding distr...

📖 Read original article


388. LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding ​

Author: Junsung Hwang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.24555v1 Announce Type: cross Abstract: Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read in full at every decode step. Attention keys are locally low-rank though globally high-rank: shared low-rank bases discard page-specific directio...

📖 Read original article


389. BettiSplit: Topology-Guided Privacy-Aware Split Learning Against Feature Inversion and Gradient Leakage ​

Author: Akarsh K. Nair, Muhammad Arifur Rahman, David Brown, Mufti Mahmud
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2607.24556v1 Announce Type: cross Abstract: Split learning enables collaborative model training by partitioning neural networks across clients and servers. However, improper split placement can lead to severe privacy leakage through intermediate representations. In this work, we propose a topo...

📖 Read original article


390. EgoPlay: Event-Triggered Video Editing for Egocentric Streams ​

Author: Jinjie Mai, Gordon Guocheng Qian, Willi Menapace, Arpit Sahni, Chaoyang Wang, Ashkan Mirzaei, Runjia Li, Sergey Tulyakov, Bernard Ghanem, Peter Wonka, Rameen Abdal
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.24560v1 Announce Type: cross Abstract: We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered...

📖 Read original article


391. The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding ​

Author: Jiameng Zhang, Srikanth Madikeri
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.24570v1 Announce Type: cross Abstract: Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violations occur within each video. Processing every frame is infeasible at scale, so systems are constrained to spar...

📖 Read original article


392. CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding ​

Author: Jinlong Yang, Wenhao Zhang, Kuanwei Lin, Sijie Cheng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.24582v1 Announce Type: cross Abstract: Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference procedure to every example regardless of difficulty. This uniform strategy invokes unnecessary tool-a...

📖 Read original article


393. D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models ​

Author: Bianca Raimondi, Davide Evangelista, Maurizio Gabbrielli, Elena Loli Piccolomini
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.24586v1 Announce Type: cross Abstract: Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to be internally represented by the model. We study hallucination detection from the geometry of hidden ...

📖 Read original article


394. Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review ​

Author: Zhenhan Gao, Marvin Mu~noz Bar'on, Umm-e Habiba, Daniel Graziotin, Stefan Wagner
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.HC

arXiv:2607.24601v1 Announce Type: cross Abstract: Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains hard to understand. Developers struggle to assess the validity of LLM-generated reviews, making it difficult to g...

📖 Read original article


395. Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair ​

Author: Xueping Gao, Jianwei Yang, Qiang Yang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.24604v1 Announce Type: cross Abstract: Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval...

📖 Read original article


396. Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents ​

Author: Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov, Ildar Iskhakov, Matvey Kukuy
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.24625v1 Announce Type: cross Abstract: Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional taint trackin...

📖 Read original article


397. Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects ​

Author: Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty, Iryna Gurevych, Subhabrata Dutta
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.24645v1 Announce Type: cross Abstract: The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can ...

📖 Read original article


398. A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility ​

Author: Rajat Sainju, Dariusz Jarosz, Hairong Shang, Michael Prince, Ryan M. Aydelott, Mathew J. Cherukara, Yine Sun, Michael D. Borland
Published: 7/28/2026, 4:00:00 AM
Categories: physics.acc-ph, cs.AI, cs.IR

arXiv:2607.24663v1 Announce Type: cross Abstract: Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system data. We prese...

📖 Read original article


399. Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification ​

Author: Francisco Mena, Dino Ienco, Roberto Interdonato, Cassio F. Dantas, Simon Besnard
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.24683v1 Announce Type: cross Abstract: Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real-world scenarios subject to operational constraints, such as sensor failures or privacy restrictions, lead to i...

📖 Read original article


400. Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines ​

Author: Jhonatan Tavori, Gur-Eyal Sela, Ion Stoica, Gil Zussman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.CR, cs.DC

arXiv:2607.24692v1 Announce Type: cross Abstract: Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline together with a higher-accuracy slow path that runs higher-compute methods on stronger, remote hardware, so its results can be retur...

📖 Read original article


401. Efficient LLM-Generated Shuttling Compilers for Complex Trapped-Ion Architectures ​

Author: Fabian Kreppel, Reza Salkhordeh, Ferdinand Schmidt-Kaler, Andr'e Brinkmann
Published: 7/28/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.ET

arXiv:2607.24714v1 Announce Type: cross Abstract: Trapped-ion quantum computers rely on shuttling compilers, which cast an input algorithm into a sequence of ion-qubit movements within a given architecture. We present the first study in which a single frontier large language model (LLM), Claude Opus...

📖 Read original article


402. DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data ​

Author: Zhen Huang, Yikun Wang, Shijie Xia, Pengfei Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.24717v1 Announce Type: cross Abstract: Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixed processing strategy at the corpus or domain level and apply it uniformly to many examples, without...

📖 Read original article


403. The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation ​

Author: Tianyi Men, Zhuoran Jin, Kang Liu, Jun Zhao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.24720v1 Announce Type: cross Abstract: Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning abilit...

📖 Read original article


404. KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability ​

Author: Krithi Shailya, Ananya Lakshmi Ravi, Venkatanathan K. V., Sowmya S. Sundaram, Gokul S. Krishnan, Aditi Anand, Balaraman Ravindran
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.24730v1 Announce Type: cross Abstract: Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician trust. In clinical workflows, chest X-ray classifiers are increasingly paired with Vision-Language Models (VLMs)...

📖 Read original article


405. Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation ​

Author: Bingnan Li, Haozhe Wang, Haozhong Xiong, Fangtai Wu, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.24731v1 Announce Type: cross Abstract: On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains po...

📖 Read original article


406. ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding ​

Author: Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2607.24743v2 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medica...

📖 Read original article


407. Procedural Content Generation via Generative Artificial Intelligence ​

Author: Xinyu Mao, Wanli Yu, Kazunori D Yamada, Michael R. Zielewski
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2407.09013v2 Announce Type: replace Abstract: The attempt to utilize machine learning in PCG has been made in the past. In this survey paper, we investigate how generative artificial intelligence (AI), which saw a significant increase in interest in the mid-2010s, is being used for PCG. We rev...

📖 Read original article


408. ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins ​

Author: Safayat Bin Hakim, Muhammad Adil, Alvaro Velasquez, Houbing Herbert Song
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.LG, cs.SC

arXiv:2501.08561v5 Announce Type: replace Abstract: Digital twins are increasingly used to monitor and optimize industrial systems, yet many existing frameworks remain difficult to interpret, slow to adapt, and limited in their ability to incorporate explicit domain knowledge. This paper presents AN...

📖 Read original article


409. Robustness and Cybersecurity in the EU Artificial Intelligence Act ​

Author: Henrik Nolte, Miriam Rateike, Mich`ele Finck
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.CY, cs.LG

arXiv:2502.16184v3 Announce Type: replace Abstract: The EU Artificial Intelligence Act (AIA) establishes different legal principles for different types of AI systems. While prior work has sought to clarify some of these principles, little attention has been paid to robustness and cybersecurity. This...

📖 Read original article


410. Hybrid AI-Physical Modeling for Penetration Bias Correction in X-band InSAR DEMs: A Greenland Case Study ​

Author: Islam Mansour, Georg Fischer, Ronny Haensch, Irena Hajnsek
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2504.08909v2 Announce Type: replace Abstract: Digital elevation models derived from Interferometric Synthetic Aperture Radar (InSAR) data over glacial and snow-covered regions often exhibit systematic elevation errors, commonly termed "penetration bias." We leverage existing physics-based mode...

📖 Read original article


411. PD$^3$: A Project Duplication Detection Framework via Adapted Multi-Agent Debate ​

Author: Dezheng Bao, Yueci Yang, Chutian Yu, Xin Chen, Zeguo Fei, Xiang Yuan, Lijun Zhang, Jiangqian Huang, Zhengxuan Jiang, Daoze Zhang, Junru Chen, Yang Yang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2505.17492v2 Announce Type: replace Abstract: Project duplication detection is critical for project quality assessment because it helps avoid investment in repeated proposals. Existing methods usually cast it as ranking and rely on surface matching or direct large language models judging, ofte...

📖 Read original article


412. LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants ​

Author: Haochen Huang, Yue Su, Xin Sun, Moonisa Ahsan, Mohammad Aliannejadi, Irene Viola, Zhaochun Ren, Chuang Yu, Aneta Lisowska, Artem Belopolsky, Koen Hindriks, Pablo Cesar, Junxiao Wang, Jiahuan Pei
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV

arXiv:2507.05515v4 Announce Type: replace Abstract: Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precise object state detection are required. In this work, we explore LEGO...

📖 Read original article


413. Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models ​

Author: Vishal Raman, Vijai Aravindh R, Abhijith Ragav
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2507.21438v2 Announce Type: replace Abstract: Ontologies and knowledge graphs require continuous evolution to remain comprehensive and accurate, but manual curation is labor intensive. Large Language Models (LLMs) possess vast unstructured knowledge but struggle with maintaining structured con...

📖 Read original article


414. Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty ​

Author: Rui Wang, Qihan Lin, Jiayu Liu, Qing Zong, Tianshi Zheng, Dadi Guo, Haochen Shi, Peixuan Han, Weiqi Wang, Yangqiu Song
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2508.08992v4 Announce Type: replace Abstract: Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty. Although recent studies have develo...

📖 Read original article


415. TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning ​

Author: Chuang Jiang, Mingyue Cheng, Xiaoyu Tao, Qingyang Mao, Jie Ouyang, Qi Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2509.06278v4 Announce Type: replace Abstract: Table reasoning requires models to jointly perform comprehensive semantic understanding and precise numerical operations. Although recent large language model (LLM)-based methods have achieved promising results, most of them still rely on a single-...

📖 Read original article


416. Decentralized Causal Discovery using Judo Calculus ​

Author: Sridhar Mahadevan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2510.23942v2 Announce Type: replace Abstract: We describe a theory and implementation of an intuitionistic decentralized framework for causal discovery using judo calculus, which is formally defined as j-stable causal inference using j-do-calculus in a topos of sheaves. In real-world applicati...

📖 Read original article


417. OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows ​

Author: Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.HC

arXiv:2510.24411v3 Announce Type: replace Abstract: Computer-using agents powered by Vision-Language Models (VLMs) have demonstrated human-like capabilities in operating digital environments like mobile platforms. While these agents hold great promise for advancing digital automation, their potentia...

📖 Read original article


418. Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering ​

Author: Zijian Fu, Changsheng Lv, Xianlin Zhang, Mengshi Qi, Huadong Ma
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.23304v2 Announce Type: replace Abstract: In this paper, we propose a novel Multi-Modal Scene Graph with Kolmogorov-Arnold Expert Network for Audio-Visual Question Answering (SHRIKE). The task aims to mimic human reasoning by extracting and fusing information from audio-visual scenes, with...

📖 Read original article


419. Multi-agent DRL-based Lane Change Decision Model for Cooperative Platooning in Mixed Traffic ​

Author: Zeyu Mu, Shangtong Zhang, B. Brian Park
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.11809v2 Announce Type: replace Abstract: Connected automated vehicles (CAVs) possess the ability to communicate and coordinate with one another, enabling cooperative platooning that enhances both energy efficiency and traffic flow. However, during the initial stage of CAV deployment, the ...

📖 Read original article


420. Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models ​

Author: Moule Lin, Shuhao Guan, Andrea Patane, David Gregg, Goetz Botterweck
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.21003v3 Announce Type: replace Abstract: Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is especially severe when fine-tuned on small datasets due to the inherent tendency toward miscalibration. In th...

📖 Read original article


421. The Lattice Representation Hypothesis of Large Language Models ​

Author: Bo Xiong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.01227v3 Announce Type: replace Abstract: We propose the Lattice Representation Hypothesis of large language models: a symbolic backbone that grounds conceptual hierarchies and logical operations in embedding geometry. Our framework unifies the Linear Representation Hypothesis with Formal ...

📖 Read original article


422. PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms ​

Author: Wei Wang, Tianyu Shi, Shuai Zhang, Boyang Xia, Zequn Xie, Chenyu Zeng, Qi Zhang, Lynn Ai, Yaqi Yu, Kaiming Zhang, Feiyue Tang, Lei Ding
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2603.27476v2 Announce Type: replace Abstract: AI-powered people search platforms are increasingly used in recruiting, sales prospecting, and professional networking, yet no widely accepted benchmark exists for evaluating their performance. We introduce PeopleSearchBench, an open-source benchma...

📖 Read original article


423. LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent ​

Author: Bince Qu, Wanli Li, Bo Pan, Jianyu Zhang, Zheng Liu, Pan Zhang, Wei Chen, Bo Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.17931v5 Announce Type: replace Abstract: Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains constrained by two coupled challenges: hand-crafted synthetic data fails to elicit genuine real-worl...

📖 Read original article


424. From Context to Skills: Can Language Models Learn from Context Skillfully? ​

Author: Shuzheng Si, Haozhe Zhao, Yu Lei, Qingyi Wang, Dingwei Chen, Zhitong Wang, Zhenhailong Wang, Kangyang Luo, Zheng Wang, Gang Chen, Fanchao Qi, Minjia Zhang, Maosong Sun
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.27660v4 Announce Type: replace Abstract: Many real-world tasks require language models (LMs) to reason over complex contexts that exceed their parametric knowledge. This calls for context learning, where LMs directly learn relevant knowledge from the given context. An intuitive solution i...

📖 Read original article


425. GeoDecider: An Evidence-Grounded Agent for Geological Interpretation via Deliberative Reasoning ​

Author: Xiaoyu Tao, Mingyue Cheng, Jiahao Wang, Yitong Zhou, Qingyang Mao, Yimin Dou, Qi Liu, Shijin Wang, Enhong Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.03383v2 Announce Type: replace Abstract: Geological interpretation infers subsurface properties and structures from indirect geophysical observations. Well-log classification provides a measurable setting by assigning geological classes to depth-indexed petrophysical records. The task is ...

📖 Read original article


426. From Pixels to Prompts: Vision-Language Models ​

Author: Khang Nhat Hoang Vo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.07544v3 Announce Type: replace Abstract: When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so long ago. Teaching machines to see was already hard. Teaching them to read and generate language was already har...

📖 Read original article


427. Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas ​

Author: Nils A. Herrmann, Leander Girrbach, Kirill Bykov, Zeynep Akata
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.09159v2 Announce Type: replace Abstract: Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors". Prior work has used such directions as static handles for behavioral steering. We in...

📖 Read original article


428. What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation ​

Author: Xiang Wang, Wei Wei
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.26795v2 Announce Type: replace Abstract: Chain-of-thought (CoT) prompting enhances large language model performance, yet what drives these gains remains unclear. We study this question from a probe-time perspective: holding CoT rationales fixed, we test which textual properties matter for...

📖 Read original article


429. On the Detection of Commutative Factors in Factor Graphs: Necessary and Sufficient Conditions ​

Author: Malte Luttermann, Ralf M"oller, Marcel Gehrke
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.DS, cs.LG

arXiv:2605.26908v2 Announce Type: replace Abstract: Exploiting the indistinguishability of objects in a probabilistic graphical model such as a factor graph is key to lifted probabilistic inference algorithms and allows for tractable probabilistic inference problems with respect to domain sizes. A c...

📖 Read original article


430. WIRE: Profiling Witnessed Within-Policy Instruction Collisions in LLM Agents ​

Author: Lu Yan, Xuan Chen, Xiangyu Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.27784v2 Announce Type: replace Abstract: LLM agents are governed by long-lived prompt policies, where individually reasonable stand- ing rules can jointly govern the same pre- generation state. Existing instruction-following evaluations usually ask whether a model satis- fies explicit con...

📖 Read original article


431. Universal Quantum Transformer ​

Author: Sungyong Chung, Alireza Talebpour
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, quant-ph

arXiv:2606.00045v3 Announce Type: replace Abstract: Classical continuous-space neural networks fundamentally struggle to lock into exact formal rules, whether mathematical, such as modular arithmetic and non-Abelian group algebra, or linguistic, such as systematic compositional generalization. To ap...

📖 Read original article


432. Tree-Based Formalization of Multi-Agent Complementarity in Human-AI Interactions ​

Author: Andrea Ferrario
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, math.CO

arXiv:2606.04779v2 Announce Type: replace Abstract: Complementarity is the case in which a human--AI interaction (HAI) outperforms the best prediction benchmark available among its members. Although this idea is central in HAI research, formal work on complementarity remains limited. Existing framew...

📖 Read original article


433. Some hypotheses on how chatbots work in problem-solving-driven conversations. Large Language Models as confirmation of the Innovation Illusion ​

Author: S. F. M. van Vlijmen, H. D. Lethe jr
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.07722v3 Announce Type: replace Abstract: We discuss the nature of chatbots as conversation partners in problem-solving. What can chatbots do and what can't they do? We develop hypotheses on how this can this be explained. Our argument draws on insights from Aggregation Dynamics, Cognitive...

📖 Read original article


434. Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery ​

Author: Syed Rifat Raiyan, Mohsinul Kabir, Hasan Mahmud, Md Kamrul Hasan, Sophia Ananiadou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.LG

arXiv:2606.08728v4 Announce Type: replace Abstract: Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it has moved from a niche problem within NLP to one of the most consequential AI frontiers. This survey provides a unified account of the fiel...

📖 Read original article


435. ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents ​

Author: Ander Alvarez, Santhiya Rajan, Samuel Mugel, Rom'an Or'us
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA

arXiv:2606.18037v2 Announce Type: replace Abstract: Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, databases, clinical records, and formulary tools. Standard factuality metrics usually test whether an ans...

📖 Read original article


436. Intent-Governed Tool Authorization for AI Agents ​

Author: Genliang Zhu, Chu Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.22916v2 Announce Type: replace Abstract: AI agents increasingly act through external tools: they read private data, construct structured payloads, submit write requests, export records, and coordinate workflows across application boundaries. Existing authorization mechanisms usually ask w...

📖 Read original article


437. The Hitchhiker's Guide to Agentic AI: From Foundations to Systems ​

Author: Haggai Roitman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, cs.LG

arXiv:2606.24937v2 Announce Type: replace Abstract: The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. The book covers the full stack from first principles to production deployment, organized around a central thesis: building great ag...

📖 Read original article


438. ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents ​

Author: Qitai Tan, Zefang Zong, Mo Li, Yipeng Shi, Yang Li, Peng Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.27814v2 Announce Type: replace Abstract: Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the early stage, but ...

📖 Read original article


439. EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures ​

Author: Bu\u{g}ra Alperen Ulu{\i}rmak, Rifat Kurban
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.SE

arXiv:2606.30219v4 Announce Type: replace Abstract: This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capabiliti...

📖 Read original article


440. Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents ​

Author: Xutao Mao, Liangjie Zhao, Leyao Wang, Rui Qian, Qiang Huang, Wentao Wang, Bo Han, Xiang Zheng, Cong Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.10526v3 Announce Type: replace Abstract: Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conversational sycophancy into a state-writing failure: accepted user-centric claims can be committed as lasting ...

📖 Read original article


441. Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations ​

Author: Samer Saab Jr, Chaouki Abdallah
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.12077v2 Announce Type: replace Abstract: Multi-agent language-model (LM) systems often determine which agents communicate, yet routing is usually treated as an implementation detail. We ask whether routing itself determines whether a population converges on a shared convention or fragment...

📖 Read original article


442. ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System ​

Author: Yutong He, Daibo Li, Guohong Li, Jiahe Geng, Zhengyang Huang, Can Ren, Zekun Zhang, Yifan Liu, Shuchen Zhu, Hengrui Zhang, Boao Kong, Ming Sun, Shu Li, Chenyi Li, Jiang Hu, Kun Yuan, Zaiwen Wen, Pingwen Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.14178v2 Announce Type: replace Abstract: Recent advances in Large Language Models have fueled autonomous AI agents capable of tackling complex scientific tasks, yet existing automated research systems remain predominantly focused on empirically driven domains with quantitative benchmarks,...

📖 Read original article


443. MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection ​

Author: Goktug Ozkan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.15166v2 Announce Type: replace Abstract: Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a synthetic benchmark and failure atlas built by a clinician. The resource labels medical ...

📖 Read original article


444. Knowledge-Centric Agents for Workflow Generation in ComfyUI ​

Author: Zhendong Li, Lei Sun, Ruibo Ming, He Zhang, Danda Pani Paudel, Luc Van Gool, Jinjin Gu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.15845v2 Announce Type: replace Abstract: Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large language model (LLM) approaches often treat this as a direct text-to-JSON g...

📖 Read original article


445. Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII) ​

Author: Li Qiwei, Wells Lucas Santo, Sarita Schoenebeck, Eric Gilbert
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.18263v2 Announce Type: replace Abstract: AI-generated non-consensual intimate imagery (AIG-NCII) is not adequately addressed in AI/ML literature regarding AI-generated media, commonly referred to as "deepfakes". While research on deepfakes currently focuses on its epistemic harms -- or ha...

📖 Read original article


446. Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability ​

Author: Taewoon Kim, Vincent Fran\c{c}ois-Lavet, Michael Cochez
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.18368v3 Announce Type: replace Abstract: Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-policy that learns which symbolic memory heuristic to apply at each decision point while keeping execut...

📖 Read original article


447. Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction ​

Author: Jialian Li, Junhong Liu, Yuchen Cao, Weiran Guo, Jiaming Song, Xutao Wang, Yi Zhao, Jiangpin Liu, Jie Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.18985v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, there is a growing demand for compact models that can serve as an on-d...

📖 Read original article


448. Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs ​

Author: Harry Rogers, Sally Shiels, Ashley Tomlinson, James Thomas, James Aylward, Nathan Gauge, Helen Higham, Alison Noble
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.19063v2 Announce Type: replace Abstract: Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and cognitive bias. Standard examiner validation via inter-rater statistic...

📖 Read original article


449. AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally ​

Author: Shaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen, Shaofan Liu, Suncong Zheng, Jian Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.19363v2 Announce Type: replace Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and leng...

📖 Read original article


450. Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks ​

Author: Christian Arzate Cruz, Stefanos Gkikas, Houshyar Asadi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.20820v2 Announce Type: replace Abstract: Body-based emotion recognition is important for real-time affective systems, but graph-based skeleton models can be computationally expensive. This paper studies whether lightweight temporal convolutional networks (TCNs) can provide an efficient an...

📖 Read original article


451. EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization ​

Author: Lihuang Fang, Yuchen Zou, kebing Jin, Jinghui Qin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2607.21013v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understanding abilities and natural ...

📖 Read original article


452. V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure ​

Author: Zhetong Zhang, Honghao Fu, Miao Xu, Yiwei Wang, Yujun Cai
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.21151v2 Announce Type: replace Abstract: As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuitively, we find that harmful videos paired with benign queries achieve higher attack success rates...

📖 Read original article


453. Expert Behavior Prior Reinforcement Learning ​

Author: Gong Gao, Weidong Zhao, Xianhui Liu, Ning Jia
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.21302v2 Announce Type: replace Abstract: Behavior prior reinforcement learning (BPRL) has emerged as a promising paradigm to improve sample efficiency in online reinforcement learning (RL) by leveraging policy priors derived from offline demonstrations. However, most existing BPRL methods...

📖 Read original article


454. Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model ​

Author: Nanbeige Lab, :, Chen Yang, Chengrui Huang, Fufeng Lan, Hanhui Chen, Hao Zhou, Huatong Song, Jiaqi Cao, Jiaying Zhu, Jinlin Niu, Kai Wang, Lisheng Huang, Qiliang Liang, Ran Le, Ruixiang Feng, Shuang Sun, Tao Gu, Tao Zhang, Tianyu Luo, Yang Song, Yun Xing, Yuntao Wen, Ziyao Xu, Zongchao Chen, Zongqiang Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.22083v2 Announce Type: replace Abstract: We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in m...

📖 Read original article


455. TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI ​

Author: Ritik Raj, Souvik Kundu, Sarbartha Banerjee, Dheemanth Joshi, Ishita Vohra, Tushar Krishna
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA

arXiv:2607.22465v2 Announce Type: replace Abstract: Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, agentic a...

📖 Read original article


456. Like a Baby: Visually Situated Neural Language Acquisition ​

Author: Alexander G. Ororbia, Ankur Mali, Mary Alexandria Kelly, David Reitter
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:1805.11546v3 Announce Type: replace-cross Abstract: We examine the benefits of visual context in training neural language models to perform next-word prediction. A multi-modal neural architecture is introduced that outperform its equivalent trained on language alone with a 2% decrease in perpl...

📖 Read original article


457. Speed Reading Tool Powered by Artificial Intelligence for Students with ADHD, Dyslexia, and Short Attention Span ​

Author: Megat Irfan Zackry Bin Ismail, Ahmad Nazran bin Yusri, Muhammad Hafizzul Bin Abdul Manap, Muhammad Muizzuddin Bin Kamarozaman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2307.14544v2 Announce Type: replace-cross Abstract: This paper presents an artificial intelligence tool designed to assist students with dyslexia, ADHD, and short attention spans in processing text-based information more efficiently. The proposed solution addresses both cognitive and visual re...

📖 Read original article


458. Fairness Interventions in Classification: A Study on AI Explainability ​

Author: Thomas Souverain, Paul 'Egr'e
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY

arXiv:2407.14766v4 Announce Type: replace-cross Abstract: This paper presents a philosophical and experimental study of fairness interventions in AI classification, centered on the explainability and transparency of corrective methods, and on the opposition between two fairness criteria, namely Demo...

📖 Read original article


459. CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases ​

Author: Xiangyan Liu, Bo Lan, Zhiyuan Hu, Yang Liu, Zhicheng Zhang, Fei Wang, Michael Shieh, Wenmeng Zhou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2408.03910v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories. This challenge has prompted research on enhancing LLM-codebase interaction at a repository scale. Curre...

📖 Read original article


460. Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks ​

Author: Rui Hu, Yifan Zhang, Zhuoran Li, Longbo Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2410.02596v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) are a novel class of generative models designed to sample from unnormalized distributions and have found applications in various important tasks, attracting great research interest in their training algori...

📖 Read original article


461. CausAdv: A Causal-based Framework for Detecting Adversarial Examples ​

Author: Hichem Debbi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, stat.ME, stat.ML

arXiv:2411.00839v4 Announce Type: replace-cross Abstract: Deep learning has led to tremendous success in computer vision, largely due to Convolutional Neural Networks (CNNs). However, CNNs have been shown to be vulnerable to crafted adversarial perturbations. This vulnerability of adversarial exampl...

📖 Read original article


462. LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models ​

Author: Nam V. Nguyen, Thong T. Doan, Luong Tran, Van Nguyen, Quang Pham
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2411.00918v5 Announce Type: replace-cross Abstract: Mixture of experts (MoE) architectures have become a cornerstone for scaling up and are a key component in most large language models such as GPT-OSS, DeepSeek-V3, Llama-4, and Gemini-2.5. However, systematic research on MoE remains severely ...

📖 Read original article


463. Sign-Symmetry Learning Rules are Robust Fine-Tuners ​

Author: Aymene Berriche, Mehdi Zakaria Adjal, Riyadh Baghdadi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2502.05925v2 Announce Type: replace-cross Abstract: Backpropagation (BP) has long been the predominant method for training neural networks due to its effectiveness. However, numerous alternative approaches, broadly categorized under feedback alignment, have been proposed, many of which are mot...

📖 Read original article


464. A Survey of Graph Transformers: Architectures, Theories and Applications ​

Author: Chaohao Yuan, Kangfei Zhao, Ercan Engin Kuruoglu, Liang Wang, Tingyang Xu, Wenbing Huang, Deli Zhao, Hong Cheng, Yu Rong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2502.16533v3 Announce Type: replace-cross Abstract: Graph Transformers (GTs) have demonstrated a strong capability in modeling graph structures by addressing the intrinsic limitations of graph neural networks (GNNs), such as over-smoothing and over-squashing. Recent studies have proposed diver...

📖 Read original article


465. Sampling Decisions: Exact Path-Space Correction, Prior Cancellation and Local-Boltzmann Guidance ​

Author: Michael Chertkov, Sungsoo Ahn, Hamidreza Behjoo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.stat-mech, cs.AI, cs.SY, eess.SY, stat.ML

arXiv:2503.14549v3 Announce Type: replace-cross Abstract: How can a cheap but biased sequential, finite-horizon sampler over a discrete space be corrected so that its terminal output follows a prescribed Gibbs distribution? We formulate Sampling Decisions as a path-space relative-entropy projection ...

📖 Read original article


466. Kinship Verification through a Forest Neural Network ​

Author: Ali Nazari, Omidreza Borzoei, Mohsen Ebrahimi Moghaddam
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2504.18910v2 Announce Type: replace-cross Abstract: Early methods used face representations in kinship verification, which are less accurate than joint representations of parents' and children's facial images learned from scratch. We propose an approach featuring graph neural network concepts ...

📖 Read original article


467. Computational Experiments in Number Theory ​

Author: Ali Saraeb
Published: 7/28/2026, 4:00:00 AM
Categories: math.NT, cs.AI

arXiv:2504.19451v4 Announce Type: replace-cross Abstract: This paper presents two concrete applications of Artificial Intelligence to algorithmic and analytic number theory. Recent benchmarks of large language models have mainly focused on general mathematics problems and the currently infeasible ob...

📖 Read original article


468. Steerable Chatbots: Exploring Personalization Control Interfaces via LLM Activation Steering ​

Author: Jessica Y. Bo, Tianyu Xu, Ishan Chatterjee, Katrina Passarella-Ward, Achin Kulshrestha, D Shin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2505.04260v3 Announce Type: replace-cross Abstract: Personalizing LLM responses typically requires users to articulate their preferences through prompting, which can be burdensome at cold start and difficult to articulate in natural language. We introduce an alternative paradigm, steerable cha...

📖 Read original article


469. AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use ​

Author: Yaotian Yang, Yiwen Tang, Yizhe Chen, Xiao Chen, Jiangjie Qiu, Hao Xiong, Haoyu Yin, Zhiyao Luo, Yifei Zhang, Sijia Tao, Wentao Li, Qinghua Zhang, Yuqiang Li, Wanli Ouyang, Bin Zhao, Xiaonan Wang, Fei Wei
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2505.12650v2 Announce Type: replace-cross Abstract: Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can explain similar contrast, and purely feed-forward models cannot verify physical validity. We present AutoMa...

📖 Read original article


470. CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning ​

Author: Jiaxun Cui, Chen Tang, Jarrett Holtz, Janice Nguyen, Alessandro G. Allievi, Hang Qiu, Peter Stone
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.MA

arXiv:2505.18334v2 Announce Type: replace-cross Abstract: Past work has demonstrated that autonomous vehicles can drive more safely if they communicate with each other. However, this communication is usually not human-understandable. Using natural language as a vehicle-to-vehicle (V2V) communication...

📖 Read original article


471. Retrieval-Augmented Generation of Ontologies from Relational Databases ​

Author: Nadeen Fathallah, Mojtaba Nayyeri, Athish A Yogi, Ratan Bahadur Thapa, Hans-Michael Tautenhahn, Anton Schnurpel, Steffen Staab
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2506.01232v2 Announce Type: replace-cross Abstract: Deriving OWL ontologies from relational database schemas supports semantic interoperability and downstream tasks such as knowledge graph population, ontology-based data access, graph-based learning, and automated reasoning. Existing approache...

📖 Read original article


472. Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages ​

Author: Ali Almutairi, Abdullah Alsuhaibani, Shoaib Jameel, Aditya Joshi, Gelareh Mohammadi, Imran Razzak
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2506.10292v2 Announce Type: replace-cross Abstract: Training deep learning networks with minimal supervision has gained significant research attention due to its potential to reduce reliance on extensive labelled data. While self-training methods have proven effective in semi-supervised learni...

📖 Read original article


473. LEDOM: Reverse Language Model ​

Author: Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu, Li Lin, Xinyi Wang, Liangming Pan, William Yang Wang, Xiaojun Wan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2507.01335v4 Announce Type: replace-cross Abstract: Autoregressive language models are trained exclusively left-to-right. We explore the complementary factorization, training right-to-left at scale, and ask what reasoning patterns emerge when a model conditions on future context to predict the...

📖 Read original article


474. CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback ​

Author: Qiushi Sun, Jinyang Gong, Lei Li, Qipeng Guo, Fei Yuan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2507.22080v2 Announce Type: replace-cross Abstract: Acquiring high-quality instruction-code pairs is essential for training Large Language Models for code generation. While automated synthesis has emerged as an alternative to expensive manual curation, current approaches often rely on rigid he...

📖 Read original article


475. Realizing Scaling Laws in Recommender Systems: A Foundation-Expert Paradigm for Hyperscale Model Deployment ​

Author: Dai Li, Kevin Course, Wei Li, Hongwei Li, Jie Hua, Yiqi Chen, Zhao Zhu, Rui Jian, Xuan Cao, Bi Xue, Yu Shi, Jing Qian, Kai Ren, Matt Ma, Qunshu Zhang, Rui Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG

arXiv:2508.02929v3 Announce Type: replace-cross Abstract: Scaling laws have been established for recommender systems, yet efficiently deploying foundation model (FM) across multiple recommendation surfaces remains a major unsolved challenge. Existing methods for transfer learning face fundamental li...

📖 Read original article


476. PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment ​

Author: Chang Hong, Minghao Wu, Qingying Xiao, Yuchi Wang, Xiang Wan, Guangjun Yu, Benyou Wang, Yan Hu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2508.05132v3 Announce Type: replace-cross Abstract: As medical LLMs transition to clinical deployment, assessing their ethical reasoning capability becomes critical. While achieving high accuracy on knowledge benchmarks, LLMs lack validated assessment for navigating ethical trade-offs in clini...

📖 Read original article


477. EGRA:Toward Enhanced Behavior Graphs and Representation Alignment for Multimodal Recommendation ​

Author: Xiaoxiong Zhang, Xin Zhou, Zhiwei Zeng, Yongjie Wang, Zhiqi Shen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2508.16170v3 Announce Type: replace-cross Abstract: MultiModal Recommendation (MMR) systems have emerged as a promising solution for improving recommendation quality by leveraging rich item-side modality information, prompting a surge of diverse methods. Despite these advances, existing method...

📖 Read original article


478. OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries ​

Author: Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
Published: 7/28/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI, cs.ET, cs.IR

arXiv:2509.02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-s...

📖 Read original article


479. Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers ​

Author: Xingyue Huang, Rishabh, Gregor Franke, Ziyi Yang, Jiamu Bai, Weijie Bai, Jinhe Bi, Zifeng Ding, Yiqun Duan, Chengyu Fan, Wendong Fan, Xin Gao, Ruohao Guo, Yuan He, Zhuangzhuang He, Xianglong Hu, Neil Johnson, Bowen Li, Fangru Lin, Siyu Lin, Tong Liu, Yunpu Ma, Hao Shen, Hao Sun, Beibei Wang, Fangyijie Wang, Hao Wang, Haoran Wang, Yang Wang, Yifeng Wang, Zhaowei Wang, Ziyang Wang, Yifan Wu, Zikai Xiao, Chengxing Xie, Fan Yang, Junxiao Yang, Qianshuo Ye, Ziyu Ye, Guangtao Zeng, Yuwen Ebony Zhang, Zeyu Zhang, Zihao Zhu, Bernard Ghanem, Philip Torr, Guohao Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2509.03059v2 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can be significantly improved through Reinforcement Learning with Verifiable Reward (RLVR), particularly in domains like mathematics and programming,...

📖 Read original article


480. A funny companion: Distinct neural responses to AI- versus human-attributed humor ​

Author: Xiaohui Rao, Hanlin Wu, Zhenguang G. Cai
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2509.10847v3 Announce Type: replace-cross Abstract: As artificial intelligence (AI) companions become capable of human-like communication, including telling jokes, understanding how people cognitively and affectively respond to AI-attributed humor becomes increasingly important. This study use...

📖 Read original article


481. Poison to Detect: Detection of Targeted Overfitting in Federated Learning ​

Author: Soumia Zohra El Mestari, Maciej Krzysztof Zuziak, Gabriele Lenzini
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2509.11974v3 Announce Type: replace-cross Abstract: Federated Learning (FL) enables collaborative model training among clients without centralising data, making it a widely adopted privacy-enhancing technology (PET). Despite its privacy benefits, FL remains vulnerable to orchestrator-driven pr...

📖 Read original article


482. From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety ​

Author: Shucheng Zhang, Yan Shi, Bingzhang Wang, Yuang Zhang, Muhammad Monjurul Karim, Kehua Chen, Chenxi Liu, Mehrdad Nasri, Yinhai Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2510.03314v2 Announce Type: replace-cross Abstract: Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infrastructure-based measures are often insufficient in dynamic urban environments. Recent advances in learni...

📖 Read original article


483. LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish ​

Author: Fred Philippy, Laura Bernardy, Siwen Guo, Jacques Klein, Tegawend'e F. Bissyand'e
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2510.07074v2 Announce Type: replace-cross Abstract: Instruction tuning has become a key technique for enhancing the performance of large language models, enabling them to better follow human prompts. However, low-resource languages such as Luxembourgish face severe limitations due to the lack ...

📖 Read original article


484. Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling ​

Author: Alexandru Meterez, Depen Morwani, Jingfeng Wu, Costin-Andrei Oncescu, Cengiz Pehlevan, Sham Kakade
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC, stat.ML

arXiv:2510.14717v2 Announce Type: replace-cross Abstract: Increasing the batch size during training -- a ''batch ramp'' -- is a promising strategy to accelerate large language model pretraining. While for SGD, doubling the batch size can be equivalent to halving the learning rate, the optimal strate...

📖 Read original article


485. Continual Knowledge Consolidation LORA for Domain Incremental Learning ​

Author: Naeem Paeedeh, Mahardhika Pratama, Weiping Ding, Jimmy Cao, Wolfgang Mayer, Ryszard Kowalczyk, Ary Shiddiqi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2510.16077v2 Announce Type: replace-cross Abstract: Domain Incremental Learning (DIL) is a sub-branch of continual learning that aims to address the never-ending arrival of new domains without catastrophic forgetting. Despite the advent of parameter-efficient fine-tuning (PEFT) approaches, pri...

📖 Read original article


486. Intuitionistic $j$-Do-Calculus in Topos Causal Models ​

Author: Sridhar Mahadevan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LO, cs.AI

arXiv:2510.17944v2 Announce Type: replace-cross Abstract: In this paper, we generalize Pearl's do-calculus to an Intuitionistic setting called $j$-stable causal inference inside a topos of sheaves. Our framework is an elaboration of the recently proposed framework of Topos Causal Models (TCMs), wher...

📖 Read original article


487. CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference ​

Author: Dong Liu, Shu Wang, Yanxuan Yu, Haisheng Wang, Ben Lengerich
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2511.21702v2 Announce Type: replace-cross Abstract: Large language models face significant computational bottlenecks during inference due to the expensive output layer computation over large vocabularies. We present CSV-Decode, a novel approach that uses geometric upper bounds to construct sma...

📖 Read original article


488. VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference ​

Author: Jiaming Tang, Yufei Sun, Yilong Zhao, Shang Yang, Yujun Lin, Zhuoyang Zhang, James Hou, Yao Lu, Zhijian Liu, Song Han
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2512.01031v2 Announce Type: replace-cross Abstract: Vision-Language-Action models (VLAs) are becoming increasingly capable across diverse robotic tasks. However, these models are typically deployed under synchronous inference, where the robot waits for model inference to complete before acting...

📖 Read original article


489. GFLAN: Generative Functional Layouts ​

Author: Mohamed Abouagour, Eleftherios Garyfallidis
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2512.16275v2 Announce Type: replace-cross Abstract: Automated floor plan generation lies at the intersection of combinatorial search, geometric constraint satisfaction, and functional design requirements -- a confluence that has historically resisted a unified computational treatment. While re...

📖 Read original article


490. Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training ​

Author: Yi Liu, Sukai Wang, Dafeng Wei, Xiaowei Cai, Linqing Zhong, Jiange Yang, Guanghui Ren, Jinyu Zhang, Maoqing Yao, Chuankang Li, Xindong He, Liliang Chen, Jianlan Luo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2512.24125v3 Announce Type: replace-cross Abstract: General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challenging for existing Vision-Language-Action (VLA) models. While la...

📖 Read original article


491. AutoFed: Personalized Federated Traffic Prediction via Adaptive Prompt ​

Author: Zijian Zhao, Yitong Shang, Sen Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2512.24625v3 Announce Type: replace-cross Abstract: Accurate traffic prediction is essential for Intelligent Transportation Systems, including ride-hailing, urban road planning, and vehicle fleet management. However, due to significant privacy concerns surrounding traffic data, most existing m...

📖 Read original article


492. The Optimal Sample Complexity of Linear Contracts ​

Author: Mikael M{\o}ller H{\o}gsgaard
Published: 7/28/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.LG

arXiv:2601.01496v3 Announce Type: replace-cross Abstract: In this paper, we settle the problem of learning optimal linear contracts from data in the offline setting, where agent types are drawn from an unknown distribution and the principal's goal is to design a contract that maximizes her expected ...

📖 Read original article


493. TextBridgeGNN: Pre-training Graph Neural Network for Cross-Domain Recommendation via Text-Guided Transfer ​

Author: Yiwen Chen, Yiqing Wu, Huishi Luo, Fuzhen Zhuang, Deqing Wang, Zhao Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2601.02366v3 Announce Type: replace-cross Abstract: Graph-based recommendation has achieved great success in recent years. The classical graph recommendation model utilizes ID embedding to store essential collaborative information. However, this ID-based paradigm faces challenges in transferri...

📖 Read original article


494. Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models ​

Author: Christopher Ormerod
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.02580v2 Announce Type: replace-cross Abstract: Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibration. This study introduc...

📖 Read original article


495. Infinite-Precision Autoregressive Modeling for Vector Graphics and Layouts ​

Author: Yeonsang Shin, Insoo Kim, Bongkeun Kim, Keonwoo Bae, Bohyung Han
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2601.05680v2 Announce Type: replace-cross Abstract: While Transformer-based autoregressive models excel in data generation, their token discretization strategy inherently limits their precision in continuous domains. We analyze the scalability limitations of existing discretization-based appro...

📖 Read original article


496. A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning ​

Author: Jiali Cheng, Ziheng Chen, Chirag Agarwal, Hadi Amiri
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CV

arXiv:2601.09624v2 Announce Type: replace-cross Abstract: Machine unlearning is becoming essential for building trustworthy and compliant language models. Yet unlearning success varies considerably across individual samples: some are reliably erased, while others persist despite the same procedure. ...

📖 Read original article


497. Ordering-based Causal Discovery via Generalized Score Matching ​

Author: Vy Vo, He Zhao, Trung Le, Edwin V. Bonilla, Dinh Phung
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.16249v3 Announce Type: replace-cross Abstract: Learning DAG structures from purely observational data remains a long-standing challenge across scientific domains. An emerging line of research leverages the score of the data distribution to initially identify a topological order of the und...

📖 Read original article


498. MANGO: A Global Single-Date Paired Dataset for Mangrove Segmentation ​

Author: Junhyuk Heo, Beomkyu Choi, Hyunjin Shin, Darongsae Kwon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2601.17039v2 Announce Type: replace-cross Abstract: Mangroves are critical for climate-change mitigation, requiring reliable monitoring for effective conservation. While deep learning has emerged as a powerful tool for mangrove detection, its progress is hindered by the limitations of existing...

📖 Read original article


499. Physics-Encoded Inverse Modeling for Arctic Snow Depth Estimation ​

Author: Akila Sampath, Vandana P. Janeja, Jianwu Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.17074v5 Announce Type: replace-cross Abstract: Accurate estimation of unobserved quantities in time-varying inverse problems remains challenging when observations are sparse and only indirectly related to the target variable. In Arctic climate applications, snow depth over sea ice is not ...

📖 Read original article


500. Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis ​

Author: Junyi Zou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.18350v5 Announce Type: replace-cross Abstract: Large language models fine-tuned via a two-stage pipeline (domain adaptation followed by instruction alignment) can exhibit non-trivial interference after adapter merging, including the re-emergence of explicit reasoning traces under strict d...

📖 Read original article


501. Action-Sufficient Goal Representations ​

Author: Jinu Hyeon, Woobin Park, Hongjoon Ahn, Taesup Moon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.22496v2 Announce Type: replace-cross Abstract: In offline goal-conditioned reinforcement learning (GCRL), hierarchical approaches decompose long-horizon tasks into high-level subgoal prediction and low-level action execution. A critical design choice in such architectures is the goal repr...

📖 Read original article


Author: Quang Truong, Yu Song, Donald Loveland, Mingxuan Ju, Tong Zhao, Neil Shah, Jiliang Tang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.01553v3 Announce Type: replace-cross Abstract: Link prediction is a core challenge in graph machine learning, demanding models that capture rich and complex topological dependencies. While Graph Neural Networks (GNNs) are the standard solution, state-of-the-art pipelines often rely on exp...

📖 Read original article


503. Towards Isolated Interventions via Almost Orthogonal Features in Language Models ​

Author: Moritz Miller, Florent Draye, Bernhard Sch"olkopf
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2602.04718v3 Announce Type: replace-cross Abstract: A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not ...

📖 Read original article


504. Multi-Task GRPO: Reliable LLM Reasoning Across Tasks ​

Author: Shyam Sundhar Ramesh, Xiaotong Ji, Matthieu Zimmer, Sangwoong Yoon, Zhiyong Wang, Haitham Bou Ammar, Aurelien Lucchi, Ilija Bogunovic
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2602.05547v2 Announce Type: replace-cross Abstract: RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deployment requires reliable performance across diverse tasks. A straightforward multi-task adaptation of GRPO...

📖 Read original article


505. How College Students Use AI to Navigate Course Readings: Evidence from an Eight-Week Study ​

Author: Yue Fu, Joel Wester, Niels Van Berkel, Alexis Hiniker
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY

arXiv:2602.09907v3 Announce Type: replace-cross Abstract: College students increasingly use AI chatbots to support academic reading, yet we lack granular understanding of how these interactions shape their reading experience and cognitive engagement. We conducted an eight-week longitudinal study wit...

📖 Read original article


506. LLM-Based Scientific Equation Discovery via Physics-Informed Token-Regularized Policy Optimization ​

Author: Boxiao Wang, Kai Li, Tianyi Liu, Chen Li, Junzhe Wang, Yifan Zhang, Jian Cheng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.10576v2 Announce Type: replace-cross Abstract: Symbolic regression aims to distill mathematical equations from observational data. Recent approaches have successfully leveraged Large Language Models (LLMs) to generate equation hypotheses, capitalizing on their vast pre-trained scientific ...

📖 Read original article


507. The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs ​

Author: Samir Sadok, Laurent Girin, Xavier Alameda-Pineda
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2602.15491v2 Announce Type: replace-cross Abstract: Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly within the same latent space. As a result, they are poorly robust to a global variation of the input sig...

📖 Read original article


508. AIFL: A Global Daily Streamflow Forecasting Model Using a Deterministic LSTM Pre-trained on ERA5-Land and Fine-tuned on IFS ​

Author: Maria Luisa Taccari, Kenza Tazi, Ois'in M. Morrison, Andreas Grafberger, Juan Colonese, Corentin Carton de Wiart, Christel Prudhomme, Cinzia Mazzetti, Matthew Chantry, Florian Pappenberger
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.app-ph

arXiv:2602.16579v2 Announce Type: replace-cross Abstract: Reliable global streamflow forecasting is essential for flood preparedness and water resource management, yet data-driven models often suffer from a performance gap when transitioning from historical reanalysis to operational forecast product...

📖 Read original article


509. Voice-Driven Semantic Perception for UAV-Assisted Emergency Networks ​

Author: Nuno Saavedra, Pedro Ribeiro, Andr'e Coelho, Rui Campos
Published: 7/28/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.SD

arXiv:2602.17394v2 Announce Type: replace-cross Abstract: Unmanned Aerial Vehicle (UAV)-assisted networks are increasingly foreseen as a promising approach for emergency response, providing rapid, flexible, and resilient communications in environments where terrestrial infrastructure is degraded or ...

📖 Read original article


510. Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting ​

Author: Xinghong Fu, Yanhong Li, Georgios Papaioannou, Yoon Kim
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.17634v2 Announce Type: replace-cross Abstract: Learning time series foundation models has been shown to be a promising approach for zero-shot time series forecasting across diverse time series domains. Insofar as scaling has been a critical driver of performance of foundation models in ot...

📖 Read original article


511. From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications ​

Author: Philipp Steigerwald, Jens Albrecht
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL, cs.CY

arXiv:2602.18443v2 Announce Type: replace-cross Abstract: Psychosocial online counselling frequently encounters generic subject lines that impede efficient case prioritisation. This study evaluates eleven large language models generating six-word subject lines for German counselling emails through h...

📖 Read original article


512. UP-Fuse: Uncertainty-guided LiDAR-Camera Fusion for 3D Panoptic Segmentation ​

Author: Rohit Mohan, Florian Drews, Yakov Miron, Daniele Cattaneo, Abhinav Valada
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2602.19349v2 Announce Type: replace-cross Abstract: LiDAR-camera fusion enhances 3D panoptic segmentation by leveraging camera images to complement sparse LiDAR scans, but it also introduces a critical failure mode. Under adverse conditions, degradation or failure of the camera sensor can sign...

📖 Read original article


513. Exact and Asymptotically Complete Robust Verifications of Neural Networks via Ising Solvers ​

Author: Wenxin Li, Wenchao Liu, Weihao Li, Chuan Wang, Qi Gao, Yin Ma, Hai Wei, Kai Wen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.optics, quant-ph

arXiv:2603.00408v2 Announce Type: replace-cross Abstract: We present an Ising-compatible framework for formal neural-network robustness verification under bounded input perturbations. For piecewise-linear activations, the Exact Logarithmic PWL Model (Log-PWL) provides an exact, sound, and complete f...

📖 Read original article


514. SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems ​

Author: Haochang Hao, Yifan Xu, Xinzhuo Li, Yingqiang Ge, Lu Cheng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2603.03536v2 Announce Type: replace-cross Abstract: Current LLM-based conversational recommender systems (CRS) primarily optimize recommendation accuracy and user satisfaction. We identify an underexplored vulnerability in which recommendation outputs may negatively impact users by violating p...

📖 Read original article


515. DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces ​

Author: Mohammad Sadil Khan, Muhammad Usama, Rolandos Alexandros Potamias, Didier Stricker, Muhammad Zeshan Afzal, Jiankang Deng, Ismail Elezi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.05607v2 Announce Type: replace-cross Abstract: Computer-Aided Design (CAD) relies on structured and editable geometric representations, yet existing generative methods are constrained by small annotated datasets with explicit design histories or boundary representation (BRep) labels. Mean...

📖 Read original article


516. Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning ​

Author: Kanishkha Jaisankar, Pranav M. Pawar, Diana Susan Joseph, Raja Muthalagu, Mithun Mukherjee, Dnyaneshawar Mantri, Ramjee Prasad
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.09255v2 Announce Type: replace-cross Abstract: Deep learning and computer vision techniques have become increasingly important in the development of self-driving cars. These techniques play a crucial role in enabling self-driving cars to perceive and understand their surroundings, allowin...

📖 Read original article


517. Designing Service Systems from Textual Evidence ​

Author: Ruicheng Ao, Hongyu Chen, Siyang Gao, Hanwei Li, David Simchi-Levi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC, stat.ML

arXiv:2603.10400v2 Announce Type: replace-cross Abstract: Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing policy, or the most effective quality control procedure. In many service systems, the primary evidence of ...

📖 Read original article


518. A General Deep Learning Framework for Wireless Resource Allocation under Discrete Constraints ​

Author: Yikun Wang, Yang Li, Yik-Chung Wu, Rui Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, math.IT

arXiv:2603.19322v2 Announce Type: replace-cross Abstract: While deep learning (DL)-based methods have achieved remarkable success in continuous wireless resource allocation, efficient solutions for problems involving discrete variables remain challenging. This is primarily due to the zero-gradient i...

📖 Read original article


519. Stability of AI Governance Systems: A Coupled Dynamics Model of Public Trust and Social Disruptions ​

Author: Jiaqi Lai, Hou Liang, Weihong Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC, cs.MA

arXiv:2603.20248v2 Announce Type: replace-cross Abstract: AI systems are increasingly entrenched in public governance, yet scholarship lacks formal tools to determine when deviations of public trust in algorithmic institutions dissipate and when they grow into collapse. Stability refers here to asym...

📖 Read original article


520. Coherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure ​

Author: Camilo Chac'on Sartori
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2603.28371v2 Announce Type: replace-cross Abstract: When an agent can articulate why something works, we typically take this as evidence of genuine understanding. This presupposes that effective action and correct explanation covary, and that coherent explanation reliably signals both. I argue...

📖 Read original article


521. AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models ​

Author: Mozhgan Pourkeshavarz, Tianran Liu, Nicholas Rhinehart
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.LG

arXiv:2603.28963v2 Announce Type: replace-cross Abstract: Simulation with realistic traffic agents is essential for validating autonomous driving systems. Existing data-driven simulators learn agent behavior from higher-level abstractions such as 3D bounding boxes and polylines, inferred by upstream...

📖 Read original article


522. Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments ​

Author: Zonghan Li, Feng Ji
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.ET

arXiv:2604.02458v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-world experiments are costly or infeasible. The treatment-effect estimates are often evaluated using stati...

📖 Read original article


523. SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills ​

Author: Yinghan Hou, Zongyou Yang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2604.06550v3 Announce Type: replace-cross Abstract: Agent skills combine natural-language instructions with executable code while inheriting an agent's filesystem, credential, and network access. Attacks can span prose and files, whereas regex and code-only analyzers cover only one modality. S...

📖 Read original article


524. Sheaf-Laplacian Obstruction and Projection Hardness for Cross-Modal Compatibility on a Modality-Independent Site ​

Author: Tibor Sloboda
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.07632v2 Announce Type: replace-cross Abstract: Cross-modal representations vary in how easily they can be aligned, and compatibility is generally non-transitive: two modalities may align through an intermediate modality at lower complexity than through a direct map. We introduce a referen...

📖 Read original article


525. Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex ​

Author: Alexandru Brateanu, Tingting Mu, Codruta Ancuti, Cosmin Ancuti
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.10359v3 Announce Type: replace-cross Abstract: Low-light image enhancement (LLIE) aims to restore natural visibility, color fidelity, and structural detail under severe illumination degradation. State-of-the-art (SOTA) LLIE techniques often rely on large models and multi-stage training, l...

📖 Read original article


526. How Transformers Learn to Plan via Multi-Token Prediction ​

Author: Jianhao Huang, Zhanpeng Zhou, Renqiu Xia, Baharan Mirzasoleiman, Weijie Su, Wei Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.11912v2 Announce Type: replace-cross Abstract: While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global structure in reasoning tasks. Multi-token prediction (MTP) has recently emerged as a promising alternative, y...

📖 Read original article


527. Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus ​

Author: Zijian Zhao, Jing Gao, Sen Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA

arXiv:2604.13472v2 Announce Type: replace-cross Abstract: Cooperative multi-agent reinforcement learning (MARL) is widely used to address large joint observation and action spaces by decomposing a centralized control problem into multiple interacting agents. However, such decomposition often introdu...

📖 Read original article


528. Adaptive receptive field-based spatial-frequency feature reconstruction network for fine-grained few-shot image classification ​

Author: Linyue Zhang, Wenyi Zeng, Zicheng Pan, Yongsheng Gao, Changming Sun, Jun Hu, Lixian Liu, Weichuan Zhang, Tuo Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.16936v2 Announce Type: replace-cross Abstract: Feature reconstruction techniques are widely applied for few-shot fine-grained image classification (FSFGIC). Our research indicates that one of the main challenges facing existing feature-based FSFGIC methods is how to choose the size of the...

📖 Read original article


529. Principles and Guidelines for Randomized Controlled Trials in AI Evaluation ​

Author: Christopher Kelly, Angelica Chowdhury, Alexandra Campili, Bimpe Ayoola, Devin Barbour, Thomas Chen Dawson, Ze Shen Chin, Rokas Gipi\v{s}kis
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC, cs.LG

arXiv:2605.02050v2 Announce Type: replace-cross Abstract: This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established practices from disciplines with established RCT traditions, including software engineering, economics, clin...

📖 Read original article


530. Are Flat Minima an Illusion? ​

Author: Michael Timothy Bennett
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.05209v2 Announce Type: replace-cross Abstract: Flat minima are an account of why deep networks generalise. However flatness is a matter of form (parameters), while generalisation is of function. The same function can be a result of many different parameterisations. I demonstrate this by r...

📖 Read original article


531. Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory ​

Author: Daniel Goldstein, Navneel Singhal, Eugene Cheah
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2605.09877v4 Announce Type: replace-cross Abstract: Recall presents a difficult choice: transformers have a linearly growing memory that slows each successive token, while linear RNNs typically have fixed costs but limited recall. We present Key-Value Means ("KVM"), a novel block-recurrence fo...

📖 Read original article


532. A Cascaded Edge-Cloud Architecture for Automated Diabetic Retinopathy Screening ​

Author: Nishi Doshi, Shrey Shah
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2605.14108v2 Announce Type: replace-cross Abstract: Diabetic Retinopathy (DR) is one of the leading causes of preventable blindness, and automated screening can help extend specialist capacity in resource-constrained clinical workflows. Cloud-based deep learning systems can provide strong grad...

📖 Read original article


533. ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing ​

Author: Bla\v{z} Rolih, Matic Fu\v{c}ka, Filip Wolf, Luka \v{C}ehovin Zajc
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.15375v2 Announce Type: replace-cross Abstract: Remote sensing change detection (RSCD) localises changes between two images of the same geographic region. Most state-of-the-art methods are trained with a per-pixel discriminative objective that classifies each spatial location independently...

📖 Read original article


534. Agentic Graph Retrieval-Augmented Generation for Auditable Commercial Registry Analysis ​

Author: Arthur Capozzi, Dirk Helbing
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2605.18770v2 Announce Type: replace-cross Abstract: Public commercial registries are formally open, yet their practical analysis remains difficult because relevant facts are scattered across millions of records that combine structured metadata, multilingual legal notices, temporal events, and ...

📖 Read original article


535. When Good Equations Get Bad Scores: Improving Symbolic Regression Through Better Parameter Optimization ​

Author: Boxiao Wang, Kai Li, Zhiwei Chen, Yang Huang, Runxiang Wang, Ziwen Zhang, Yifan Zhang, Jian Cheng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.23272v3 Announce Type: replace-cross Abstract: Symbolic Regression (SR) plays a central role in scientific knowledge discovery by distilling mathematical equations from observational data. Most existing SR methods function within a bi-level optimization framework: an outer loop that searc...

📖 Read original article


536. Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Task-Oriented Review with Practical Design Guidelines ​

Author: Chathura Wimalasiri, Yuchong Yao, Kishor Nandakishor, Marimuthu Palaniswami
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.23995v5 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning transferable representations from unlabeled data. However, SSL performance depends not only on model a...

📖 Read original article


537. KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing ​

Author: Yijia Fang, Yiqing Feng, Bingyu Li, Mingxun Zhou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2605.29524v2 Announce Type: replace-cross Abstract: Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoint is actually serving the advertised model. We introduce KBF, a low-cost black-box auditing...

📖 Read original article


538. Causal Density Functions ​

Author: Sridhar Mahadevan
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, cs.LG

arXiv:2606.00754v2 Announce Type: replace-cross Abstract: We study the full density ratio between a specified intervention regime $P_a$ and an observational regime $P_0$, $\rho_a=dP_a/dP_0$, under the prerequisite $P_a\ll P_0$. We call the regime-indexed ratio a causal density function when $P_a$ is...

📖 Read original article


539. Enhancing MedSAM with a Lightweight Box Predictor for Medical Image Segmentation ​

Author: Amirhossein Movahedisefat, Amirreza Fateh, Mohammad Reza Mohammadi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.04705v2 Announce Type: replace-cross Abstract: Semantic segmentation in medical imaging is a critical yet challenging task due to data scarcity and high variability across modalities. While foundation models like the Segment Anything Model (SAM) show promise, they often struggle with medi...

📖 Read original article


540. CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks ​

Author: Yijia Shao, Zora Zhiruo Wang, Neel Ahuja, Yicheng Wang, Bowen Liu, Diyi Yang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY

arXiv:2606.09833v2 Announce Type: replace-cross Abstract: AI agents are reshaping the workspace, leading to drastic change of how humans work. Despite the considerable potential of human-agent collaboration both in preserving human agency and generating economic value, this paradigm remains largely ...

📖 Read original article


541. $\tau$-Rec: A Verifiable Benchmark for Agentic Recommender Systems ​

Author: Bharath Sivaram Narasimhan, Karthik R Narasimhan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2606.10156v3 Announce Type: replace-cross Abstract: As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. Current benchmarks often rely on "LLM-as-a-judge" evaluations, which introduce subjectivity, high costs ...

📖 Read original article


Author: Yan Dai, Maryam Farboodi, Negin Golrezaei, Sepehr Shahshahani
Published: 7/28/2026, 4:00:00 AM
Categories: econ.TH, cs.AI, cs.GT, cs.LG, stat.ML

arXiv:2606.12260v2 Announce Type: replace-cross Abstract: How can we design a market of human-generated content for use in training AI models that both enables technological progress and preserves individual incentives for high-quality content creation? Existing approaches take polar positions: a "f...

📖 Read original article


543. Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents ​

Author: Zihao Wang, Yiming Li, Yutong Wu, Kangjie Chen, Zheyu Liu, Fok Kar Wai, Pin-Yu Chen, Vrizlynn L. L. Thing, Bo Li, Dacheng Tao, Tianwei Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CY, cs.HC, cs.MM

arXiv:2606.13385v2 Announce Type: replace-cross Abstract: LLM-based web agents are increasingly deployed in real-world settings such as e-commerce, where they interact extensively with untrusted web content while executing actions that carry direct financial consequences. This makes them vulnerable ...

📖 Read original article


544. Aligning Quantum Operators with Large Language Models ​

Author: Rogerio Feris, Yunchao Liu, Pengyuan Li, Hang Hua, David Kremer
Published: 7/28/2026, 4:00:00 AM
Categories: quant-ph, cs.AI

arXiv:2606.13811v2 Announce Type: replace-cross Abstract: Can Large Language Models (LLMs) understand and reason about quantum operators? Despite their remarkable capabilities in mathematics and symbolic reasoning, LLMs remain inherently blind to quantum representations such as unitary matrices. In ...

📖 Read original article


545. daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization ​

Author: Dayuan Fu, Mohan Jiang, Tongyu Wang, Dian Yang, Jiarui Hu, Liming Liu, Jinlong Hou, Pengfei Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2606.16497v3 Announce Type: replace-cross Abstract: GPU kernel optimization represents a paradigm where functional correctness is assumed and execution efficiency is the objective. We present daVinci-kernel, a reinforcement learning framework that couples skill discovery with skill exploitatio...

📖 Read original article


546. Enhancing Pathological VLMs with Cross-scale Reasoning ​

Author: Chi Phan, Tianyi Zhang, Qiaochu Xue, Yufeng Wu, Dan Hu, Zeyu Liu, Sudong Wang, Yueming Jin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.17412v4 Announce Type: replace-cross Abstract: Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnification to cellular morphology at higher magnification for accurate diagnosis. While existing pathologic...

📖 Read original article


547. SwitchBraidNet: Quantisation-Aware Lightweight Architecture for Hybrid Brain-Computer Interface ​

Author: Gourav Siddhad, Yogesh Kumar Meena
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.ET

arXiv:2606.18816v2 Announce Type: replace-cross Abstract: Hybrid brain-computer interfaces (BCIs) that integrate motor imagery (MI) and steady-state visual evoked potentials (SSVEP) provide high-dimensional neural decoding but typically exceed the computational limits of embedded hardware. To addres...

📖 Read original article


548. TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2 ​

Author: Yijin Wang, Shuyi Wang, Wenhan Zhang, Yuqi Ouyang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.19259v2 Announce Type: replace-cross Abstract: Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation models become increasingly capable of synthesizing realistic textual content and structured visual design...

📖 Read original article


549. Latent Confounded Causal Discovery via Lie Bracket Geometry ​

Author: Sridhar Mahadevan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.19610v2 Announce Type: replace-cross Abstract: We study causal discovery from observational and interventional regimes when latent variables may affect the measured system. Our first algorithm, BRIDGE (Bracket Residuals for Interventional Discovery and Geometric Estimation), combines a de...

📖 Read original article


550. Infinitesimal Causality ​

Author: Sridhar Mahadevan
Published: 7/28/2026, 4:00:00 AM
Categories: math.CT, cs.AI, math.ST, stat.TH

arXiv:2606.24621v2 Announce Type: replace-cross Abstract: Interventions can be varied continuously in many causal models. Differentiating a specified smooth intervention protocol produces vector fields on a statistical model, and their Lie brackets describe the noncommutativity of the corresponding ...

📖 Read original article


551. EmotionAI: A Privacy-Preserving Computational Intelligence Pipeline for Speech-Emotion-Grounded Conversational Analysis ​

Author: Wai Laam Mak, Isibor Kennedy Ihianle, Pedro Machado
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2606.24941v2 Announce Type: replace-cross Abstract: Reviewing recorded interviews for affective cues such as composure and agitation is slow and subjective, and cloud services that could automate the task require sensitive audio to leave the device. EmotionAI is a fully local Computational Int...

📖 Read original article


552. Neural Machine Translation for Low-Resource Tangkhul--English ​

Author: Chormi Zimik Vashai, Agniva Maiti
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.25365v2 Announce Type: replace-cross Abstract: We present a study on low-resource machine translation for the Tangkhul-English (nmf-en) language pair. Tangkhul is a severely under-resourced Tibeto-Burman language spoken primarily in Manipur, India, with virtually no prior natural language...

📖 Read original article


553. LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution ​

Author: Adib Sakhawat, Syed Rifat Raiyan, Tahsin Islam, Takia Farhin, Hasan Mahmud, Md Kamrul Hasan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL

arXiv:2606.28335v2 Announce Type: replace-cross Abstract: We argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional distribution $\mathbb{P}($position$\mid$context$)$ over a real political space. We evaluate nine current LL...

📖 Read original article


554. PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF ​

Author: Doo Hwan Hwang, Kee-Eung Kim
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.29758v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) for Large Language Models increasingly relies on critic-free methods as a practical alternative to actor--critic training. Despite their simplicity, existing critic-free approaches propagate a...

📖 Read original article


555. From Materials Database to Materials Bank: Assetizing Data for AI Driven Materials Innovation ​

Author: Chenyao Ma, Di Zhang, Weibo Gong, Wei Du, Rui Su, Yuhang Chen, Kan Xu, Huan Gu, Limin Li, Piao Ma, Zhenghao Li, Hao Li
Published: 7/28/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, physics.chem-ph

arXiv:2606.31366v3 Announce Type: replace-cross Abstract: Driven by high-throughput experimentation, computational modeling, and artificial intelligence (AI), materials data has expanded at an unprecedented rate. Conventional materials databases function only as passive repositories, archiving raw e...

📖 Read original article


556. DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning ​

Author: Hengyu Fu, Tianyu Guo, Zixuan Wang, Hanlin Zhu, Jason D. Lee, Jiantao Jiao, Stuart Russell, Song Mei
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.00341v2 Announce Type: replace-cross Abstract: Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought (CoT). However, many questions require the model to internalize the multi-step reasoning within a sing...

📖 Read original article


557. From World Models to World Action Models: A Concise Tutorial for Robotics ​

Author: Xiaoxiong Zhang, Xiong Zeng, Wei Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.SY, eess.SY

arXiv:2607.00836v4 Announce Type: replace-cross Abstract: Rather than providing an exhaustive survey, this paper presents a concise tutorial on world models and world action models for robotics. After reading the tutorial, readers should have a clear understanding of what constitutes a "world", how ...

📖 Read original article


558. The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits ​

Author: Gemma Galdon Clavell, Pablo Accuosto, Usman Gohar
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.02201v3 Announce Type: replace-cross Abstract: The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains fragmented across competing risk taxonomies that catalog risks without showing how an audit is executed...

📖 Read original article


559. DOSE-I: A Multimodal Biosignal Dataset of Procedural Sedation for Endoscopy -- Technical Report ​

Author: Jakob Garbe, Jan W. Kantelhardt, Katja Seeliger, Thomas Schmid
Published: 7/28/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2607.02570v2 Announce Type: replace-cross Abstract: In this document, we describe characteristics and technical details of the multimodal biosignal dataset DOSE-I of procedural sedation for endoscopy published on zenodo. The DOSE-I dataset includes 78.5 hours of recording in 171 records rangin...

📖 Read original article


560. QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting ​

Author: Shah Nawaz Haider, Steve Austin, Arnab Barua, Sarowar Morshed Shawon, Hadaate Ullah
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.02632v2 Announce Type: replace-cross Abstract: Time-series forecasting supports decisions in finance, en-ergy, transportation, public health, and industrial monitoring. Recent foundation models improve transfer across forecast-ing tasks, but many depend on centralized data and Trans-forme...

📖 Read original article


561. Multi-Turn On-Policy Distillation with Prefix Replay ​

Author: Baohao Liao, Hanze Dong, Christof Monz, Xinxing Xu, Li Dong, Furu Wei
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, stat.ML

arXiv:2607.04763v3 Announce Type: replace-cross Abstract: We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because eac...

📖 Read original article


562. Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition ​

Author: Andrei Florian, Cynthia Jayne Amol, Hope Kerubo Ombaba, Xiaoyu Cui, Boniface Mwau, Biatus Maina Kamau, Lilian Diana Awuor Wanzare, Christiane Fellbaum, Happy Buzaaba
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.04814v2 Announce Type: replace-cross Abstract: Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale. A promising direction is to leverage the linguistic relatedness between a low-resource targe...

📖 Read original article


563. Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation ​

Author: Haozhe Wang, Weijia Feng, Jinpeng Yu, Che Liu, Ping Nie, Fangzhen Lin, Jiaming Liu, Ruihua Huang, Jimmy Lin, Wenhu Chen, Cong Wei
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.05382v4 Announce Type: replace-cross Abstract: Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutoff events, and more. This world-knowledge b...

📖 Read original article


564. GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model ​

Author: Peter Bohm, Saimunur Rahman, Abdelwahed Khamis, Sagun Man Singh Shrestha, Chris McCool, Peyman Moghadam
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.06882v2 Announce Type: replace-cross Abstract: Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets. We ask whether this recipe i...

📖 Read original article


565. IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation ​

Author: Yiting Wang, Jingyi Zhang, Wenhu Zhang, Ke Chao, Yves Liang, Kun Cheng, Kang Zhao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.09133v2 Announce Type: replace-cross Abstract: While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step iterative solvers incurs severe inference latency. Few-step distillation targeting the Classifier-Free Gui...

📖 Read original article


566. Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control ​

Author: Boyu Li, Linjie Qiu, Lin-Ping Yuan, Duotun Wang, Yue Jiang, Zeyu Wang, Hongbo Fu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.GR

arXiv:2607.10405v2 Announce Type: replace-cross Abstract: Controlling attributes is a critical step toward achieving the final creative outcome, yet current approaches fall short in supporting users in the iterative refinement of generative content. We propose Spatula, a proof-of-concept system that...

📖 Read original article


567. How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study ​

Author: Yunbo Lyu, David Williams, Jieke Shi, Zhensu Sun, Chao Peng, Zhou Yang, Federica Sarro, David Lo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.HC

arXiv:2607.10856v2 Announce Type: replace-cross Abstract: The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks with limited human intervention, has been marked by rapid advances and adoption, but little is known abou...

📖 Read original article


568. The One-Word Census: Answer-Choice Conformity Across 44 Language Models ​

Author: Tapan Parikh
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2607.12796v2 Announce Type: replace-cross Abstract: When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same answer every other model picks? Asked to "pick a word -- any word," 44 models chose "serendipity" 41% o...

📖 Read original article


569. The Hitchhiker's Guide to Monoculture ​

Author: Gordon Burtch
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.SE

arXiv:2607.13077v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the software artifacts that developers create. Whether this occurs in practice is unclear because developers...

📖 Read original article


570. Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models ​

Author: Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim, Suyoun Kim, Bo-Ru Lu, Qingming Tang, Ankur Gandhe, Hung-yi Lee, Chieh-Chi Kao, Chao Wang
Published: 7/28/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL, cs.LG, cs.SD

arXiv:2607.13408v2 Announce Type: replace-cross Abstract: Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global simila...

📖 Read original article


571. Anatomically Faithful but Temporally Diffuse: Auditing Attribution for Left-Ventricular Ejection-Fraction Estimation from Echocardiography ​

Author: Hyunkyung Han, Min Jung Kim
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.13738v4 Announce Type: replace-cross Abstract: Deep video models estimate left-ventricular ejection fraction (EF) from echocardiography with near-expert accuracy, and post-hoc attribution is increasingly used to certify that such models look at the right place. Because EF is defined by th...

📖 Read original article


572. Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents ​

Author: Paul Kassianik, Blaine Nelson, Yaron Singer
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.15263v3 Announce Type: replace-cross Abstract: Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but incomp...

📖 Read original article


573. Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection ​

Author: Haochen Zhao, Yongxiu Xu, Xinkui Lin, Dong Xie, Jiarui Lu, Yuqi Qian, Yubin Wang, Hongbo Xu, Gaopeng Gou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.18080v2 Announce Type: replace-cross Abstract: Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world misinformation often exh...

📖 Read original article


574. SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions ​

Author: Hoang-Thang Ta
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.18290v2 Announce Type: replace-cross Abstract: In recent years, Kolmogorov-Arnold Networks (KANs) have attracted increasing attention due to their effectiveness in machine learning and scientific computing, offering a new paradigm for neural network design. In this paper, we present SechK...

📖 Read original article


575. Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary ​

Author: Jan Kirin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.18553v3 Announce Type: replace-cross Abstract: Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes? We test both questions in a frozen 2.6B looped transformer, Ouro-RLTT. On GSM8K, a strict pre-answer pr...

📖 Read original article


576. Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts ​

Author: Haodi Fan, Zucong Lan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.18970v2 Announce Type: replace-cross Abstract: Agent Skills have become persistent behavioral artifacts across independent AI agent systems. They combine natural-language task specifications with metadata and optional references, scripts, assets, hooks, package manifests, tests, and compa...

📖 Read original article


577. MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents ​

Author: Guofeng Zhang, Yizeng Quan, Huaiyi Fang, Jianwei Lv, Jinyao Liu, Xunxu Duan, Lening An, Yu Ouyang, Junfeng Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.18999v2 Announce Type: replace-cross Abstract: Evaluating multi-turn medical consultation agents requires judging the diagnostic support provided by the histories they elicit through interaction. Yet coupled evaluation lets each policy both elicit the history and generate the terminal dia...

📖 Read original article


578. SLPO: Scaling Latent Reasoning via a Surrogate Policy ​

Author: Runyang You, Zhiyuan Liu, Yongqi Li, Wenjie Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.19691v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must ...

📖 Read original article


579. G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection ​

Author: Yechan Kim, JongHyun Park, Dongho Yoon, Namhoon Jung, Moongu Jeon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.19942v2 Announce Type: replace-cross Abstract: This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited view...

📖 Read original article


580. Generative AI floods and dilutes the market for books ​

Author: Tuhin Chakrabarty, Xinyue Liu, Jane C. Ginsburg, Paramveer Dhillon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2607.20349v2 Announce Type: replace-cross Abstract: Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-quality ``slop'' that buyers will ignore, and are assumed to carry little commercial weight. We test that assumption with full-te...

📖 Read original article


581. PhantomFill: When the Form Demands an Answer, Language Models Invent One ​

Author: Rana Muhammad Usman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.20492v2 Announce Type: replace-cross Abstract: Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself causes hallucination. We ask thirteen models the same question about the same input and cha...

📖 Read original article


582. Adaptive Multi-Horizon Reinforcement Learning ​

Author: Manoosh Samiei, Doina Precup, Paul Masset
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.20656v2 Announce Type: replace-cross Abstract: Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL), this trade-off is typically controlled through a fixed discount factor, which imposes a s...

📖 Read original article


583. Emergent Compositional Skills in Mixture-of-Experts VLAs ​

Author: Shlok Shah, Rhiaan Jhaveri, Tharun Kumar Tiruppali Kalidoss, Chirayu Nimonkar, Ishaan Javali, Dhruv Shah
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2607.20771v2 Announce Type: replace-cross Abstract: We consider the problem of learning compositional robot policies end-to-end from expert demonstrations, without any pre-specified notion of task decomposition or hierarchy. We ask whether a VLA trained with a simplified Mixture-of-Experts (Mo...

📖 Read original article


584. Robostral Navigate ​

Author: Abhijeet Somani, Aditi Kabra, Adrian Valente, Alan Jeffares, Albert Jiang, Aleksandr Timashov, Alexandre Sablayrolles, Amelie Heliou, Andre Jonasson, Andrew Bai, Andrew Ehrenberg, Andrew Zhao, Angele Lenglemetz, Anmol Agarwal, Antonia Calvi, Arata Suzuki, Aylin Guliz Akkus, Aysenur Karaduman, Baptiste Bout, Baptiste Roziere, Baudouin De Monicault, Benjamin Holzschuh, Benjamin Lefaudeux, Bernhard Stadlbauer, Blazej Osinski, Camille Le Scao, Chaoran Yu, Chen-Yo Sun, Christian Wallenwein, Christophe Renaudin, Clemence Lanfranchi, Corentin Barreau, Corentin Sautier, Cristiana-Diana Diaconu, Cyprien Courtot, Daniel Marczak, Darius Dabert, Diego de Las Casas, Dominik Nuss, Dylan Rubini, Dzmitry Soupel, Emilien Fugier, Erik Aas, Etienne Millon, Eujeong Choi, Fabian Paischer, Fabian Schlager, Faruk Ahmed, Federico Baldassarre, Filip Szatkowski, Gabrielle Berrada, Gaetan Ecrepont, Gaetan Lepage, Gaspard Blanchet, Gaspard Donada-Vidal, Gauthier Delerce, Gauthier Guinet, Genevieve Hayes, Georgii Novikov, Giada Pistilli, Gianluca Galletti, Guillaume Breton, Guillaume Martin, Gunjan Dhanuka, Gunshi Gupta, Han Zhou, Hasan Furkan Vural, Indraneel Mukherjee, Ivan Cuevas Salazar, Jan Ludziejewski, Jason Rute, Jean-Hadrien Chabran, Jean-Malo Delignon, Jie Zhang, Joachim Studnia, Joep Barmentlo, Johannes Brandstetter, John Harvill, Jonas Amar, Jonas Schweizer, Josselin Somerville, Julien Denize, Julien Tauran, Kartik Khandelwal, Kush Jain, Larissa Laich, Laura Calem, Laurence Aitchison, Laurent Callot, Leo Cotteleer, Leonard Blier, Lingxiao Zhao, Louis Martin, Louis Serrano, Lucile Saulnier, Luis Montero, Maarten Buyl, Marcin Mozejko, Margaret Jennings, Mathieu Schmitt, Mathilde Guillaumin, Matthieu Dinot, Matthieu Futeral, Mauro Comi, Max Mynter, Maxim Berman, Maxime Darrin, Maxime Louis, Maximilian Augustin, Maximilian Muller, Mert Unsal, Mia Chiquier, Michael Pilcer, Michal Pietruszka, Michal Zajac, Mikhail Biriuchinskii, Minwoo Kang, Morgane Riviere, Namit Katariya, Nathan Grinsztajn, Neeraj Aggarwal, Neha Gupta, Ola Mysiak, Oliver Leicht, Olivier Bousquet, Parag Jain, Patricia Wang, Patrick von Platen, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Pavan Kumar Reddy, Philomene Chagniot, Pierre Stock, Pierre-Andre Savalle, Piotr Milos, Prateek Gupta, Pravesh Agrawal, Quentin Desreumaux, Quentin Torroba, Quercus Hernandez, Ram Ramrakhya, Randall Isenhour, Ranjit Parva, Raul Perez Pelaez, Remi Delacourt, Rishi Shah, Rohin Arora, Romain Sauvestre, Roman Soletskyi, Sagar Vaze, Samuel Humeau, Sanchit Gandhi, Sandeep Subramanian, Sarthak Mittal, Saskia Adaime, Sebastian Kaltenbach, Shashwat Dalal, Sherif Waly, Shrimai Prabhumoye, Siddharth Gandhi, Simon Sorg, Soham Ghosh, Sophie Marbach, Stanislas Lange, Sumukh Aithal, Szymon Antoniak, Teven Le Scao, Thibaut Lavril, Thomas Coste, Thomas Foubert, Thomas Robert, Thomas Wang, Tianyu Zhang, Tim Lawson, Timothee Lacroix, Tobias Kronlachner, Tom Bewley, Tomas Hodan, Tuhin Das, Tyler Wang, Van Phung, Vedant Nanda, Victor Jouault, Victor Letzelter, Victor Paltz, Victor Poucheret, Vincent Pfister, Virgile Richard, Vladislav Bataev, Wassim Bouaziz, Wen Ding Li, William Marshall, Xinghui Li, Xingran Guo, Xinyu Yang, Yann Dreze, Yihan Wang, Zaccharie Ramzi, Zhenlin Xu, Zsofia Csakany, Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.20785v2 Announce Type: replace-cross Abstract: Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, ...

📖 Read original article


585. Error Certificates for KV-Cache Eviction via Randomized Design ​

Author: Peng Xie
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.21475v2 Announce Type: replace-cross Abstract: Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest. We prove that this design cannot know what it destroyed: evicted values can be altered so that everything the serving system retains is u...

📖 Read original article


586. Neptuna: A Comprehensive Machine Learning Framework for Benchmarking Complex Multiphase Flows ​

Author: Harish Ramachandran, Bj"orn Kimpel, Thomas Paula, Josef Winter, Steffen Schmidt, Nikolaus Adams
Published: 7/28/2026, 4:00:00 AM
Categories: physics.flu-dyn, cs.AI

arXiv:2607.22280v2 Announce Type: replace-cross Abstract: Compressible multiphase flows involving shocks and material interfaces arise in applications such as bubble collapse and droplet breakup, where strong nonlinear interactions produce complex interface deformation, mixing, and multiscale dynami...

📖 Read original article


587. Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science ​

Author: Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL

arXiv:2607.22513v2 Announce Type: replace-cross Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethn...

📖 Read original article