Skip to content

arXiv cs.AI - 2026-08-03 ​

261 items collected.


1. OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems ​

Author: Konstantinos I. Roumeliotis, Ranjan Sapkota
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28629v1 Announce Type: new Abstract: The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the architectural understanding of Agentic AI, particularly in separating inference, orchestration, and execution layers ...

📖 Read original article


2. Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review ​

Author: Vaibhava Lakshmi Ravideshik, Mayank Kejriwal
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI-generated papers remains an open challenge. We propose and implement a rigorous...

📖 Read original article


3. LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann Hypothesis ​

Author: Alizer Wong, Zixin Zeng, Yi Tan, Wenyuan Li, Xuhang Chen, Xingru Lai, Yang Shi, Liangsi Lu, Yanhui Chen
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28632v1 Announce Type: new Abstract: Major mathematical conjectures still depend heavily on expert intuition, so a unified method for the systematic generation and validation of conjectures with substantial mathematical potential remains unavailable. We present a three stage pipeline for ...

📖 Read original article


4. ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning ​

Author: Fei Ding, Yongkang Zhang, Runhao Liu, Yuhao Liao, Zijian Zeng
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.28642v1 Announce Type: new Abstract: Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. We argue that under bounded context windows, the core bottleneck is not trajectory compress...

📖 Read original article


5. TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter ​

Author: Oliver Savolainen, Emanuele Bastianelli, Hosein Azarbonyad
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28657v1 Announce Type: new Abstract: Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert users. This work addresses the challenge by introducing a Task-Aware Prompt Rewriter (TAPR), a model that reformu...

📖 Read original article


6. Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding ​

Author: Yuxuan Hu, Yuhao Wang, Tianbo Huang, Chao Zhang, Ziwei Liu, Lihua Zhang, Xiangyu Zhao
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28659v1 Announce Type: new Abstract: Cross-domain sequential recommendation (CDSR) aims to model users' dynamic interest transitions and sequential patterns across multiple domains. Recently, generative recommendation (GR) has emerged. It first learns semantic identifiers (SIDs) from item...

📖 Read original article


7. An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents ​

Author: Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28662v1 Announce Type: new Abstract: Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfaces under several name variants, relationships duplicate, and distinct i...

📖 Read original article


8. How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories ​

Author: Hui Wei, Junda Wu, Sheldon Yu, Sizhe Zhou, Yizhu Jiao, Ming Zhong, Bowen Jin, Tong Yu, Shijia Pan, Jiawei Han, Julian McAuley
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.28674v1 Announce Type: new Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods rely on output-level signals or collapse processing depth into a single trajectory...

📖 Read original article


9. Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support ​

Author: Shayndhan Sivanathan, Shravan Nageswaran, Mehdi Zadem, Ryaan Sultan, Nicolas von Mallinckrodt, Max Solovyev, Alexey Matyushkin, Sumon Sadhu, Gabriele C DeLuca, Sanjeeva Jeyaretna, James Hillis, Manoj Ramachandran, Prakash Jayakumar
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidan...

📖 Read original article


10. ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding ​

Author: Xinkui Zhao, Enbo Chen, Yifan Zhang, Chang Liu, Guanjie Cheng, Naibo Wang, Yueshen Xu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2607.28678v1 Announce Type: new Abstract: Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However, existing agentic memory approaches often discard fine-grained dentity c...

📖 Read original article


11. Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO ​

Author: Sheryl Paul, Vidisha Kudalkar, Anand Balakrishnan, Lars Lindemann, Alberto Speranzon, Jyotirmoy V. Deshmukh
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2607.28679v1 Announce Type: new Abstract: Multi-agent planning problems arise in a variety of engineering applications, such as multi-robot wildfire fighting and unmanned aerial inspection in factories. A particular challenge is the existence of spatio-temporal (i.e., when and/or where an agen...

📖 Read original article


12. Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery ​

Author: Zhan'ao Yao, Liang Yin, Zhihao Gao, Boxuan Zhang, Xiaoyu Wu, Linjing Li, Rongyan Wang, Tingwei Chen, Youwei Wang, Xiaolin Zhao, Jiahui Shi, Jianjun Liu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28684v1 Announce Type: new Abstract: Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult to determine whether a model is discovering laws from data or merely recalling answers from its train...

📖 Read original article


13. Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks ​

Author: Youting Wang, Xiao Han, Dingyan Shang, Yuan Tang, Bowen Liu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official im...

📖 Read original article


14. SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition ​

Author: Yuqi Tang, Chenyi Zhou, Libin Wang, Keyan Ding, Qiang Zhang, Huajun Chen
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.28692v1 Announce Type: new Abstract: Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance on predefined tool spaces with static semantics limits their applicability to o...

📖 Read original article


15. EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses ​

Author: Jiahui Li, Ruili Fang, Zishuai Liu, Yutong Guo, Nan Yang, Wenzhan Song, Jin Lu, Fei Dou
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28788v1 Announce Type: new Abstract: Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks are poorly suited to this setting: they restrict prediction to closed code sets, exclude free-text notes, and supe...

📖 Read original article


16. Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures ​

Author: Harsh Raj, Vipul Gupta, Anas Mahmoud, Razvan-Gabriel Dumitru, Darvin Yi, Aakash Sabharwal, Yunzhong He
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28802v1 Announce Type: new Abstract: Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment problem: the same visible failure may call for model...

📖 Read original article


17. Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions ​

Author: Pranav Narayanan Venkit, Akshara Prabhakar, Yu Li, Daniel Lee, Chien-Sheng Wu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2607.28818v1 Announce Type: new Abstract: As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse',...

📖 Read original article


18. Fragility of Value under Imperfect Alignment ​

Author: Winter Cross
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28881v1 Announce Type: new Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect...

📖 Read original article


19. Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design ​

Author: Manisha Dubey, Rimvydas Rubavicius, N. Siddharth, Subramanian Ramamoorthy
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, stat.ML

arXiv:2607.28894v1 Announce Type: new Abstract: Computational cognitive modeling seeks to infer latent cognitive mechanisms underlying observed behavior. Bayesian inverse planning provides a principled framework for such inference, but its success depends critically on the experimental environment. ...

📖 Read original article


20. NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability ​

Author: Duo Xu, Faramarz Fekri
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28942v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scientific discovery. In these settings, agents must act based on limited observatio...

📖 Read original article


21. MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations ​

Author: Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful b...

📖 Read original article


22. Scaling Scientific Discovery Environments for Turn-Level Agentic RL ​

Author: Yucheng Xu, Keyi Zhang, Yuyang Yu, Min Zhang, Shiyuan Meng, Pei Chu, Zhongying Tu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28990v1 Announce Type: new Abstract: Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remains constrained by th...

📖 Read original article


23. MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents ​

Author: Zeying Hao, Hao Guo, Mengtao Xu, Yimin Hu, Yuheng Song, Zesheng Zhou, Jinsong Lan, Xiaoyong Zhu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29002v1 Announce Type: new Abstract: Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone. However, existing benchmarks largely rely on text-only or synthetic re...

📖 Read original article


24. Evidence-Grounded Constraint Checking in Construction Documents ​

Author: Rashid Mushkani, Hugo Berard, Shin Koseki
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29058v1 Announce Type: new Abstract: Professional-document review is a constraint-checking problem in which decisions depend on relations among text, geometry, pages, and document revisions. We present an evidence-grounded pipeline that normalizes extracted facts, executes four-state rule...

📖 Read original article


25. On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness ​

Author: Matthew Nguyen, Kyle Cox, Austin Meek, Iv'an Arcuschin
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29062v1 Announce Type: new Abstract: Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is possible to monitor it. However, in some cases, models do not verbalize i...

📖 Read original article


26. A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation ​

Author: Keita Kinjo
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, stat.ML

arXiv:2607.29077v1 Announce Type: new Abstract: Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input required to obtain a desired output. Although CEs are conventionally formulated as a distance-minimization problem,...

📖 Read original article


27. Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration ​

Author: Yanbin Fang, Xuan Wei, Wei Chen
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29087v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in enterprise settings, yet individual models remain bounded by model-specific capability limitations. These heterogeneous boundaries pose a deployment challenge, but also create an opportunity: st...

📖 Read original article


28. CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents ​

Author: Blaise Delattre, Cong Wang, Yang Cao
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29190v1 Announce Type: new Abstract: Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates generally authorize the observed return and action, leaving the decision unprotected against small errors...

📖 Read original article


29. MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft ​

Author: Jianxin Gao, Beini Hu, Runze Li, Wanli Peng, Ruohan Lei, Jinyuan Zhang, Linna Deng, Tianyi Yu, Zining Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29218v1 Announce Type: new Abstract: With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unfortunately, most existing benchmarks evaluate them under fixed game mechanics. High performance in these settings...

📖 Read original article


30. Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL ​

Author: Ruiming Liang, Yi Zhong, Yizhen Yuan, Yinan Zheng, Tianyi Tan, Tianyue Wang, Haiyun Guo, Jinqiao Wang, Xianyuan Zhan
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29246v1 Announce Type: new Abstract: Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinforcement learning (RL) has become an increasingly important problem for L...

📖 Read original article


31. Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents ​

Author: Minghui Pan, Jiayuxuan Yang, Yuanyuan Yuan, Yu Jiang, Zhenpeng Chen
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29254v1 Announce Type: new Abstract: AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions. Yet LLMs often become substantially less safe when deployed as agents, and the ...

📖 Read original article


32. MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation ​

Author: Hang Yan, Zhangxuan GU, Beitong Zhou, Jiaxuan Chen, Runze Li, Yusong Hu, Shuheng Shen, Changhua Meng
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29320v1 Announce Type: new Abstract: Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments. However, existing agents are typically domain-specific, limiting the deployment and user experience. This motiv...

📖 Read original article


33. Beyond Component Testing: Validating Agentic AI Systems ​

Author: Fabio Orazio Mirto, Luca D'Agati, Giuseppe Tricomi, Stefano Silvestri, Francesco Longo, Antonio Puliafito, Giovanni Merlino
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SE

arXiv:2607.29405v1 Announce Type: new Abstract: Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches validation practice beyond component testing and one-shot input--output evaluation, because acceptable ...

📖 Read original article


34. ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models ​

Author: Penglin Zhu, Jungang Xu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29431v1 Announce Type: new Abstract: Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that ...

📖 Read original article


35. Beyond Retrieval: Analytic Memory for Multimodal Agents ​

Author: Zhoujin Tian, Yao Tian, Hao Zhang, Cheng Chen, Yakun Li, Lei Zhang, Xiaofang Zhou
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29440v1 Announce Type: new Abstract: Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions. Existing systems largely emphasize \emph{retrieval memory}, organizing interaction histories through...

📖 Read original article


36. Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember ​

Author: Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Haoyu Wu, Minghui Wu, Chenxu Zhao, Ante Wang, Guannan He, Changwei Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29468v1 Announce Type: new Abstract: Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gradients yet do not explicitly shape future practice. External skill memories preserve procedural expe...

📖 Read original article


37. AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction ​

Author: Rui Zou, Yutao Zhu, Mengqi Wei, Ji-Rong Wen
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29549v1 Announce Type: new Abstract: Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challenging. Existing representative methods mainly revise outputs through natural-language reflection or a...

📖 Read original article


38. COntExt: Towards Context-Aware Ontology Extension from Operational Metrics ​

Author: Hussain Hussain, Stefan Sch"oberl, Angelika Schneider, Verena Geist
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29553v1 Announce Type: new Abstract: Organizations increasingly define operational metrics in structured, machine-readable formats to monitor systems, processes, and compliance. These metric definitions implicitly encode domain knowledge, such as referencing concepts, properties, and rela...

📖 Read original article


39. LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback ​

Author: Manith Adikari, Bei Peng, Samuele Vinanzi, Angelo Cangelosi
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2607.29559v1 Announce Type: new Abstract: Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decision-making tasks often involve multiple, competing objectives, such as performance versus efficiency, where ground...

📖 Read original article


40. DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat ​

Author: Ismayil Ismayilov, Atakan Kara, Kaan Oktay
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.29577v1 Announce Type: new Abstract: Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interaction...

📖 Read original article


41. AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers ​

Author: Tianyu Huai, Tingshuo Fan, Xinchi Chen, Yining Zheng, Yuxin Wang, Shuang Chen, Jie Zhou, Xuanjing Huang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper replication, or final...

📖 Read original article


42. Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics ​

Author: Yimin Chen, Brian Fricke, Bo Shen, Jamie Lian, Mingkan Zhang, James Lo, Yun Zhang, Shi Ye, Jiajing Huang, Han Hu, Chujie Lu, Rui Tang, George Zhuang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29657v1 Announce Type: new Abstract: Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency, and maintenance effectiveness. However, effective deployment of FDD solutions in buildings requires structured domain knowledge that c...

📖 Read original article


43. ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction ​

Author: Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We prese...

📖 Read original article


44. Scaffolding Critical Engagement with GenAI: Transforming Ethnic Minority Preparatory Students' Collaborative Discourse in Prompt Engineering Tasks ​

Author: Deliang Wang, Cunling Bian
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.28630v1 Announce Type: cross Abstract: Generative AI (GenAI) holds significant promise for advancing educational equity among ethnic minority students by broadening access to learning resources and mitigating linguistic barriers. However, these benefits are counterbalanced by the risk of ...

📖 Read original article


45. Topology-Aware Data Movement for Disaggregated GPU Inference ​

Author: Sanjeev Rao Ganjihal
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.PF

arXiv:2607.28633v1 Announce Type: cross Abstract: Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run on separate GPU pools, the KV cache must be transferred between them. For a 70B model this is 2.6 GB per request...

📖 Read original article


46. The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models ​

Author: Plawan Kumar Rath
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2607.28639v1 Announce Type: cross Abstract: We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias. On unambiguous tasks (BBQ-disambig), response-based distillation from a Gemma-2-9B teacher improves context-following: for the most biased ...

📖 Read original article


47. The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load? ​

Author: Dahlia Shehata, Ming Li
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.28641v1 Announce Type: cross Abstract: We introduce the \textit{Agentic Formalism Trap} and the Evaluative Dissonance Index ($D_E$), quantifying how LLM-as-a-Judge systems conflate structural proceduralism with semantic truth under adversarial load. Analyzing 22,500 trajectories across 3 ...

📖 Read original article


48. Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework ​

Author: Prerna Luthra
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.28644v1 Announce Type: cross Abstract: Creativity in computational systems is often evaluated as an objective property of artifacts, with existing Computational Creativity (CC) frameworks assessing creative merit at the level of outputs or systems rather than interpretive context. However...

📖 Read original article


49. Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation ​

Author: Fan Wu, Cuiyun Gao, Yiming Huang, Yang Xiao, Yujia Chen, Qing Liao
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.28645v1 Announce Type: cross Abstract: Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildable codebase with shared components and working navigation. This project-level s...

📖 Read original article


50. ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning ​

Author: Thang Doan Viet, Anh Nguyen Hoang, Tinh Luong Son, Anh Hoang Thi Ngoc, Huyen Giang Thi Thu, Tai Le Quy
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.28647v1 Announce Type: cross Abstract: This paper presents ConnectED, a human-centered AI system that supports the full instructional lifecycle in Vietnamese education by linking curriculum-aligned lesson design, interactive student learning, and feedback-driven refinement. Built on VietE...

📖 Read original article


51. Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations ​

Author: Hainiu Xu, Zhaoyue Sun, Hanqi Yan, Jinhua Du, Caroline Catmur, Yulan He
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL

arXiv:2607.28648v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on modifying cognitive appraisals, the subjective interpretation of events that elicit negative emotions, which is ty...

📖 Read original article


52. COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention ​

Author: Zonghuan Li, Litian Li, Arthur Mercier, Gara Dorta, Balint Dioszegi, Jose Morales-Vargas, Chenxu Hao, Ivan Kondyurin, Vanessa Begemann, Nale Lehmann-Willenbrock, Bernd Dudzik, Saunaq Chakrabarty, Sotiris Vacanas, Laura Cabrera-Quir'os, Anne L. J. ter Wal, Vitaliy Popov, Jorge Castro-God'inez, Chirag Raman, Stephanie Tan, Hayley Hung
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.28649v1 Announce Type: cross Abstract: COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an international conference. It captures ecologically valid social interactions in a weakly scripted setting consisting of two...

📖 Read original article


53. Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration ​

Author: Rana Abou Khamis, Hala Assal, Ashraf Matrawy
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.28650v1 Announce Type: cross Abstract: While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration of GenAI into system administration may involve deeper shifts in professional practice that are not yet fully ...

📖 Read original article


54. HenTwin: A Multimodal Digital Twin Framework for Longitudinal Biological State Monitoring in Laying Hens ​

Author: Yashan Dhaliwal, Shreya Rao, Suresh Neethirajan
Published: 8/3/2026, 4:00:00 AM
Categories: q-bio.OT, cs.AI, stat.AP

arXiv:2607.28652v1 Announce Type: cross Abstract: Early-life monitoring in laying hens remains constrained by fragmented single-modality sensing and the absence of formal system-level state representations. HenTwin, a multimodal digital twin framework implemented as a five-layer IoT architecture, fo...

📖 Read original article


55. Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation ​

Author: Claudia Grosser, Maike Heuer, Denis Krompass, Thomas A. Runkler
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.28658v1 Announce Type: cross Abstract: Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, evaluating federated pre-training remains challenging because differences in client participation and...

📖 Read original article


56. Sensitivity Analysis of GRU, LSTM and Transformer Encoder in Classification of Automated Driving Systems ​

Author: Bidhya Shrestha, Christos Papadopoulos
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.28665v1 Announce Type: cross Abstract: Automated driving systems (ADSs) are becoming ubiquitous. Future Software Defined Vehicles (SDVs) may be able to run multiple ADSs, both native and aftermarket such as Comma.ai's Openpilot. Monitoring systems to independently verify which automated d...

📖 Read original article


57. Guarantees on Dynamical System Distinguishability for LLM Token Generation ​

Author: Mohamed Akrout, Dan Wilson
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.DS

arXiv:2607.28667v1 Announce Type: cross Abstract: Recent work has shown that classifying large language models (LLMs)' responses can be distinguished by modeling token embeddings as trajectories of a black-box dynamical system (DS) and comparing prediction residuals of two DSs. Despite the empirical...

📖 Read original article


58. LAWFUL: Law-Aligned Witness for Faithful Use of Latents ​

Author: Kevin Chen, Kenneth W. Parker, Anish Arora
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.28672v1 Announce Type: cross Abstract: When a neural network predicts a physical system accurately, has it learned the governing law as formal, structured knowledge, and if so, does the network's internal computation actually use that representation throughout the law's domain of validity...

📖 Read original article


59. MPP-GNN: Subject-Adaptive Community Detection for fMRI-Based Alzheimer's Disease Classification ​

Author: Yang Zhang, Xiao Zhou, Jonathan Warrell, Avram Holmes, Xuan Zhang, Mark Gerstein
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.NC

arXiv:2607.28681v1 Announce Type: cross Abstract: Functional magnetic resonance imaging (fMRI) is a widely used technique for studying the brain. Recent methods that utilize graph neural networks (GNNs) for analysis of brain functional connectivity have shown great potential for the classification o...

📖 Read original article


60. Metaphor-Induced Algorithmic Steering: Cross-Domain Procedural Transfer in LLM Code Generation ​

Author: Zhibo Hu, Chen Wang, Yanfeng Shu, Hye-young Paik, Liming Dong, Liming Zhu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.28683v1 Announce Type: cross Abstract: Large language models benefit from elements in natural language, such as metaphors and analogies in training data and inference input to achieve generalisability across different domains. However, these language elements may also lead to unwanted beh...

📖 Read original article


Author: Mohammad Asif, Azizuddin Khan, Mohd Azam, Anurag Rajkumar Bombarde
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.ET

arXiv:2607.28687v1 Announce Type: cross Abstract: As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades, yet routine assessment often misses its earliest signs. This article critically synthesizes recent technologi...

📖 Read original article


62. Reflected UAS: Corrected Deterministic Stability and Direct CTMC Drift Calculation ​

Author: Krishna Subedi
Published: 8/3/2026, 4:00:00 AM
Categories: cs.PF, cs.AI, math.PR

arXiv:2607.28688v1 Announce Type: cross Abstract: We analyze Reflected UAS routing for heterogeneous multi-server queues at fixed parameters under subcritical load. The deterministic surrogate is a reflected ODE on the nonnegative orthant, not the unconstrained drift equation. This reflected ODE has...

📖 Read original article


63. Code Is the Body: Agent-Owned Software Bodies for Recursive Evolution and Descent ​

Author: Roy Zhao (Paul G. Allen School of Computer Science & Engineering, University of Washington), Zhenyu Zhao (Independent Researcher)
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.28691v1 Announce Type: cross Abstract: Personalized AI agents are often configurable without giving users control over the artifacts that determine their future behavior. We present OurArk, an architecture for persistent personal agents centered on an agent-owned software body: an identit...

📖 Read original article


64. SEDR-Seq2P: A Lightweight Dilated Residual Sequence-to-Point Network for Multi-Task Industrial NILM ​

Author: Hatem Haddad, Feres Jerbi, Issam Smaali
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.28693v1 Announce Type: cross Abstract: Industrial NILM remains challenging because measurement noise and widespread concurrent machine operation reduce the generalization of models tuned on residential data. This work adopts a one-to-many, multi-task disaggregation setting, in which a sin...

📖 Read original article


65. Predicting Steel Fatigue Life from Micrographs Using Physics-Informed Deep Learning ​

Author: Aryuemaan Kumar Chowdhury
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2607.28695v1 Announce Type: cross Abstract: Here is the plain text version optimized for arXiv's submission form. Custom macros (like \CV and \SI) have been converted to standard text/math so they render correctly on the webpage: Evaluating the fatigue life of structural steels conventionally ...

📖 Read original article


66. WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization ​

Author: Fanzhe Wei, Li Liu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AR, cs.AI

arXiv:2607.28699v1 Announce Type: cross Abstract: KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the request it is serving right now. We give it a provably sound runtime meter, a "DTrace for KV quantization": a pe...

📖 Read original article


67. A user's guide to PINNs in geometric analysis: lessons from the asymptotic Plateau problem ​

Author: Tancredi Schettini Gherardini
Published: 8/3/2026, 4:00:00 AM
Categories: math.DG, cs.AI, hep-th, math.AP

arXiv:2607.28733v1 Announce Type: cross Abstract: This proceedings contribution elaborates on the findings of arXiv:2605.26234v2: a joint work with Marco Usula, where we introduced a machine learning framework based on physics-informed neural networks (PINNs), aimed at constructing near-minimal disc...

📖 Read original article


68. DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing ​

Author: Sowjanya Puligadda, Mengdie Zhang, Ali Zamani, Dhruva Dixith Kurra, Eric Chen, Juan Marcano
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.28750v1 Announce Type: cross Abstract: As mobile applications grow in complexity, traditional End-to-End (E2E) testing frameworks struggle with UI volatility, maintenance overhead, and cross-platform scalability. This paper presents DragonCrawl, an AI-driven mobile testing system for cont...

📖 Read original article


69. SCMA: Structure-Conditioned and Metal-Aware Flow Matching for CT Metal Artifact Reduction ​

Author: Heran Wang, Jianing Sun, Xu Jiang, Genwei Ma, Xing Zhao, Jigang Duan
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.28759v1 Announce Type: cross Abstract: In X-ray CT, metallic objects cause beam hardening, photon starvation, and scattering, leading to projection inconsistency, streaks, dark bands, and structural distortions that compromise clinical diagnosis and quantitative analysis. Existing metal a...

📖 Read original article


70. WaiT for the Signal: Simple Frequency-Aware Flow-Matching ​

Author: Krunoslav Lehman Pavasovic, Th'eophane Vallaeys, St'ephane Mallat, Giulio Biroli, Luke Zettlemoyer, Brian Karrer, Jakob Verbeek
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, stat.ML

arXiv:2607.28760v1 Announce Type: cross Abstract: As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for generation quality. However, standard flow matching treats all spatial frequencies uniformly, ignoring the natu...

📖 Read original article


71. Stratified Negation in RDF Rules: A Correct Approach (Extended Version) ​

Author: Nils K"uchenmeister, Alex Ivliev, D"orthe Arndt, Markus Kr"otzsch
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.DB

arXiv:2607.28778v1 Announce Type: cross Abstract: Combining RDF rule languages, such as N3 or SHACL Rules, with default negation is challenging. Existing methods to stratify negation often fail for RDF rules, since individual triples do not carry enough information to meaningfully restrict potential...

📖 Read original article


72. Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation ​

Author: Philipp D. Siedler, Jordan Sassoon
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.28801v1 Announce Type: cross Abstract: Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples. We introduce a dataset-centric meta-evaluation fram...

📖 Read original article


73. Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing ​

Author: Weiying Chen, Junlong Shen, Zhexuan Tang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.28814v1 Announce Type: cross Abstract: In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance, a move that can fail in two opposite ways: capitulation (abandoning the change agenda to preserve rapport) or c...

📖 Read original article


74. Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity ​

Author: Naman Saxena, Mudit Gaur, Vaneet Aggarwal
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.28849v1 Announce Type: cross Abstract: Bilevel reinforcement learning (RL) is an important framework within the literature of RL that can be used to formalize various categories of problems, such as meta-learning, hierarchical task decomposition, and reinforcement learning from human feed...

📖 Read original article


75. A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging ​

Author: Diego J. Torrej'on, Luna Y. Hern'andez, Javier S'anchez
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.28858v1 Announce Type: cross Abstract: Automatic brain tumor segmentation from magnetic resonance imaging (MRI) has become a fundamental task in computer-assisted diagnosis, treatment planning, and disease monitoring. Although numerous deep learning architectures have recently been propos...

📖 Read original article


76. TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text ​

Author: Chengshuai Zhao, Pingchuan Ma, Dawei Li, Bohan Jiang, Zhiyuan Yu, Zhen Tan, Huan Liu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CR, cs.LG

arXiv:2607.28862v1 Announce Type: cross Abstract: The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultaneously raising growing concerns about unauthorized data exploitation and privacy leakage. Unlearnable examples ...

📖 Read original article


77. Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug? ​

Author: Xiaonan Xu, Wenjing Wu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.28871v1 Announce Type: cross Abstract: When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect. We measure how often that treatment is warranted. BSG-VA (buggy-state/candidate-state/gold-fix validation analysis) captures each validatio...

📖 Read original article


78. RareSense: Rarity-Aware Similarity Search for Anomaly Retrieval in Transactional Data ​

Author: Sidahmed Benabderrahmane, Talal Rahwan
Published: 8/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG

arXiv:2607.28879v1 Announce Type: cross Abstract: Similarity search over sparse set-valued data is often dominated by frequent background attributes because classical measures such as Jaccard, cosine, and Hamming compare objects through atomic overlap. IDF (Inverse document frequency) weighting part...

📖 Read original article


79. To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing ​

Author: Amir M. Ebrahimi, Mohammed Mehedi Hasan, Aaditya Bhatia, Gopi Krishnan Rajbahadur, Ahmed E. Hassan
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2607.28887v1 Announce Type: cross Abstract: Large language models increasingly write and repair production code, yet evidence is mounting that their test-passing patches leave codebases harder to maintain. We identify one concrete source: deletion avoidance, the systematic tendency to retain c...

📖 Read original article


80. Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use ​

Author: Alex Liu, Min Sun, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.28889v1 Announce Type: cross Abstract: Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models (LLMs) are frequently proposed as analytic assistants. The open questions are not whether LLMs can ...

📖 Read original article


81. Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth ​

Author: Alex Liu, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He, Min Sun
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.28890v1 Announce Type: cross Abstract: Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes human coding is the standard to approximate. This study provides empirical evidence that the presumpt...

📖 Read original article


82. TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models ​

Author: Aryan Vijay Bhosale, Harshit Rajgarhia, Abhishek Mukherji, Dinesh Manocha
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL

arXiv:2607.28896v1 Announce Type: cross Abstract: Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about the same audio? Curr...

📖 Read original article


83. Design Concept: Scaffolding Geopolitical Reflection Among Tech Workers ​

Author: Sydney Reis
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY

arXiv:2607.28904v1 Announce Type: cross Abstract: This paper presents a speculative Human-Computer Interaction design proposal for encouraging geopolitical reflexivity amongst tech workers at geopolitically relevant technology companies. Recent scholarship in International Relations and Science and ...

📖 Read original article


84. Gated Q-learning: Add Off-Policy Bias to Taste ​

Author: Brett Daley
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.28916v1 Announce Type: cross Abstract: Multistep credit assignment is critical for sample-efficient reinforcement learning, yet managing off-policy bias in Q-learning remains a fundamental challenge. For 30 years, practitioners have been limited to a binary choice: eliminate the bias at t...

📖 Read original article


85. FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation ​

Author: Martin Lukk (University of Toronto)
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2607.28934v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, fi...

📖 Read original article


86. DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models ​

Author: Omid Ahmadieh, Nima Karimian
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.28936v1 Announce Type: cross Abstract: Facial biometric identification relies on the distinctiveness of user attributes within a high-dimensional embedding space. However, the decision boundaries of deep face recognition (FR) systems are often sufficiently narrow that they can be conflate...

📖 Read original article


87. Retrieval-Driven Training-Free AI-Generated Video Attribution ​

Author: Renxi Cheng, Chaolei Han, Jie Gui, Hongsong Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.28955v1 Announce Type: cross Abstract: AI-generated videos are becoming increasingly realistic and difficult to distinguish from authentic ones, which facilitates malicious misuse and poses growing threats to cybersecurity and social governance. Attributing AI-generated videos to their sp...

📖 Read original article


88. Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates ​

Author: Weiyi He, Yuping Lin, Jiliang Tang, Yue Xing
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.28959v1 Announce Type: cross Abstract: Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large language models (LLMs). While existing mitigation strategies, e.g., latent a...

📖 Read original article


89. A robust association between LLM use and scientific productivity: Assessing stopping-time selection ​

Author: Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin
Published: 8/3/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.CY

arXiv:2607.28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effe...

📖 Read original article


90. RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images ​

Author: Renxi Cheng, Jie Gui, Hongsong Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.28974v1 Announce Type: cross Abstract: The rapid advancement of image generation models has made it increasingly difficult for people to distinguish AI-generated images from real ones. To prevent the potential risks associated with the misuse of fake images, AI-generated image detection h...

📖 Read original article


91. PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learning under Explicit Limits ​

Author: Namkyung Yoon, Sanghong Kim, Hwangnam Kim
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.28982v1 Announce Type: cross Abstract: Recent language models achieve strong performance across a variety of tasks, but conventional adaptation applies updates uniformly across training samples regardless of their local update benefit. We propose PARALLEL, a prefrontal-aligned reinforceme...

📖 Read original article


92. Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning ​

Author: Duy Tran Thanh, Thien-Phuc Doan, Long Nguyen-Vu, Ngo Tan Vu Khanh
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.MA

arXiv:2607.28986v1 Announce Type: cross Abstract: Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers. Existing retrieval-augmented methods score image-text alignme...

📖 Read original article


93. Point2Radio: A Foundation Model for Cross-Scene Radio Fields from Material-Aware Point Clouds ​

Author: Chaozheng Wen, Chenghong Bian, Hongze Chen, Jun Zhang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.CV, cs.LG

arXiv:2607.28994v1 Announce Type: cross Abstract: High-fidelity radio fields are typically simulated for every scene--transmitter configuration or fitted separately to each scene, failing to exploit propagation structures shared across environments. We present Point2Radio, a foundation model that le...

📖 Read original article


94. Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving ​

Author: Jiwei Yang, Zhengxian Chen, Chaosheng Huang, Jun Li
Published: 8/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.29031v1 Announce Type: cross Abstract: Existing autonomous-driving world models typically perform dense prediction of future videos, occupancy states, BEV representations, or agent motion. We argue that planning need not reconstruct the complete future world, but only focus on scene featu...

📖 Read original article


95. Improving scDiffusion with Sparsity-Biased Classifier-Free Guidance ​

Author: Yu Song, Hao Sun, Ikuko Nishikawa, Yen-Wei Chen
Published: 8/3/2026, 4:00:00 AM
Categories: q-bio.GN, cs.AI

arXiv:2607.29043v1 Announce Type: cross Abstract: Single-cell RNA sequencing (scRNA-seq) has become an essential tool in modern cellular biology, and generating accurate synthetic scRNA-seq data is becoming increasingly important. Although diffusion models have achieved promising results in conditio...

📖 Read original article


96. Learning Lookahead Lemmas for Neural Network Verification ​

Author: Liam Davis, Haoze Wu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO

arXiv:2607.29051v1 Announce Type: cross Abstract: State-of-the-art neural network verifiers use the branch-and-bound procedure as their core solving mechanism. We introduce an inprocessing framework for neural network verification driven by the lookahead procedure. Under this framework, lookahead de...

📖 Read original article


Author: Hanxiao Lu, Tianyi Zhang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA

arXiv:2607.29055v1 Announce Type: cross Abstract: Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. In case of incorrect or unsatisfactory outputs, users have to manually locate agent mistakes by inspecting agent trajectories (i.e., {\em failure attribution}) and provide fe...

📖 Read original article


98. Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives ​

Author: Sudhir Bharati, Rajendra K C Khatri, Sudip Bharati
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.29064v1 Announce Type: cross Abstract: Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclear how well large language models (LLMs) reproduce official crash coding. This study benchmarked six ...

📖 Read original article


Author: Theekshana Samaradiwakara, Nisansa de Silva, George C. Lobb
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.29066v1 Announce Type: cross Abstract: Deception detection has critical implications for legal proceedings, law enforcement, and online security. Although human judgment is limited in accuracy and scalability, Natural Language Processing (NLP) offers a data-driven alternative. We present ...

📖 Read original article


100. Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients ​

Author: Shengkun Zhu, Jinshan Zeng, Zhihua Allen-Zhao, Mayi Xu, Quanqing Xu, Wei Ren, Qiang Yang, Yang Liu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.29071v1 Announce Type: cross Abstract: Federated learning of foundation models faces a fundamental resource-asymmetry challenge: the institutions holding the most valuable domain-specific data cannot host billion-parameter models. Existing heterogeneous federated approaches attempt to bri...

📖 Read original article


101. metasignal: A Python Package for Comprehensive Metacognitive Analysis and Decision-Making ​

Author: Saurabh Ranjan, Mukesh Makwana, Konstantina Sokratous, Brian Odegaard
Published: 8/3/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, stat.AP

arXiv:2607.29093v1 Announce Type: cross Abstract: Metasignal is an open-source Python package for signal detection theory (SDT) and metacognitive measurement. It implements the 17 metacognitive measures evaluated by Rahnev (2025), together with the reference variables d' (perceptual sensitivity), re...

📖 Read original article


102. DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs ​

Author: Ziwei Cheng, Zhenhua Tan, Zhuomin Zhu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.29112v1 Announce Type: cross Abstract: Audio-visual speech recognition (AVSR) relies on effective fusion of audio and visual modalities, yet existing approaches treat cross-modal interaction as a single-step operation without structured iterative refinement. We present DoubleHelix, a mult...

📖 Read original article


Author: Sen Zhao, Cheng Liu, Shuyin Xia, Zhiyuan Liu, Yi Liu, Yi Wang, Wei Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SI, cs.AI

arXiv:2607.29115v1 Announce Type: cross Abstract: Link prediction aims to identify potential or future connections within a given graph structure. Position information is essential for link prediction, as it distinguishes homogeneous nodes through their relative relationships, facilitating the accur...

📖 Read original article


104. InferQ: A Database-Oriented Benchmark for Quantum Circuits Simulation ​

Author: Andrei Ilinescu, Aadi Patwardhan, Rihan Hai
Published: 8/3/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.DB

arXiv:2607.29134v1 Announce Type: cross Abstract: Recent work suggests that relational database management systems (RDBMSs) can execute quantum circuit simulation by compiling the simulation into SQL workloads (primarily join-and-aggregate tensor contractions). While early results are promising, the...

📖 Read original article


105. HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators ​

Author: Jiaquan Zhang, Shuxu Chen, Haifan Meng, Yi Lu, Zhihan Lyu, Fan Mo, Wei Dong, Yang Yang, Chaoning Zhang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.29135v1 Announce Type: cross Abstract: Neural operators provide fast surrogates for time-dependent partial differential equations (PDEs) by applying a learned evolution operator recursively to its own predictions, but this autoregressive rollout feeds every prediction error back as input,...

📖 Read original article


106. Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership ​

Author: Pawe{\l} Borsukiewicz, Daniele Lunghi, Wendk^uuni C. Ou'edraogo, Jacques Klein, Tegawend'e F. Bissyand'e
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.29144v1 Announce Type: cross Abstract: Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the generators that produce these datasets are trained on real faces, so synthetic data may still reveal their real sou...

📖 Read original article


107. Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations ​

Author: Johannes Mae{\ss}, Leon Werner, J. Thorben Frank, Winfried Ripken, Martin Michajlow, Joshua Futterer, Klaus-Robert M"uller, Stefan Chmiela
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.29158v1 Announce Type: cross Abstract: We introduce implicit machine learning force fields (I-MLFFs), which replace explicit stacks of neural network layers with self-consistent fixed-point equations. In molecular simulations, this formulation enables intermediate representations to be re...

📖 Read original article


108. Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory ​

Author: Jinghan Xu, Yiyong Xiao, Wanru Shao, Hankai Liu, Xinjin Li
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.29167v1 Announce Type: cross Abstract: Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations into persistent action context. We identify memory provenance laundering: during LLM-based memory consolidation, a...

📖 Read original article


109. ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency ​

Author: Wenda Yu, Tianshi Wang, Fengling Li, Xin Li, Jingjing Li, Lei Zhu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.29169v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the temporal alignment among visual observations, robot states, and executed actions. We introduce ActFo...

📖 Read original article


110. CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning ​

Author: Yuxin Chen, Hari Srikanth, Nathan Jew, Menglin Wu, Pengcheng Wang, Junli Ren, Masayoshi Tomizuka, Peng Xu, Jinyu Xie, Thomas Tian
Published: 8/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.29172v1 Announce Type: cross Abstract: While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-source, limiting downstream users' ability to adapt them to new tasks, embodiments, and deployment settin...

📖 Read original article


111. MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation ​

Author: Rongchao Xu, Lin Jiang, Dahai Yu, Ximiao Li, Guang Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.29177v1 Announce Type: cross Abstract: Utility data (e.g., electricity, water, and gas consumption), collected by ubiquitous sensors and embedded devices, often contains substantial missing values due to various factors such as device failures and data transmission issues. The data missin...

📖 Read original article


112. MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation ​

Author: Yifei Zhu, Mingyi Shi, Yangyang Cai, Miao Cheng, Yoshifumi Kitamura, Taku Komura
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.29180v1 Announce Type: cross Abstract: Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible. A natural approach is to first project motion data into a structured semantic space and then train a generative model within ...

📖 Read original article


113. SERUM: State Extraction and Refinement for User Modeling ​

Author: Andy J. Phu, James Mooney, Karin de Langis, Khanh Chi Le, Dongyeop Kang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2607.29181v1 Announce Type: cross Abstract: Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow. However, building these models from raw, unstructured screen activity remains an open challenge. We present SERUM, a multi-pass ...

📖 Read original article


114. SAF-OPD: Stable Advantage Fusion for On-Policy Distillation ​

Author: Yifan Ding, Xincheng Wei, Yoshua Y. Li, Ziheng Li, Yuquan Lu, Siyu Zhang, Dongsheng Ma, Rongxiang Weng, Xunliang Cai, Yun Chen
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.29209v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher qual...

📖 Read original article


115. MOSAIC: Masked Outsourcing of Secure AI Computations ​

Author: James Hsin-yu Chiang, Sheila Zingg, Kari Kostiainen, Srdjan Capkun
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.29221v1 Announce Type: cross Abstract: We address the challenge of securely and efficiently outsourcing AI computations from a trusted but computationally weak client to an untrusted but powerful server, in the setting where the client holds both the input and the model, and the server mu...

📖 Read original article


116. Linear Proposal Operators and Stochastic Search Geometry in SOMA and Differential Evolution ​

Author: Vojt\v{e}ch Nov'ak, Ivan Zelinka
Published: 8/3/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2607.29228v1 Announce Type: cross Abstract: Swarm and evolutionary algorithms are usually analyzed as complete procedural systems in which nonlinear selection, replacement, and adaptation obscure simpler structure within candidate generation. This paper introduces an operator--selection factor...

📖 Read original article


117. FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution ​

Author: Peize Li, Ruimeng Zhang, Ru Zhang, Cong Huang, Kai Chen, Shanghang Zhang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.SY, eess.SY

arXiv:2607.29235v1 Announce Type: cross Abstract: Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands repeated re-grounding in real observations--not recursive rollout. Existing WAMs address this by ref...

📖 Read original article


118. Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters ​

Author: Antorweep Chakravorty
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.ET, cs.HC

arXiv:2607.29238v1 Announce Type: cross Abstract: InMyStyle is a privacy first, single user system that adapts small language models to rewrite AI-edited text towards an individual user's writing style without an instruction prompt at inference. Given a user's documents, it uses multiple local helpe...

📖 Read original article


119. When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration ​

Author: Kesheng Chen, Yamin Hu, Wenjian Luo
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.29240v1 Announce Type: cross Abstract: In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence of an atypical state. For example, a model may report that a visibly six-fingered hand has five fingers. We sho...

📖 Read original article


120. RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems ​

Author: Haoran Ling, Yuecheng Li, Zeyu Song, Jing Yao, Shuwen Kang, Chi Lu, Wenjin Wu, Peng Jiang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2607.29241v1 Announce Type: cross Abstract: Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes. While LLM-based agents can automate this trial-and-error process, allowing the LLM to both select...

📖 Read original article


121. TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation ​

Author: Alessandra Zerillo, Stefano Cannata, Diego Bellavia, Daniele Ciriello, Simone Manini, Salvatore Pasta, Caterina Gandolfo
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.29243v1 Announce Type: cross Abstract: Computed tomography angiography (CTA) is crucial for preprocedural TAVI planning, providing the anatomical information required for prosthesis sizing and vascular access assessment. As the volume of TAVI procedure increases, improving efficiency and ...

📖 Read original article


122. CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation ​

Author: Mengting Chen, Yanshu Sun, Wanting Liang, Beidi Luan, Rui Sun, Dezhi Chen, Jing Li, Zuo Bai
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.29252v1 Announce Type: cross Abstract: Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measu...

📖 Read original article


123. OsteoCAD: A Human-in-the-Loop Cloud-Edge Framework for Bone Tumor Segmentation ​

Author: Maximo Rodriguez-Herrero, Dante D. Sanchez-Gallegos, Heriberto Aguirre-Meneses, Marco Antonio N'u~nez-Gaona, J. L. Gonzalez-Compean, Jesus Carretero
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.29266v1 Announce Type: cross Abstract: Artificial Intelligence (AI) and Deep Learning (DL) have notably advanced medical image analysis, yet many health- care organizations struggle to adopt them due to limited com- putational resources and specialized expertise. To address these barriers...

📖 Read original article


124. Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation ​

Author: Yongshi Ye, Biao Fu, Chongxuan Huang, Yidong Chen, Xiaodong Shi
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.29287v1 Announce Type: cross Abstract: Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by human translators' ability to adapt reasoning effort based on difficulty, we propose TwT (Translation with Tho...

📖 Read original article


125. The persuasive power of large language models does not depend on their perceived national origin ​

Author: Ningzhi Liu, Yannic Hinrichs, Jonas R. Kunst
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.29334v1 Announce Type: cross Abstract: Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rejected as foreign propaganda, with consequences for democratic discourse and information sovereignty. Yet, wheth...

📖 Read original article


126. DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation ​

Author: Fernando Garc'ia-Torres, Roc'io del Amor, Sandra Morales, 'Alvaro Barroso, Peter Heiduschka, Bj"orn Kemper, Valery Naranjo
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.29337v1 Announce Type: cross Abstract: Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of annotated data in medical imaging, particularly in optical coherence tomography (OCT) of mouse eyes, where manua...

📖 Read original article


127. SeekBrain: An Autonomous Multi-Agent System for Accelerating Neuroscience Discovery ​

Author: Jiamin Wu, Peishan Xiang, Jingyang Chen, Yuqing Zhu, Yuxi Li, Ling Luo, Qihao Zheng, Jialiang Zu, Yongchao Wu, Mindong Liu, Haitao Wu, Chaofan Hu, Yijie Sun, Yuqi Hang, Yu Zhu, Shuo Li, Yue Fan, Shiyang Feng, Wanghan Xu, Tianlei Zhang, Jie Zhang, Wenlong Zhang, Bo Zhang, Kai Wang, Lei Bai, Mianxin Liu, Wanli Ouyang, Jiulin Du, Chunfeng Song
Published: 8/3/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2607.29347v1 Announce Type: cross Abstract: Modern neuroscience relies on integrating multi-scale, multimodal datasets to uncover the neural principles underlying intelligence. However, analytical challenges posed by highly heterogeneous data and fragmented workflows increasingly constrain dis...

📖 Read original article


128. Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning ​

Author: Douwe den Blanken, Martin Lefebvre, Charlotte Frenkel
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.AS

arXiv:2607.29353v1 Announce Type: cross Abstract: With the ever-increasing pervasiveness of smart edge devices, the demand is growing for applications that can be tailored to users (e.g., custom keyword spotting) or patients (e.g., adaptive health monitoring). Yet, most edge devices rely on fixed in...

📖 Read original article


129. Cross-Lingual Transfer for Machine Translation in Turkic Languages ​

Author: Omer Burak Cinar, Mehmet Mert Dalkilic, Cagri Toraman
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.29355v1 Announce Type: cross Abstract: Cross-lingual transfer is central to low-resource machine translation, but its behavior within closely related language families remains insufficiently characterized. We study transfer among five Turkic languages; Turkish, Azerbaijani, Uzbek, Kazakh,...

📖 Read original article


130. Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens ​

Author: Yi Luo, Rongzhi Gu, Jixun Yao
Published: 8/3/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.LG, cs.SD

arXiv:2607.29363v1 Announce Type: cross Abstract: Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation. Representations with higher frame rates or greater capacity can preserve more signal detail, but...

📖 Read original article


131. Dense Temporal Contrast Synthesis via Conditioned Latent Transport ​

Author: Smriti Joshi, Apostolia Tsirikoglou, Daniel M. Lang, Richard Osuala, Noah M'arquez Varaa, Alejandro Guzman, Grzegorz Skorupko, Sebastian Ibarra Arregui, Lidia Garrucho, Akane Ohashi, Dimitra Ntoula, Eugen Divjak, O\u{g}uz Lafc{\i}, Jan C. Peeken, Julia A. Schnabel, Fredrik Strand, Oliver Diaz, Karim Lekadir
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.29394v1 Announce Type: cross Abstract: Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based contrast agents (GBCAs) restricts use in contraindicated populations, prolongs scan protocols, and presents env...

📖 Read original article


132. Explore Beyond the Boundary Using Entropic Information ​

Author: Bumgeun Park, Donghwan Lee
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.29419v1 Announce Type: cross Abstract: In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guiding the learning process. Addressing this issue requires extensive exploration in the state space to...

📖 Read original article


133. AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair ​

Author: Michael Fu, Qiyue Mei, Patanamon Thongtanunam, Kla Tantithamthavorn
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CR

arXiv:2607.29422v1 Announce Type: cross Abstract: Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Recent agentic AI approaches have shown promising results in automated program repair. However, vulnerability repai...

📖 Read original article


134. QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models ​

Author: Xiang Chen, Yingying Zhao, Chao Li, Jiaju Han, Ben Zhang, Ang Li, Jiahuan Long, Yiwei Wei, Jiujiang Guo, Chengyin Hu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.29445v1 Announce Type: cross Abstract: Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering. However, their robustness to structured thermal perturbations and the stability of cross-modal sem...

📖 Read original article


135. TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate Fusion ​

Author: Yu Sun, Yuan Chang, Xiaohou Shi, Yan Sun
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.29459v1 Announce Type: cross Abstract: Large-scale multivariate time series from heterogeneous IoT sensors demand accurate long-term forecasting for resource scheduling and predictive maintenance. While recent time series foundation models exhibit strong generalization, they rely on stati...

📖 Read original article


Author: Jiayang Niu, Yan Wang, Jie Li, Ke Deng, Azadeh Alavi, Muhammad Usman, Yongli Ren
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.29491v1 Announce Type: cross Abstract: Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known. We introduce Drea...

📖 Read original article


137. From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale ​

Author: Chandra Maddila, Mashrur Rashik, Euna Mehnaz Khan, Smriti Jha, James Saindon, Nachi Nagappan, Peter C. Rigby
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2607.29516v1 Announce Type: cross Abstract: AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code review tools over-index on low-value suggestions such as style and best practices while under-indexing on the conc...

📖 Read original article


138. TerraNova: A Foundation Model for the Anthropocene ​

Author: Carlos Rodriguez-Pardo, Massimo Tavoni
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY, econ.EM, stat.ML

arXiv:2607.29527v1 Announce Type: cross Abstract: A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representation spans their observational breadth. We argue the obstacle is geometric: the physical Earth is measured as co...

📖 Read original article


139. ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation ​

Author: Gaetano Perrone, Simon Pietro Romano
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.29539v1 Announce Type: cross Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it remains unclear whether ...

📖 Read original article


140. MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models ​

Author: Boxiao Wang, Runxiang Wang, Kai Li, Chongming Li, Zhiwei Chen, Yifan Zhang, Jian Cheng
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.29561v1 Announce Type: cross Abstract: Symbolic Regression (SR) aims to discover analytical equations from observational data and plays a central role in scientific modeling. While recent Large Language Model (LLM) based approaches show promise, they face two limitations. First, they lack...

📖 Read original article


141. TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning ​

Author: Binnan Liu, Yechi Ma, Tian Xie, Wei Hua
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.29586v1 Announce Type: cross Abstract: The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and apply it to a new grid. Looped visual reasoners refine predictions over multiple iterations, but conventional tra...

📖 Read original article


142. FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models ​

Author: Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony D'Avirro, Benjamin Peloquin
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.HC

arXiv:2607.29602v1 Announce Type: cross Abstract: Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation...

📖 Read original article


143. A Human-Centered Validation of the Explainability-Performance Coefficient ​

Author: Christian Oliva, Luis F. Lago-Fern'andez
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2607.29614v1 Announce Type: cross Abstract: The rapid adoption of deep learning models in high-risk domains has intensified the need for trustworthy Explainable Artificial Intelligence (XAI). However, objectively evaluating explanation fidelity and aligning XAI metrics with human-centered unde...

📖 Read original article


144. When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning ​

Author: Luca Viano, Antoine Moulin, Audrey Huang, Volkan Cevher, Philip Amortila, Dylan J. Foster
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2607.29617v1 Announce Type: cross Abstract: Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer from compounding err...

📖 Read original article


145. CENDRe: Concept Extraction with Natural Domain Representations ​

Author: Antonia Holzapfel, Andres Felipe Posada Moreno, Sebastian Trimpe
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.29621v1 Announce Type: cross Abstract: Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical domains requires understanding the temporal and spectral patterns that drive their predictions. Concept extraction (CE) methods iden...

📖 Read original article


146. The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations ​

Author: Ilya Mikhelson
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2607.29624v1 Announce Type: cross Abstract: Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback. Conversely, traditional face-to-face oral examinations introduce severe construct-irrelevant variance by...

📖 Read original article


147. SATViz: Real-Time Visualization of Clausal Proofs ​

Author: Tim Holzenkamp, Kevin Kuryshev, Thomas Oltmann, Lucas W"aldele, Johann Zuber, Tobias Heuer, Ashlin Iser
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2209.05838v2 Announce Type: replace Abstract: Visual layouts of graphs representing SAT instances can highlight the community structure of SAT instances. The community structure of SAT instances has been associated with both instance hardness and known clause quality heuristics. Our tool SATVi...

📖 Read original article


148. Combining Large Language Models and Symbolic Reasoning for Multi-Robot Temporal Planning through Explainable Knowledge Bases ​

Author: Enrico Saccon, Matteo Saveriano, Edoardo Lamon, Luigi Palopoli, Marco Roveri
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.RO

arXiv:2502.19135v2 Announce Type: replace Abstract: We present PLANTOR, a framework for generating and executing multi-robot task plans from natural-language task descriptions through LLM-assisted knowledge-base construction. The approach uses large language models to synthesize a structured Prolog ...

📖 Read original article


149. Shall We Play a Game? Language Models for Open-ended Wargames ​

Author: Glenn Matlin, Isaac Song, Yixiong Hao, Parv Mahajan, Evan Montoya, Ryan Bard, Stuart R. Topp, Anthony Wen-Ming Zang, Mohammed Rehan Parwani, Soham Shetty, Mark Riedl
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2509.17192v3 Announce Type: replace Abstract: LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing several different jobs: choosing what an actor says or does, deciding what happens after an action,...

📖 Read original article


150. Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning ​

Author: Alexander Meulemans, Rajai Nasser, Maciej Wo{\l}czyk, Marissa A. Weis, Seijin Kobayashi, Blake Richards, Guillaume Lajoie, Angelika Steger, Marcus Hutter, James Manyika, Rif A. Saurous, Jo~ao Sacramento, Blaise Ag"uera y Arcas
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.22226v2 Announce Type: replace Abstract: The standard theory of model-free reinforcement learning assumes that the environment dynamics are stationary and that agents are decoupled from their environment, such that policies are treated as being separate from the world they inhabit. This l...

📖 Read original article


151. Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents ​

Author: Reuben Tan, Baolin Peng, Zhengyuan Yang, Hao Cheng, Oier Mees, Theodore Zhao, Andrea Tupini, Isar Meijer, Qianhui Wu, Yuncong Yang, Lars Liden, Yu Gu, Sheng Zhang, Xiaodong Liu, Lijuan Wang, Marc Pollefeys, Yong Jae Lee, Jianfeng Gao
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2512.03438v3 Announce Type: replace Abstract: Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally optimized using sparse, outcome-based rewards computed based on the final answers. Richer rewards compu...

📖 Read original article


152. M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities ​

Author: Ao Li, Jinghui Zhang, Luyu Li, Yuxiang Duan, Lang Gao, Mingcai Chen, Weijun Qin, Shaopeng Li, Fengxian Ji, Ning Liu, Lizhen Cui, Xiuying Chen, Yuntao Du
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.02854v2 Announce Type: replace Abstract: As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to improve answer quality and support complex reasoning. However, existing research on MAD suffers from two funda...

📖 Read original article


153. RAPiD: Reward-Guided Consistency Distillation of Diffusion Planners for Real-Time Autonomous Driving ​

Author: Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.RO

arXiv:2602.07339v2 Announce Type: replace Abstract: Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bottleneck for real-time closed-loop deployment. We present RAPiD, a reward-guided consistency distillation frame...

📖 Read original article


154. Shaping Scientific Explanations to Expert Perspectives with Persona-Conditioned Reinforcement Learning ​

Author: Susana Nunes, Tiago Guerreiro, Catia Pesquita
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2603.21846v2 Announce Type: replace Abstract: Explainable AI is increasingly important to scientific discovery. However, existing methods largely ignore that explanation quality is not universal: experts differ in how they assess evidence, prioritize mechanisms, and construct explanatory narra...

📖 Read original article


155. What Makes a Sale? Simulating End-to-End Seller--Buyer Retail Dynamics with LLM Agents ​

Author: Jeonghwan Choi, Jibin Hwang, Gyeonghun Sun, Minjeong Ban, Taewon Yun, Hyeonjae Cheon, Hwanjun Song
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2604.04468v2 Announce Type: replace Abstract: Evaluating retail strategies before deployment is difficult, as outcomes are determined across multiple stages, from seller-side persuasion through buyer-seller interaction to purchase decisions. However, existing retail simulators capture only par...

📖 Read original article


156. PEMAND: Persona-Enriched Multi-Agent Negotiation for Household Decision-Making ​

Author: Yuran Sun, Mustafa Sameen, Yaotian Zhang, Rongguan Gu, Mrunal Vibhute, Chia-yu Wu, Yuanyuan Lei, Xilei Zhao
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.10475v2 Announce Type: replace Abstract: Modeling household-level decisions is central to many real-world applications, including trip planning, residential mobility and migration, disaster management, etc. Existing studies primarily rely on classical machine learning models with limited ...

📖 Read original article


157. SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios ​

Author: Jackson Clark, Yiming Su, Saad Mohammad Rafid Pial, Yifang Tian, Lily Gniedziejko, Hans-Arno Jacobsen, Yinfang Chen, Tianyin Xu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.07161v3 Announce Type: replace Abstract: AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE). Current SRE benchmarks are limited to oversimplistic SRE tasks and are unfortunately hard to extend due to...

📖 Read original article


158. Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling ​

Author: Rongman Xu, Yifei Li, Tianzhe Zhao, Yanrui Wu, Bo Li, Hang Yan
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.15100v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning. However, maximizing their potential through inference-time scaling faces challenges in trade-off between sampling budget and reasoning quality. Current strategies rem...

📖 Read original article


159. PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization ​

Author: Wonjoong Kim, Yeonjun In, Sangwu Park, Dongha Lee, Chanyoung Park
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.17877v2 Announce Type: replace Abstract: A significant hurdle for current LLMs is the execution of complex, multi-stage tasks. Group Relative Policy Optimization (GRPO) has been emerging as a leading choice, but its reliance on sparse outcome rewards severely limits credit assignment acro...

📖 Read original article


160. The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models ​

Author: Kuan-Yen Chen, Fang-Yi Su, Shih-Yen Lin, Bao Li, Jung-Hsien Chiang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2606.05976v2 Announce Type: replace Abstract: Recent works show that LLM agents struggle to correct errors in their own reasoning traces, despite their ability to correct errors from external sources. We ask whether this reflects a capability deficit or an artifact of the role labeling. To tes...

📖 Read original article


161. A Multi-Agent System for Motor Design Optimization via an FEA-AI Hybrid Approach ​

Author: Jinseong Han, Sunwoong Yang, Namwoo Kang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2606.09037v2 Announce Type: replace Abstract: This study presents a large language model (LLM)-based multi-agent framework for interior permanent magnet synchronous motor (IPMSM) design optimization that mitigates limitations of conventional workflows: expertise-dependent problem setup and dat...

📖 Read original article


162. Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution ​

Author: Xucong Wang, Ziyu Ma, Shidong Yang, Tongwen Huang, Pengkun Wang, Yong Wang, Xiangxiang Chu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.10917v2 Announce Type: replace Abstract: Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static training environments, which hinder broader generalization. To address ...

📖 Read original article


163. ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning ​

Author: Xucong Wang, Ziyu Ma, Yong Wang, Shidong Yang, Hailang Huang, Renda Li, Pengkun Wang, Xiangxiang Chu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.13316v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods often encourage unnecessarily long reasoning rollouts, which can degra...

📖 Read original article


164. Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents ​

Author: Minghui Ma, Bin Guo, Hao Wang, Han Wang, Mengqi Chen, Jingqi Liu, Yan Liu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.29495v2 Announce Type: replace Abstract: As LLM-based conversational agents advance toward increasingly open-ended and interaction-intensive scenarios, task completion alone provides an incomplete assessment of their effectiveness. The evolution of users' internal states, including belief...

📖 Read original article


165. EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures ​

Author: Bu\u{g}ra Alperen Ulu{\i}rmak, Rifat Kurban
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.SE

arXiv:2606.30219v5 Announce Type: replace Abstract: This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capabiliti...

📖 Read original article


166. Latent Actions from Factorized Transition Effects under Agent Ambiguity ​

Author: Heejeong Nam, Chandradithya S Jonnalagadda, Harshit Aggarwal, Eric Xu, Randall Balestriero
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.30544v2 Announce Type: replace Abstract: Latent Action Models (LAMs) learn action-like proxies from observation. However, in multi-object or distractor-rich scenes, observations contain not only agent motion but also distractors, camera dynamics, and background changes, making recovery of...

📖 Read original article


167. LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents ​

Author: Jingpu Yang, Fengxian Ji, Zhengzhao Lai, Zhexuan Cui, Guangxian Ouyang, Qian Jiang, Fan Zhang, Min Peng, Qianqian Xie, Preslav Nakov, Zhuohan Xie
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.31045v2 Announce Type: replace Abstract: Scientific embodied agents are increasingly capable of carrying out laboratory procedures, but executing these procedures safely in dynamic laboratory environments remains challenging. Current safety approaches often overlook the intermediate step ...

📖 Read original article


168. Solution Space Path Planning: A Real-Time Human-Centered Path Planning Algorithm for En-Route Air Traffic Control ​

Author: Yiyuan Zou, Wenying Lyu, Clark Borst
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.RO, cs.SY, eess.SY

arXiv:2607.00064v2 Announce Type: replace Abstract: As technology advances, various algorithms have been proposed for air traffic management, yet their operational adoption in tactical control remains limited. This gap motivates a human-centered design emphasizing algorithmic interpretability, contr...

📖 Read original article


169. The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale ​

Author: Wenhui Chen, Jianlin Chen, Ziyao Lin, Chi Man Vong
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI, cs.IT, math.IT

arXiv:2607.14144v2 Announce Type: replace Abstract: The Platonic Representation Hypothesis (PRH) holds that as models scale, representations of heterogeneous networks converge toward a shared model of reality. We propose its sequel and boundary, the Capability Convergence Hypothesis (CCH): under a f...

📖 Read original article


170. NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning ​

Author: Hui Yang, Jiaoyan Chen, Yiping Song, Renate Schmidt, Wen Zhang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.15776v2 Announce Type: replace Abstract: OWL ontologies provide a formal knowledge representation framework that enables semantic reasoning, and have been widely adopted across domains such as healthcare and bioinformatics. In practice, however, real-world ontologies are often incomplete,...

📖 Read original article


171. Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs ​

Author: Harry Rogers, Sally Shiels, Ashley Tomlinson, James Thomas, James Aylward, Nathan Gauge, Helen Higham, Alison Noble
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.19063v3 Announce Type: replace Abstract: Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and cognitive bias. Standard examiner validation via inter-rater statistic...

📖 Read original article


172. CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents ​

Author: Qijia He, Jiayi Cheng, Chenqian Le, Rui Wang, Xunmei Liu, Yixian Chen, Jie Mei, Zhihao Wang, Xupeng Chen, Yuhuan Chen, Tao Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.19338v2 Announce Type: replace Abstract: Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap mo...

📖 Read original article


173. AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction ​

Author: Qinfeng Li, Yuntai Bao, Xinyan Yu, Hongze Chen, Yanmin Liu, Wenqi Zhang, Xuhong Zhang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.21106v2 Announce Type: replace Abstract: Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory meth...

📖 Read original article


174. Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning ​

Author: Guanqun Zhao, Zijun Xie, Binbin Zheng, Enlei Gong, Jiafeng Lu, Yehan Yang, Aoqi Hu, Zeyu Chen
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22186v2 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause pol...

📖 Read original article


175. DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training ​

Author: Hanlin Du, Zhiyuan Yan, Yungang Bao, Sa wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.22614v2 Announce Type: replace Abstract: RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU realloc...

📖 Read original article


176. From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement ​

Author: Qinsi Wang, Jing Shi, Huazheng Wang, Kun Wan, Yiran Wu, Bo Liu, Qingyun Wu, Hai Helen Li, Yiran Chen, Handong Zhao, Wentian Zhao
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23802v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathemat...

📖 Read original article


177. Reason-Mediated Behavioral Models for Auditing LLM Social Simulators ​

Author: Atharva Pandey, Gautam Jajoo
Published: 8/3/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.24649v2 Announce Type: replace Abstract: Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can matc...

📖 Read original article


178. Information Processing by Neuron Populations in the Central Nervous System: A Theory of the Mathematical Structure of Data and Operations ​

Author: Martin N. P. Nilsson
Published: 8/3/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.LG, cs.NE

arXiv:2309.02332v3 Announce Type: replace-cross Abstract: In the mammalian central nervous system, neurons are organized into populations communicating by spike trains propagating along axonal bundles. How such populations encode and transform information is only partially understood. In this study ...

📖 Read original article


179. On the Expressive Power of Sparse Geometric MPNNs ​

Author: Yonatan Sverdlov, Nadav Dym
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2407.02025v5 Announce Type: replace-cross Abstract: Motivated by applications in chemistry and other sciences, we study the expressive power of message-passing neural networks for geometric graphs, whose node features correspond to 3-dimensional positions. Recent work has shown that such model...

📖 Read original article


180. Revisiting Multi-Permutation Equivariance through the Lens of Irreducible Representations ​

Author: Yonatan Sverdlov, Ido Springer, Nadav Dym
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2410.06665v4 Announce Type: replace-cross Abstract: This paper explores the characterization of equivariant linear layers for representations of permutations and related groups. Unlike traditional approaches, which address these problems using parameter-sharing, we consider an alternative meth...

📖 Read original article


181. Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook ​

Author: Florinel-Alin Croitoru, Andrei-Iulian Hiji, Vlad Hondru, Nicolae Catalin Ristea, Paul Irofti, Marius Popescu, Cristian Rusu, Radu Tudor Ionescu, Fahad Shahbaz Khan, Mubarak Shah
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.MM, cs.SD, eess.AS

arXiv:2411.19537v4 Announce Type: replace-cross Abstract: We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content. We identify various kinds of deepfakes and construct taxonomies of deepfake generation and detection method...

📖 Read original article


182. Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints ​

Author: Pavel Kolev, Marin Vlastelica, Georg Martius
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO

arXiv:2501.04426v2 Announce Type: replace-cross Abstract: Offline diversity maximization under imitation constraints can transform demonstration data into a set of distinct behavioral policies, improving robustness to distribution shift without additional environment interaction. In practice, howeve...

📖 Read original article


183. Dimensionality reduction for homological stability and global structure preservation ​

Author: Alexander Kolpakov, Igor Rivin
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MS

arXiv:2503.03156v4 Announce Type: replace-cross Abstract: We propose DiRe, a force-directed dimensionality reduction framework designed to preserve global structure and homological features while remaining practical on modern hardware. The method combines an initial embedding with a graph-based layo...

📖 Read original article


184. Reproducing Human Individual Motor Signatures: A Data-Driven Approach for Repetitive Motion ​

Author: Angelo Di Porzio, Marco Coraggio
Published: 8/3/2026, 4:00:00 AM
Categories: cs.GR, cs.AI, cs.LG, cs.SY, eess.SY

arXiv:2503.15225v3 Announce Type: replace-cross Abstract: The deployment of autonomous virtual avatars (in extended reality) and robots in human group activities---such as rehabilitation therapy, sports, and manufacturing---is expected to increase as these technologies become more pervasive. Designi...

📖 Read original article


185. StaQ: a Finite Memory Approach to Discrete Action Policy Mirror Descent ​

Author: Alex Davey, Alena Shilova, Brahim Driss, Riad Akrour
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2506.13862v2 Announce Type: replace-cross Abstract: In Reinforcement Learning (RL), regularization with a Kullback-Leibler divergence that penalizes large deviations between successive policies has emerged as a popular tool both in theory and practice. This family of algorithms, often referred...

📖 Read original article


186. Towards White-Box Deep Wireless Sensing ​

Author: Xie Zhang, Yina Wang, Chenshu Wu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2507.21799v2 Announce Type: replace-cross Abstract: The empirical success of deep learning has spurred its application to the radio-frequency (RF) domain, leading to significant advances in Deep Wireless Sensing (DWS). However, most existing DWS models remain black boxes, with ad-hoc architect...

📖 Read original article


187. Patch-Based 3D Variational Autoencoder for Super-Resolution of Turbulent Channel Flow ​

Author: Anuraj Maurya
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.flu-dyn

arXiv:2507.22082v2 Announce Type: replace-cross Abstract: Direct numerical simulation (DNS) accurately resolves all spatio-temporal scales of wall-bounded turbulence but becomes prohibitively expensive as the Reynolds number increases. Super-resolution (SR) provides a practical alternative by recons...

📖 Read original article


188. RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment ​

Author: Marcos Fuster-Pena, David de-Fitero-Dominguez, Antonio Garcia-Cabot, Eva Garcia-Lopez
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2507.22580v2 Announce Type: replace-cross Abstract: Automated Program Repair (APR) seeks to automatically correct software bugs without requiring human intervention. However, existing tools tend to generate patches that satisfy test cases without fixing the underlying bug, those are known as o...

📖 Read original article


189. "Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness ​

Author: Jonathan A. Karr Jr., Benjamin F. Herbst, Matthew L. Sisk, Xueyun Li, Ting Hua, Matthew Hauenstein, Georgina Curto, Nitesh V. Chawla
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL

arXiv:2508.13187v5 Announce Type: replace-cross Abstract: Homelessness is a persistent social challenge, impacting millions worldwide. Over 876,000 people experiencing homelessness (PEH) were recorded in the U.S. in 2025. Social bias is a significant barrier to alleviating homelessness, shaping publ...

📖 Read original article


190. Adaptive Policy Backbone via Shared Network ​

Author: Bumgeun Park, Donghwan Lee
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2509.22310v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has achieved impressive results across domains, yet learning an optimal policy typically requires extensive interaction data, limiting practical deployment. A common remedy is to leverage priors, such as pre-collec...

📖 Read original article


191. Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events ​

Author: Richeek Das, Kostas Daniilidis, Pratik Chaudhari
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.RO

arXiv:2509.25146v2 Announce Type: replace-cross Abstract: This paper develops a mathematical argument and algorithms for building representations of data from event-based cameras, that we call Fast Feature Field ($\text{F}^3$). We learn this representation by predicting future events from past event...

📖 Read original article


192. Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation ​

Author: Xiao He, Huangxuan Zhao, Guojia Wan, Jiancheng Pan, Yanxing Liu, Yong Luo, Juhua Liu, Yongchao Xu, Wei Zhou, Dacheng Tao, Bo Du
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.IR, cs.MM

arXiv:2510.12953v4 Announce Type: replace-cross Abstract: Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are adapted to structured adult imaging and underperform in fetal ultrasound, which poses challenges of mul...

📖 Read original article


193. Monotone and Separable Set Functions: Characterizations and Neural Models ​

Author: Soutrik Sarangi, Yonatan Sverdlov, Nadav Dym, Abir De
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2510.23634v4 Announce Type: replace-cross Abstract: Motivated by applications for set containment problems, we consider the following fundamental problem: can we design set-to-vector functions so that the natural partial order on sets is preserved, namely $S\subseteq T \text{ if and only if } ...

📖 Read original article


194. Pay for The Second-Best Service: A Game-Theoretic Approach Against Dishonest LLM Providers ​

Author: Yuhan Cao, Yu Wang, Sitong Liu, Miao Li, Yixin Tao, Tianxing He
Published: 8/3/2026, 4:00:00 AM
Categories: cs.GT, cs.AI

arXiv:2511.00847v5 Announce Type: replace-cross Abstract: The widespread adoption of Large Language Models (LLMs) through Application Programming Interfaces (APIs) induces a critical vulnerability: the potential for dishonest manipulation by service providers. This manipulation can manifest in vario...

📖 Read original article


195. Robust Bidirectional Associative Memory via Regularization Inspired by the Subspace Rotation Algorithm ​

Author: Ci Lin, Tet Yeap, Iluju Kiringa
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2511.11902v2 Announce Type: replace-cross Abstract: Bidirectional Associative Memory (BAM) trained with Bidirectional Backpropagation (B-BP) often suffers from poor robustness and high sensitivity to noise and adversarial attacks. To address these issues, we propose a novel gradient-free train...

📖 Read original article


196. AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance ​

Author: Tianling Xu, Shengzhe Gan, Leslie Gu, Yuelei Li, Fangneng Zhan, Hanspeter Pfister
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO

arXiv:2512.05131v2 Announce Type: replace-cross Abstract: Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes from pre-collected images. However, existing active reconstruc...

📖 Read original article


197. WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics ​

Author: Chenxu Liu, Yingjie Fu, Wei Yang, Ying Zhang, Tao Xie
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2601.02430v3 Announce Type: replace-cross Abstract: Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and commercial potential. However, building a benchmark for LLM-generated web apps remains challenging due ...

📖 Read original article


198. GPU-Accelerated ANNS: Quantized for Speed, Built for Change ​

Author: Hunter McCoy, Zikun Wang, Prashant Pandey
Published: 8/3/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2601.07048v5 Announce Type: replace-cross Abstract: Approximate nearest neighbor search (ANNS) is a core problem in machine learning and information retrieval applications. GPUs offer a promising path to high-performance ANNS: they provide massive parallelism for distance computations, are rea...

📖 Read original article


199. GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR ​

Author: Jiaying Zhang, Lei Shi, Jiguo Li, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.09361v4 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a key paradigm for improving large-scale reasoning models. Unlike supervised fine-tuning (SFT), RLVR exhibits distinct optimization dynamics and is sensitive to the preservation of pre-...

📖 Read original article


200. Knowledge Restoration-driven Prompt Optimization: Unlocking LLM Potential for Open-Domain Relational Triplet Extraction ​

Author: Xiaonan Jing, Gongqing Wu, Xingrui Zhuo, Lang Sun, Jiapu Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.15037v2 Announce Type: replace-cross Abstract: Open-domain Relational Triplet Extraction (ORTE) aims to mine structured knowledge without predefined relation schemas. Large Language Models (LLMs) have advanced ORTE toward a prompt-driven paradigm through powerful in-context learning. Howe...

📖 Read original article


201. When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering ​

Author: Mahdi Astaraki, Mohammad Arshi Saloot, Ali Shiraee Kasmaee, Hamidreza Mahyar, Soheila Samiee
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2601.19827v5 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative retrieval-reasoning loops meaningfully outperform static RAG, particularly in scientific domains requiring...

📖 Read original article


202. Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation ​

Author: Shun Qian, Bingquan Liu, Chengjie Sun, Zhen Xu, Baoxun Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.22546v2 Announce Type: replace-cross Abstract: The recent advancements in Large Language Models (LLMs) have attracted interest in exploring their in-context learning abilities and chain-of-thought capabilities. However, there are few studies investigating the specific traits related to th...

📖 Read original article


203. AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles ​

Author: Wenkai Fan, Shurui Zhang, Xiaolong Wang, Haowei Yang, Tsz Wai Chan, Xingyan Chen, Junquan Bi, Zirui Zhou, Jia Liu, Kani Chen
Published: 8/3/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2602.10429v2 Announce Type: replace-cross Abstract: AIvilization v0 is a publicly deployed large-scale artificial society that couples a resource-constrained sandbox with a unified LLM-agent architecture, aiming to sustain long-horizon autonomy while remaining executable under a rapidly changi...

📖 Read original article


204. Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization ​

Author: Theophilus Amaefuna, Hitesh Vaidya, Anshuman Chhabra, Ankur Mali
Published: 8/3/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, cs.LG, math.IT

arXiv:2603.00910v3 Announce Type: replace-cross Abstract: Layer-wise capacity in large language models is highly non-uniform: some layers contribute disproportionately to loss reduction, whereas others are nearly redundant. Existing layer-scoring methods provide sensitivity estimates but do not give...

📖 Read original article


205. Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery ​

Author: Chaoqun Yang, Xinyu Lin, Shulin Li, Wenjie Wang, Ruihan Guo, Fuli Feng, Tat-Seng Chua
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2603.03322v2 Announce Type: replace-cross Abstract: Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery. However, rigorously evaluating an AI's capacity for knowledge discovery remains a critical challenge. Existing b...

📖 Read original article


206. Stem: Rethinking Causal Information Flow in Sparse Attention ​

Author: Lin Niu, Xin Luo, Linchuan Xie, Yifu Sun, Guanghua Yu, Jianchen Zhu, S Kevin Zhou
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.06274v2 Announce Type: replace-cross Abstract: The quadratic computational complexity of self-attention remains a fundamental bottleneck for scaling Large Language Models (LLMs) to long contexts, particularly during the pre-filling phase. In this paper, we rethink the causal attention mec...

📖 Read original article


207. Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models ​

Author: Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Abdullah Ibne Hanif Arean, Juena Ahmed Noshin
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.06828v2 Announce Type: replace-cross Abstract: We uncover a behavioral law of long-horizon vision-language models: models that maintain temporally grounded beliefs generalize better. Standard benchmarks measure only final-answer accuracy, which obscures how models use visual information; ...

📖 Read original article


208. Wrong Code, Right Structure: Learning Netlist Representations from Imperfect LLM-Generated RTL ​

Author: Siyang Cai, Cangyuan Li, Haoyu Gao, Kun Wang, Yinhe Han, Ying Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR

arXiv:2603.09161v2 Announce Type: replace-cross Abstract: Learning effective netlist representations is fundamentally constrained by the scarcity of labeled datasets, as real designs are protected by Intellectual Property (IP) and costly to annotate. Existing work therefore focuses on small-scale ci...

📖 Read original article


209. ELISA: An Interpretable Hybrid Generative AI Agent for Expression-Grounded Discovery in Single-Cell Genomics ​

Author: Omar Coser
Published: 8/3/2026, 4:00:00 AM
Categories: q-bio.GN, cs.AI

arXiv:2603.11872v3 Announce Type: replace-cross Abstract: Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI systems lack direct access to transcriptomic representations while expression foundation models remain...

📖 Read original article


210. Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation ​

Author: Hanwen Shen, Ting Ying, Jiajie Lu, Shanshan Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2603.13683v4 Announce Type: replace-cross Abstract: Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias prompts. We demonstrate via out-of-distribution (OOD) detection that these high-bias prompts cause a distr...

📖 Read original article


211. Demystifying Video Reasoning ​

Author: Ruisi Wang, Zhongang Cai, Fanyi Pu, Junxiang Xu, Wanqi Yin, Maijunxian Wang, Ran Ji, Chenyang Gu, Bo Li, Ziqi Huang, Hokin Deng, Dahua Lin, Ziwei Liu, Lei Yang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.16870v3 Announce Type: replace-cross Abstract: Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabilities. Prior work attributes this to a Chain-of-Frames (CoF) mechanism, where reasoning is assumed t...

📖 Read original article


212. OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation ​

Author: Haoyang Fang, Shuai Zhang, Yifei Ma, Hengyi Wang, Cuixiong Hu, Katrin Kirchhoff, Bernie Wang, George Karypis
Published: 8/3/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL, cs.LG

arXiv:2603.17205v3 Announce Type: replace-cross Abstract: Domain-specific finetuning is essential for dense retrievers, yet not all data pairs contribute equally to the learning process. We introduce OPERA, a data pruning framework that exploits this heterogeneity to improve both the effectiveness a...

📖 Read original article


213. Agentic Harness for Real-World Compilers ​

Author: Yingwei Zheng, Cong Li, Shaohua Li, Yuqun Zhang, Zhendong Su
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2603.20075v2 Announce Type: replace-cross Abstract: Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements enable automated bug repair, compiler bugs pose unique challenges due to their complexity, deep cross-doma...

📖 Read original article


214. Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning ​

Author: Jiajun Hu, Nuria Armengol Urpi, Jin Cheng, Stelian Coros
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.25464v2 Announce Type: replace-cross Abstract: Zero-shot reinforcement learning (RL) algorithms aim to learn a family of policies from a reward-free dataset, and recover optimal policies for any reward function directly at test time. Naturally, the quality of the pretraining dataset deter...

📖 Read original article


215. Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations ​

Author: Xiao Ni, Yiwei Wang, Tianjun Feng, Lauren Xiaoyan Lu, Yitong Wang, Congyi Zhou
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2603.29888v2 Announce Type: replace-cross Abstract: In collaboration with Alibaba, we study how a generative AI assistant affects service performance in e-commerce after-sales operations. In a large-scale field experiment, human agents providing digital chat support were randomly assigned acce...

📖 Read original article


216. ActionParty: Multi-Subject Action Binding in Generative Video Games ​

Author: Alexander Pondaven, Ziyi Wu, Igor Gilitschenski, Philip Torr, Sergey Tulyakov, Fabio Pizzati, Aliaksandr Siarohin
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2604.02330v2 Announce Type: replace-cross Abstract: Recent advances in video diffusion have enabled the development of "world models" capable of simulating interactive environments. However, these models are largely restricted to single-agent settings, failing to control multiple agents simult...

📖 Read original article


217. Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation ​

Author: Geert Trooskens (XY.AI Labs, Palo Alto, CA), Aaron Karlsberg (XY.AI Labs, Palo Alto, CA), Anmol Sharma (XY.AI Labs, Palo Alto, CA), Lamara De Brouwer (XY.AI Labs, Palo Alto, CA), Max Van Puyvelde (Stanford University School of Medicine, Stanford, CA), Matthew Young (XY.AI Labs, Palo Alto, CA), John Thickstun (Cornell University, Ithaca, NY), Gil Alterovitz (Brigham and Women's Hospital / Harvard Medical School, Boston, MA), Walter A. De Brouwer (Stanford University School of Medicine, Stanford, CA)
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2604.05150v2 Announce Type: replace-cross Abstract: We study compiled AI, a paradigm in which large language models generate executable code artifacts during a compilation phase, after which workflows execute deterministically without further model invocation. This paradigm has antecedents in ...

📖 Read original article


218. Evaluating the Alignment Between GeoAI Explanations and Domain Knowledge in Satellite-Based Flood Mapping ​

Author: Hyunho Lee, Wenwen Li
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.26051v2 Announce Type: replace-cross Abstract: The increasing number of satellites has improved the temporal resolution of Earth observation, making satellite-based flood mapping a promising approach for operational flood monitoring. Deep learning-based approaches for flood mapping using ...

📖 Read original article


219. TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning ​

Author: Siyang Li, Yize Chen, Zijie Zhu, Yuxin Pan, Yan Guo, Ming Huang, Hui Xiong
Published: 8/3/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.CV, cs.LG

arXiv:2605.00015v2 Announce Type: replace-cross Abstract: Time Series Foundation Models (TSFMs) have demonstrated strong generalization capability and data efficiency in time series forecasting through large-scale pretraining. However, adapting TSFMs to downstream forecasting tasks remains challengi...

📖 Read original article


220. Escaping Mode Collapse in LLM Generation via Geometric Regulation ​

Author: Xin Du, Kumiko Tanaka-Ishii
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cond-mat.dis-nn, cs.AI, nlin.CD

arXiv:2605.00435v3 Announce Type: replace-cross Abstract: Mode collapse is a persistent challenge in generative modeling and appears in autoregressive text generation as behaviors ranging from explicit looping to gradual loss of diversity and premature trajectory convergence. We take a dynamical-sys...

📖 Read original article


221. Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs ​

Author: Michael Rottoli, Subhankar Roy, Stefano Paraboschi
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.04215v3 Announce Type: replace-cross Abstract: Diffusion-based Large Language Models (D-LLMs) represent a promising frontier in generative AI, offering fully parallel token generation that can lead to significant throughput advantages and superior GPU utilization over the traditional auto...

📖 Read original article


222. Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping ​

Author: Gabriel Jeanson, David-Alexandre Duclos, William Larriv'ee-Hardy, No'e Cochet, Mat\v{e}j Boxan, Anthony Desch^enes, Fran\c{c}ois Pomerleau, Philippe Gigu`ere
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.RO

arXiv:2605.05627v2 Announce Type: replace-cross Abstract: Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained. While Uncrewed Aerial Vehicles (UAVs) offer scalable data collection, the transit...

📖 Read original article


223. Detecting AI-Generated Videos with Spiking Neural Networks ​

Author: Minsuk Jang, Yujin Yang, Hee-Seon Kim, Minseok Son, Younghun Kim, Changick Kim
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.05895v2 Announce Type: replace-cross Abstract: Modern AI-generated videos are photorealistic at the single-frame level, leaving inter-frame dynamics as the main remaining axis for detection. Existing detectors typically handle this temporal evidence in three ways: feeding the full frame s...

📖 Read original article


224. A Nonlinear Singular Value Theory for Neural Networks ​

Author: Brian Charles Brown, Mauricio Munoz, Robert Bridges, David Grimsman, Sean Warnick
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.06938v2 Announce Type: replace-cross Abstract: Recently Brown et al. [2025] established a singular value decomposition (SVD) for maps (especially nonlinear) satisfying certain norm conditions. We prove that most modern neural architectures admit this nonlinear SVD (NLSVD) representation--...

📖 Read original article


225. DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain ​

Author: Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang, Bei Chen, Yufang Hou
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2605.07699v2 Announce Type: replace-cross Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently ambiguous domain policies that admit multiple valid interpretations. Despite the prevalence o...

📖 Read original article


226. Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning ​

Author: Isabella Luong, Joyee Chen, Sankalpa Ghose, David Williams-King, Linh Le, Allen Lu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG

arXiv:2605.16301v3 Announce Type: replace-cross Abstract: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBen...

📖 Read original article


227. When Bits Break Recourse: Counterfactual-Faithful Quantization ​

Author: Chaymae Yahyati, Ismail Lamaakal, Khalid El Makkaoui, Ibrahim Ouahbi
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2605.17160v2 Announce Type: replace-cross Abstract: Model quantization is widely used to reduce memory, latency, and deployment cost, and is typically judged by whether predictive accuracy is preserved. In decision systems that provide algorithmic recourse, however, accuracy preservation is no...

📖 Read original article


228. AI4BayesCode: From Natural Language Descriptions to Validated Modular Stateful Bayesian Samplers ​

Author: Jungang Zou, Alex Ziyu Jiang, Qixuan Chen
Published: 8/3/2026, 4:00:00 AM
Categories: stat.CO, cs.AI, cs.LG

arXiv:2605.18476v2 Announce Type: replace-cross Abstract: Coding and computation remain major bottlenecks in Markov chain Monte Carlo (MCMC) workflows, especially as modern sampling algorithms have become increasingly complex and existing probabilistic programming systems remain limited in model sup...

📖 Read original article


229. DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation ​

Author: Bo Ye, Xinyu Cui, Jian Zhao, Tong Wei, Min-Ling Zhang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.21028v5 Announce Type: replace-cross Abstract: Autoregressive long video generation often adopts bounded-memory streaming for efficiency, typically combining local windows for short-term continuity with static early-frame sinks as long-range anchors. However, this fixed allocation keeps e...

📖 Read original article


230. MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing ​

Author: Han Chen, Zining Zhang, Wenqi Pei, Bingsheng He, Ming Wu, Jason Zeng, Michael Heinrich, Wei Wu, Hongbao Zhang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.MA

arXiv:2605.23986v2 Announce Type: replace-cross Abstract: Memory is a fundamental component for long-context LLM agents, supporting persistent state across interactions through a continuous serve-and-update lifecycle. Despite substantial prior work, many stateful systems retain sequential autoregres...

📖 Read original article


231. PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis ​

Author: Lucas Tamic, Ilan Jaffeux-Cheniout, Xavier Marjou
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.05176v2 Announce Type: replace-cross Abstract: While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-specific constraints in telecommunications customer support remain limited. In addition, da...

📖 Read original article


232. Multi-Scale Feature Attention Network for Polymer Classification Using Terahertz Spectroscopy ​

Author: Roshni Mahtani, Il'an Carretero, Daniel Moreno-Paris, Aldo Moreno-Oyervides, Laura Monroy, Oscar El'ias Bonilla-Manrique, Roc'io del Amor
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.06554v3 Announce Type: replace-cross Abstract: Reliable polymer identification is essential for ensuring the quality and safety of recycled plastics, yet conventional sorting and spectroscopic techniques often struggle to deliver robust discrimination. Terahertz (THz) spectroscopy offers ...

📖 Read original article


233. APPO: Agentic Procedural Policy Optimization ​

Author: Xucong Wang, Ziyu Ma, Yong Wang, Yuxiang Ji, Shidong Yang, Guanhua Chen, Pengkun Wang, Xiangxiang Chu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.12384v2 Announce Type: replace-cross Abstract: Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing methods assign credit over coarse heuristic units, such as tool-cal...

📖 Read original article


234. Creative Integration: A Decidable Criterion of Creativity ​

Author: Yoshinori Nomura
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.13977v2 Announce Type: replace-cross Abstract: "Integrative" solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes the world cheaper to describe -- from a tidy re-description. Building on the lineage that treats creativ...

📖 Read original article


235. Implicit Reasoning for Large Language Model-based Generative Recommendation ​

Author: Yinhan He, Liam Collins, Bhuvesh Kumar, Jundong Li, Neil Shah, Donald Loveland
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.14142v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly adopted as backbones for Generative Recommendation (GR), promising access to pretrained world knowledge. Yet reliably invoking this knowledge for GR remains poorly understood. A key obstacle is th...

📖 Read original article


236. The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence ​

Author: David Nordfors
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2606.21008v2 Announce Type: replace-cross Abstract: The metanym game is a competitive word game for LLMs that measures structural intelligence against established cognitive-science constructs. No content is given in advance; the contestants create all of it -- a new kind of analogy test, analo...

📖 Read original article


237. SqLinear: Balanced Square Partitioning Makes Linear Interaction Sufficient for Large-Scale Traffic Forecasting ​

Author: Yongfeng Su, Hongwen Li, Zijian Zhang, Ziquan Fang, Lu Chen, Christian S. Jensen, Hong Gao, Yinjun Han
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.21072v2 Announce Type: replace-cross Abstract: Traffic prediction is a core task in intelligent transportation systems and urban-scale decision making. Despite the effectiveness of mainstream neural network-based methods, their deployment in real-world settings with thousands of traffic s...

📖 Read original article


238. DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training ​

Author: Haoning Wang, Yiwei Liu, Haisen Luo, Dan Liu, Junxi Yin, Haotian Wang, Lei Zhang, Xiaoyu Tian, Shuaiting Chen, Yuansheng Song, Baoyan Guo, Xiongfei Yan, Bolan Yang, Chengwei Liu, Ming Cui, Jiong Chen
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.30345v2 Announce Type: replace-cross Abstract: Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in complex reasoning tasks. Existing self-distillation and reinforcement learning methods lack explicit mechanis...

📖 Read original article


239. ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ​

Author: Zijun Xie, Binbin Zheng, Enlei Gong, Jihua Liu, Yuyang You, Lingfeng Liu, Jiayao Tang, Guanqun Zhao, Aoqi Hu, Zeyu Chen
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.31650v4 Announce Type: replace-cross Abstract: Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Context-management methods make such rollouts feasible by simplifying past interactions through deletion,...

📖 Read original article


240. MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding ​

Author: Wenda Wang, Yihan Tong, Yuwei Hu, Xuchen Pan, Zhewei Wei, Yaliang Li, Bolin Ding
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, q-bio.BM

arXiv:2607.01982v2 Announce Type: replace-cross Abstract: Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a new trend in tasks such as molecular design and drug discovery. However, these models struggle to fully ...

📖 Read original article


241. BeatEdit: Symbolic Music Generation as Explicit Editing ​

Author: Haoyu Gu, Lekai Qian, Haowu Zhou, Qi Liu, Shuai Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.11124v2 Announce Type: replace-cross Abstract: Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete sequences from scratch, with limited support for selective modification. Edit-based methods have proven ...

📖 Read original article


242. AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis ​

Author: Zhenqi Jia, Yuan Zhao, Aruukhan, Rui Liu, Haizhou Li
Published: 8/3/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2607.15755v2 Announce Type: replace-cross Abstract: Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predef...

📖 Read original article


243. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning ​

Author: Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.16057v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic too...

📖 Read original article


244. EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration ​

Author: Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee
Published: 8/3/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.18529v2 Announce Type: replace-cross Abstract: Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimoda...

📖 Read original article


245. CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization ​

Author: Xinting Liao, Behnoosh Zamanlooy, Masoumeh Shafieinejad, David B. Emerson, Ruinan Jin, Deval Pandya, Xiaoxiao Li
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.18622v2 Announce Type: replace-cross Abstract: Textual Collaborative Prompt Optimization (TCPO) extends TextGrad (Yuksekgonul et al., 2025) to a decentralized setting by allowing multiple clients to jointly improve prompts for large language models (LLMs) while keeping their data locally....

📖 Read original article


246. Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning ​

Author: Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.19345v2 Announce Type: replace-cross Abstract: Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier. However, we identify a critical failure mod...

📖 Read original article


247. HijackKV: New Threat in Position-Independent KV Cache Reuse ​

Author: Yichi Zhang, Zhiqi Wang, Huan Zhang, Yuchen Yang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2607.19957v2 Announce Type: replace-cross Abstract: Key-Value (KV) cache reduces inference latency in large language models (LLMs). Traditional prefix-based reuse has low cache hit rates across inference requests because it requires exact token and position matches. To improve efficiency, rece...

📖 Read original article


248. ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing ​

Author: Yueyi Liu, Chi Zhang, Sen Cui, Miao Liu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.21529v2 Announce Type: replace-cross Abstract: Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimi...

📖 Read original article


249. Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS ​

Author: Viktoriia Makovska, George Fletcher
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.22657v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing machine-unlearning algorithms can suppress this behavior. We introduce Level-based Evaluati...

📖 Read original article


250. Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric ​

Author: Nikolaos Kekatos, Panagiotis Katsaros, Alexios Lekidis, Theodoros Nestoridis, Tom Nianios
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LO, cs.RO

arXiv:2607.23532v2 Announce Type: replace-cross Abstract: Swarms of LLM-assisted autonomous robots are increasingly proposed for cooperative intelligence, surveillance, and reconnaissance (ISR) in contested environments. A growing class of their assurance failures arises not within any single platfo...

📖 Read original article


251. DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory ​

Author: Xingyang Yu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, hep-th

arXiv:2607.23614v2 Announce Type: replace-cross Abstract: We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, an...

📖 Read original article


252. Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature ​

Author: Tanjin He, Aikaterini Vriza, Logan Ward, Xu Huang, Yiming Chen, Anubhav Jain, Gerbrand Ceder, Rajeev S. Assary, Ian T. Foster, Maria K. Y. Chan
Published: 8/3/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, cs.CL, cs.DL, cs.IR

arXiv:2607.23886v3 Announce Type: replace-cross Abstract: X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described ...

📖 Read original article


253. Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls ​

Author: Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Khandaker Rifah Tasnia, Md Hasibul Amin, Sifat Rahman Ahona, Juena Ahmed Noshin
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.24343v2 Announce Type: replace-cross Abstract: Language-model agents act through structured tool calls whose arguments carry very different risks: untrusted content may legitimately shape an email body but should never set a recipient, account, command, or credential. Existing conformal r...

📖 Read original article


254. LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings ​

Author: Brittany Harbison, Ashok K. Goel
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.24435v2 Announce Type: replace-cross Abstract: Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining prevalence and agreement diagno...

📖 Read original article


255. A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks ​

Author: Du Yin, Xiachong Lin, Yue Tan, Jinliang Deng, Estrid He, Hao Xue, Flora D. Salim
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.25875v2 Announce Type: replace-cross Abstract: Traffic forecasting is important for efficient traffic management and route planning in smart cities. Existing traffic forecasting studies typically assume fixed sensor graphs, overlooking the continuous evolution of real-world traffic networ...

📖 Read original article


256. Progressive Multimodal Alignment for Continual Instruction Tuning ​

Author: Duzhen Zhang, Yahan Yu, Qiaoyi Su, Jiahua Dong, Tielin Zhang
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.26947v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In Multimodal Continual Instruction Tuning (MCIT), however, shifti...

📖 Read original article


257. Benchmarking LLM Competence on Logical Inference over Probability Operators ​

Author: Nayera Hasan, Jack Greff, Alvin Grissom II
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.27405v2 Announce Type: replace-cross Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such a...

📖 Read original article


258. SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups ​

Author: Yikun Bai, Binghang Lu, Yikai Liu, Elaheh Akbari, Soheil Kolouri, Linxuan Wang, Ping He, Shuchan Wang, Ruqi Zhang, Guang Lin
Published: 8/3/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.27431v2 Announce Type: replace-cross Abstract: Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Existing diffusion and flow-matching models produce high-quality backbones on SE(3)^N, but inference requir...

📖 Read original article


259. LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents ​

Author: Jingya Wang, Yuyang Gao, Liuzhenghao Lv, Yonghong Tian, Yuyang Liu
Published: 8/3/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.27690v2 Announce Type: replace-cross Abstract: We introduce LabEvolver, a training-free framework that equips safe and grounded wet-lab agents with episodic memory from execution experience. LabEvolver couples a state-grounded inner trial loop for adaptive perception, online planning, and...

📖 Read original article


260. Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation ​

Author: Yuhang Zhu, Mingxuan Du, Benfeng Xu, Jie Gao, Lingyun Yu, Hongtao Xie
Published: 8/3/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.27816v2 Announce Type: replace-cross Abstract: Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essentia...

📖 Read original article


261. On a joint simultaneous learning of relevant feature subsets and subspaces in regression-like problems ​

Author: Illia Horenko
Published: 8/3/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG

arXiv:2607.28080v2 Announce Type: replace-cross Abstract: We extend a recently introduced Entropy-Optimal Manifold Clustering (EOMC) to allow for a joint simultaneous identification of subsets and subspaces of relevant features in nonstationary and nonlinear regression problems. It is shown that the...

📖 Read original article