Skip to content

arXiv cs.AI - 2026-08-07 ​

307 items collected.


1. Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services ​

Author: Xi Wang, Kun Li, Xianyao Ling, Gang Yin, Liang Zhang, Jiang Wu, Wenbo Lei, Jun Xu, Annie Wang, Fu Zhang, Weizhe Wang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.05159v1 Announce Type: new Abstract: Enterprise operations extensively rely on multiple heterogeneous business systems and information applications, which also result in severe data silos and process fragmentation. Enterprises have invested considerable financial and material resources in...

📖 Read original article


2. The Ignition Index: Measuring Global Workspace Dynamics in Language Models ​

Author: Saman Rahbar
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.05160v1 Announce Type: new Abstract: We introduce the Ignition Index (I), a validated scalar metric that operationalizes Global Workspace Theory's (GWT) all-or-none ignition prediction in transformer language models. The metric fits a four-parameter sigmoid to per-layer linear probe accur...

📖 Read original article


3. Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models ​

Author: Dayu Wang, Jiaye Yang, Weikang Li, Jiahui Liang, Yang Li, Deguo Xia, Jizhou Huang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.05168v1 Announce Type: new Abstract: Large language models often fail on reasoning tasks despite possessing the capability to solve them. We argue that many such failures arise from localized reasoning bugs in intermediate steps rather than from global incompetence. We show that these bug...

📖 Read original article


4. From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction ​

Author: Esra Zihni, Katryna Cisek, Hamzah Ziadeh, Hendrik Knoche, Robert Mikulik, John D. Kelleher
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05203v1 Announce Type: new Abstract: Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by the misalignment of model explanations with clinicians' reasoning. Motivated by a clinician user stu...

📖 Read original article


5. SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse ​

Author: Jialuo Chen, Minghe Wang, Lingqi Jiang, Jianan Ma, Xinhao Deng, Xiaohu Du, Ruixiao Lin, Yunhao Feng, Linkang Du, Jingyi Wang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05204v1 Announce Type: new Abstract: LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. As skills become marketplace artifacts, auditing their reuse is no ...

📖 Read original article


6. Abstract Event Causal Rules: Induction and Application ​

Author: Ziwei Zheng, Peiqiong Chen, Bang Wang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05205v1 Announce Type: new Abstract: Event-centric intelligent analytical systems heavily depend on explicit causal event knowledge for risk early warning, decision-making support and narrative comprehension. Nevertheless, existing instance-level causal pairs suffer severe generalization ...

📖 Read original article


7. Otter: A Time-Aware, History-Conditioned Human Chess AI ​

Author: Tarun Kumar S
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05206v1 Announce Type: new Abstract: Otter is a 15.3M-parameter human chess AI that predicts human move selection by modeling play as a time-aware, sequential process rather than treating each position in isolation. It combines two conditioning signals: (1) a move history encoder that con...

📖 Read original article


8. SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents ​

Author: Zhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answers. Diagnosing such ...

📖 Read original article


9. PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads ​

Author: Ao Fu, Yi Zhou
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.SD

arXiv:2608.05218v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth motion is often over-smoothed and may violate hard articulatory constraints such as bilabial closures, producing the ...

📖 Read original article


10. When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents ​

Author: Junzhuo Liu, Weiwei Li, Jun Ling, Peng Wang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05219v1 Announce Type: new Abstract: Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only references, such as successful trajectories. In interact...

📖 Read original article


11. Small Foundation Models of Human Cognition and Behaviour ​

Author: Nick Oh, Fernand Gobet
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2608.05224v1 Announce Type: new Abstract: Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train...

📖 Read original article


12. Project2Task: Graph-Guided Project-Level Planning for Autonomous Research ​

Author: Huirui Xu, Runtao Xu, Shuo Ren, Jiajun Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05225v1 Announce Type: new Abstract: Research agents can increasingly search literature, propose hypotheses, generate code, run experiments, and draft manuscripts from a single topic. However, a research project is not merely a larger task: it is a long-horizon agenda that must be advance...

📖 Read original article


13. TriQua: Reconciling Granularity and Context in Factuality Evaluation ​

Author: Jin Liu, Steffen Thoma, Achim Rettinger
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.05228v1 Announce Type: new Abstract: The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.e., one sentence conveying one unit of information, often omit essential context, while broader statements lack the granularity needed for...

📖 Read original article


14. Coherence-Oriented Dream Scene Visualisation ​

Author: Azra A\c{c}{\i}l, Simon Colton
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.05233v1 Announce Type: new Abstract: Dreams can be emotionally intense but difficult to communicate. We describe the Dream Scene Visualiser (DSV) system which turns written dream descriptions into a temporal sequence of four panel images visualising the dream. This starts with a large lan...

📖 Read original article


15. Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforcement Learning ​

Author: Muyang Ye, Tian Lan, Feihu Jiang, Yongshi Ye, Wuyunsiqin, Bin Zhu, Qianghuai Jia, Zhao Xu, Weihua Luo, Ye Wang, Jinyang Zhang, Longyue Wang, Lingfeng Bao
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05245v1 Announce Type: new Abstract: Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path toward self-evolution in expert domains. Existing self-evolving skill methods construct skills internally from th...

📖 Read original article


16. LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs ​

Author: Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang, Jiang Feng, Min Yang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities. To...

📖 Read original article


17. WorldClaw: Agentic 3D Open-World Generation at Scale ​

Author: Chunchao Guo, Jinpeng Li, Yang Li, Zilong Huang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.05248v1 Announce Type: new Abstract: Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present ...

📖 Read original article


18. Posture and Sustainment Optimization Under Adversarial Uncertainty ​

Author: Amelie Norris, Alyssa Lee, Natan Vidra, Spurthi Setty
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05256v1 Announce Type: new Abstract: Pre-commitment posture, the assignment of military assets to theater locations before conflict scenarios resolve, is a critical and formally unsolved problem in joint operational planning. Current practice relies on greedy heuristics that maximize valu...

📖 Read original article


19. OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality ​

Author: Yidian Chen, Yingzi Gu, Natan Vidra, Spurthi Setty, Sharon Zheng
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05263v1 Announce Type: new Abstract: Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipeline failed, where a cascade began, or which routing decision caused the breakdown. OrchestraBench eval...

📖 Read original article


20. Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks ​

Author: Nathan S Johnson, Ian Abshire
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci, cs.LG

arXiv:2608.05266v1 Announce Type: new Abstract: Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and synchrotron beamlines. Research into agentic control of physical infrastructure is nascent and there are...

📖 Read original article


21. CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction ​

Author: Jose A. Bird
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05359v1 Announce Type: new Abstract: CASCADE is an agentic framework that predicts downstream transcriptional effects of gene perturbation from precomputed ARACNe regulatory networks, exposed via MCP. Prior work validates such tools by checking whether predicted genes are known cancer gen...

📖 Read original article


22. Counterfactual Analysis via Large Language Models ​

Author: Zonghao Yang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, q-fin.GN

arXiv:2608.05367v1 Announce Type: new Abstract: Counterfactual analysis aims to predict potential outcomes under hypothetical scenarios, offering valuable insights for decision-making. This paper investigates the application of large language models (LLMs), specifically the GPT-3.5 model, for counte...

📖 Read original article


23. DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data ​

Author: Ruilin Wang, Bo-Hong Wang, Elizabeth Kourbatski, Jun Bai, Hegang Chen, Ziyang Song, Gilles Boire, Marie Hudson, Yue Li
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA

arXiv:2608.05375v1 Announce Type: new Abstract: Clinical machine learning (ML) has the potential to support high-stakes medical decision-making, but reliable deployment is often constrained by scarce, heterogeneous, and temporal complexity. Developing effective ML pipelines for such data remains tim...

📖 Read original article


24. C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models ​

Author: Swapnanil Mukherjee, Agyeya Negi, Tanuja Ganu, Ponnurangam Kumaraguru
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05381v1 Announce Type: new Abstract: Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning. We introduce C$^3$PO, a benchmark of 3,404 samples s...

📖 Read original article


25. Adaptive Arena-based Contestable Argumentative Network-of-Experts for Open-Ended Care Plan Coordination ​

Author: Truong Thanh Hung Nguyen, Hoang-Loc Cao, Phuc Ho, Phuc Truong Loc Nguyen, Ren'e Richard, Hung Cao
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.05391v1 Announce Type: new Abstract: Care plan coordination demands synthesizing heterogeneous clinical, functional, and psychosocial information across multiple professional disciplines, where monolithic LLM pipelines cannot perform in a transparent or safe manner. We present CANOE (Cont...

📖 Read original article


26. Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index ​

Author: Benjamin Barlog, Hudson Craig, Zedong Peng
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05411v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learning, effective help depends not only on correctness, but also on whether a response matches the learn...

📖 Read original article


27. Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation ​

Author: Mackenzie Jorgensen, Jo Reilly, Alex Sutherland, Miri Zilka
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05418v1 Announce Type: new Abstract: AI tools are being increasingly adopted in policing in the UK and worldwide. Racial bias is a known and well-documented risk, yet representatives of affected communities are rarely included in decisions about AI adoption. We present results from a mixe...

📖 Read original article


28. SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications ​

Author: Yixuan Wang, Licheng Luo, Yu Fu, Kaidi Xu, Yue Dong, Mingyu Cai
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05439v1 Announce Type: new Abstract: Translating natural language instructions into machine-interpretable formal specifications enables robots and autonomous systems to plan, reason, and formally verify their behavior. However, existing translation models typically generate a specificatio...

📖 Read original article


29. Stochasticity Is Not the Hard Part: Reduction and Complexity in Instructional Sequencing over Prerequisite DAGs ​

Author: Zonglin Han (Department of Computer Science, University of California, Davis), Yichen Chen (Department of Computer Science, University of California, Davis), Jiawen Jiang (International Digital Economy College, Minjiang University), Tongan Shi (School of Computer Science and Artificial Intelligence, Liaoning Normal University), Kristian A. Stevens (Department of Computer Science, University of California, Davis)
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.DS

arXiv:2608.05455v1 Announce Type: new Abstract: When a student must learn concepts connected by prerequisite dependencies, when does the order of instruction matter, and what does it cost to find the best one? We study instructional sequencing as a stochastic shortest-path problem in which attemptin...

📖 Read original article


30. Recursive Synthesis for Long-Horizon Terminal Tasks ​

Author: Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05466v1 Announce Type: new Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consistent...

📖 Read original article


31. Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability ​

Author: Ahmed Hassoon, Mark Dredze
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ML

arXiv:2608.05490v1 Announce Type: new Abstract: Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When such an analysis turns out to be wrong, someone must determine which operation caused it. A recent ap...

📖 Read original article


32. EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents ​

Author: Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation is part of the task i...

📖 Read original article


33. Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging ​

Author: Yu Gu, Zhi Zheng, Yunpeng Ba, Xialiang Tong, Mingxuan Yuan, Zhenkun Wang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05541v1 Announce Type: new Abstract: Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning. However, directly applying ES to billion-parameter LLMs is highly ineffective. In such high-dimensional para...

📖 Read original article


34. SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution ​

Author: Zhi Han, Chenxi Zeng, Liuhaichen Yang, Zihan Guo, Ming Zhou, Yang Li
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05573v1 Announce Type: new Abstract: LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the ...

📖 Read original article


35. StepReflect: Structured UI Transition Reflection for Mobile GUI Agents ​

Author: Linqiang Guo (Peter), Wei Liu (Peter), Li Gu (Peter), Yang Wang (Peter), Tse-Hsun (Peter), Chen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05587v1 Announce Type: new Abstract: Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended multimodal reasoning after each action, which is costly and poorly matched to the structured nature of GUI state...

📖 Read original article


36. Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows ​

Author: Nimisha Karnatak, Max Van Kleek, Nigel Shadbolt
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.05602v1 Announce Type: new Abstract: Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what condition...

📖 Read original article


37. Measuring and Detecting Harmful AI Sycophancy ​

Author: Bohan Jiang, Dawei Li, Yasin Silva, Huan Liu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.05624v1 Announce Type: new Abstract: Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This paper focuses on one harmful sycophancy: preference-induced stance reversal sycophancy (PSRS), where a...

📖 Read original article


38. SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation ​

Author: Yuru Feng, Yaoqi Chen, Beidi Zhao, Qianxi Zhang, Xinjiang Wang, Jianan Lu, Zhirui Wang, Shusen Xu, Zengzhong Li, Qi Chen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05628v1 Announce Type: new Abstract: Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus require autonomous, on-demand skill evolution at test time, constrained by lim...

📖 Read original article


39. Bayesian Expected Uncertainty Reduction (B-EUR) Model: A Computational Account of What Makes Design Options Worth Trying ​

Author: Shimon Honda, Takuma Miyaguchi, Koji Koizumi, Takanori Sano, Tristan Briard, Hideyoshi Yanagisawa
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05642v1 Announce Type: new Abstract: This paper proposes the Bayesian Expected Uncertainty Reduction (B-EUR) model, which formalizes the value of trying a candidate design action as its expected reduction of epistemic uncertainty about action--outcome relations. The model addresses one pa...

📖 Read original article


40. Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning ​

Author: Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.05643v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity. Verifier-base...

📖 Read original article


41. A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decomposition ​

Author: Jiaheng Chen, Jiaxing Li, Tinghe Zhang, Chaopeng Guo
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2608.05673v1 Announce Type: new Abstract: Trajectory prediction has shifted toward structured formulations with explicit social modeling. However, existing methods inadequately distinguish the functional roles of social influence in trajectory planning. Observing that agents typically form mot...

📖 Read original article


42. Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Baseline, and a Dual-Head Model ​

Author: Gospel Bassey, Vincent Fakiyesi
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05685v1 Announce Type: new Abstract: Most public benchmarks for machine-condition monitoring come from test rigs, where faults are induced on purpose and every event is known. Real production fields rarely offer that. They give you sensor histories with no fault log attached, which is exa...

📖 Read original article


43. DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model ​

Author: Wenhao Lin, Chenyu Yu, Xingwei Lin, Sicong Cao, Xiang Chen, Lei Xue, Le Yu, Letian Sha, Chunming Wu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CR

arXiv:2608.05695v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services. Recent runtime guardrails mitigate ...

📖 Read original article


44. Shaping Human-AI Interactions to Provide Improvement Pathways and Balance Competing Objectives ​

Author: Keziah Naggita
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05710v1 Announce Type: new Abstract: When an AI system is deployed, the individuals who use and or are evaluated by it form beliefs about how the system operates and use those beliefs to strategically present their preferences, behaviors, or attributes. The system then responds with feedb...

📖 Read original article


45. RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation ​

Author: Shuhao Yan, Changhao He, Xi Peng, Peng Hu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05714v1 Announce Type: new Abstract: Text-to-CAD generation translates natural-language design intent into editable and executable parametric computer-aided design (CAD) codes, reducing the expertise and effort required for manual modeling. Existing methods incorporate fixed, externally s...

📖 Read original article


46. BlockPython: A Process-Aware Agent-Supported Platform for the Transition from Block-Based to Python Programming ​

Author: Jesse Yusuf Chan (Zexi Chen), Haoming Wang, Mingwei Xu, Xianlong Xu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05716v1 Announce Type: new Abstract: The transition from block-based to text-based programming requires learners to convert visible program structures into abstract textual expressions, which may create a cognitive gap between understanding computational concepts and expressing them in Py...

📖 Read original article


47. Unified Agent: Managing Interactions across Devices ​

Author: Xinshuang Liu, Runfa Blark Li, Shaoxiu Wei, Xin Lin, Truong Nguyen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.HC

arXiv:2608.05729v1 Announce Type: new Abstract: As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user's devices over time. Yet existing agent systems still fall short in this scenario. This is because observations are scattered across devices and mo...

📖 Read original article


48. Subliminal Learning is Non-Semantic Distillation ​

Author: Ethan Hadley, Eren Gultepe
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05734v1 Announce Type: new Abstract: Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models. It allows the transfer of a bias or behavior from a teacher model to a student by distilling from seemingly unrelated or random synthetic data from the...

📖 Read original article


49. When Do Prompt-Side Agent Playbooks Transfer? Accuracy, Cost, and Runtime Shift in Agent Deployment ​

Author: Weihong Lin, Lin Sun, Xiangzheng Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05778v1 Announce Type: new Abstract: Prompt-side playbooks can improve tool-using language agents without retraining, but their portability beyond the source setting is unclear. We study frozen playbook transfer under a shared distill--validate--transfer protocol. On ALFWorld, transfer is...

📖 Read original article


50. Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay ​

Author: Nossa Iyamu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05784v1 Announce Type: new Abstract: Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent's memory today records what the user said, not what the user did. We compile passively captured screen activity into agent memory w...

📖 Read original article


51. ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution ​

Author: Jiacheng Wei, Zhaoxin Fan, Xin Wen, Yuqin Lan, Dongrun Li, Wenjun Wu, Faguo Wu, Xiao Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2608.05790v1 Announce Type: new Abstract: General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockchain environments. On-chain execution is stateful, adversarial, and economically irreversible, exposi...

📖 Read original article


52. When Agentic AI Meets Integrated Sensing and Communication ​

Author: Kai Li, Conggai Li, Sarah Ali Siddiqui, Syed Sohail Ahmed, Xin Yuan, Shenghong Li, Wei Ni
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05792v1 Announce Type: new Abstract: Agentic artificial intelligence (AI) is transforming Integrated Sensing and Communication (ISAC) from a function-oriented physical-layer technology into a goal-driven, closed-loop intelligent system, a paradigm we term AISAC. Existing work on learning-...

📖 Read original article


53. When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents ​

Author: Linfang Shang, Ming Xu, Yiding Sun, Tianle Xia, Lingxiang Hu, Lan Xu, Ning Zheng
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of improving it. We formali...

📖 Read original article


54. Cautious Context Steering for Language Model Personalization ​

Author: Gihoon Kim, Jeyoung Lee, Suhan Woo, Sekwon Oh, Minsu Jeon, Hyounsoo Han, Euntai Kim
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05813v1 Announce Type: new Abstract: Personalizing language models (LMs) to individual user preferences is essential for aligning responses with diverse goals and backgrounds. Existing methods typically train a separate adapter for each user or learn a reward model whose scores depend on ...

📖 Read original article


55. ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion ​

Author: Jiafan Li, Mengxue Yang, Jiaqi Zhu, Liang Chang, Ying Li, Hongan Wang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05833v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and i...

📖 Read original article


56. Runtime Observability for Heterogeneous Attention Memory ​

Author: Fanzhe Wei, Li Liu, Ziyang Wang, Chenyu Wang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05863v1 Announce Type: new Abstract: Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each fails differently under compression. We give a runtime observability contract that c...

📖 Read original article


57. Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs? ​

Author: Yuyang Dai, Xueqing Peng, Yuxia Wang, Preslav Nakov, Zhuohan Xie
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05864v1 Announce Type: new Abstract: Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings. This makes it unclear whether models can perceive visual business evide...

📖 Read original article


58. Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation Tasks ​

Author: Jonathon Dilworth, Pedro Giesteira Cotovio, David Herron, Paul Cripps, Nigel Dewdney, Catia Pesquita, Ernesto Jim'enez-Ruiz
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05867v1 Announce Type: new Abstract: The use of ontologies and knowledge graphs is becoming increasingly widespread in the defence and national security domain. Numerous ontologies have been developed through initiatives led by academia, industry, and government. Achieving interoperabilit...

📖 Read original article


59. Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding ​

Author: Soojin Yoon, Dongha Lee
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.05876v1 Announce Type: new Abstract: User requests serve as research specifications for deep research agents, shaping what evidence to seek and how to synthesize it. In personalized deep research, these specifications must additionally reflect user goals, constraints, preferences, and eva...

📖 Read original article


60. AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents ​

Author: Weikai Xu, Yunren Feng, Haoxiang Lei, Kun Huang, Yuxuan Liu, Kang Zhao, Xiaolin Hu, Shuo Shang, Bo An
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.05891v1 Announce Type: new Abstract: Mobile GUI agents can operate apps through pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction policies. However, real trajectories are difficult to obtain for sensitive app...

📖 Read original article


61. ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation ​

Author: Akanta Das, Tasinul Islam Ahon, Ahmed Mahir Sultan Rumi, Md Mahbubur Rahman, Tausif Amim Shadly, Tanzima Hashem
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05893v1 Announce Type: new Abstract: Electrocardiography (ECG) is one of the most widely used non-invasive tools for diagnosing cardiovascular disease, but transforming multi-lead ECG recordings into reliable clinical reports remains challenging. Automating ECG report generation could red...

📖 Read original article


62. GSBF: Gaussian Splatting for Environment-Aware Beamforming ​

Author: Yijie Bian, Wei Guo, Zixin Wang, Shenghui Song, Jun Zhang, Khaled B. Letaief
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.IT, math.IT

arXiv:2608.05896v1 Announce Type: new Abstract: Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design normally requires accurate instantaneous channel state information (CSI) and iterative optimization, which incur subst...

📖 Read original article


63. CourseGraph: Finding overlaps and differences in Computer Science courses across universities ​

Author: Arthur Nijdam, Paul Stankovski Wagner, Sara Ramezanian
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2608.05910v1 Announce Type: new Abstract: Student mobility programs such as Erasmus+ enable students to take courses at other universities, broadening their academic and cultural horizons. However, this flexibility also leads to a practical challenge: ensuring that students do not take courses...

📖 Read original article


64. GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models ​

Author: Shuai Wang, Yaxin Feng, Xuekun Jiang, Shihan Tian, Ningyu Yan, Xing Shen, Chaoyang Lyu, Hui Wang, Yunsong Zhou, Hanqing Wang, Jiangmiao Pang, Yang Xiang, Xing Gao, Chunhua Shen, Weinan Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.RO

arXiv:2608.05948v1 Announce Type: new Abstract: Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit simulators of future states and interactions. However, existing evaluations of physical fidelity are ...

📖 Read original article


65. VLMs for Videogame Data Annotation ​

Author: Katrin Schmid, Iuri Frosio
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2608.05949v1 Announce Type: new Abstract: Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications. Their adoption in video games is on the other hand limited by the extreme variability of the s...

📖 Read original article


66. Training a Conditioned Video Game Agent on a VLM Annotated Dataset ​

Author: Katrin Schmid, Iuri Frosio
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2608.05954v1 Announce Type: new Abstract: Reinforcement Learning (RL) is a powerful but far from easy-to-use technique for policy learning. In the specific case of video games, access to the game engine is required to get rewards for training (e.g. to collect rewards from the environment). Fur...

📖 Read original article


67. Stability of Ranking-dependent Pair-wise Comparison Patterns in the Analytic Hierarchy Process ​

Author: Vitaliy Tsyganok, Sergii Kadenko, Oleh Andriichuk
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05958v1 Announce Type: new Abstract: The paper addresses several ranking-dependent decision support methods. Ordinal information on compared objects can be used to improve the quality of expert data during estimation and help reduce the number of comparisons that the experts need to perfo...

📖 Read original article


68. Temporal Bridges for Spatial Resolution: Enhancing Climate Data Super-Resolution with Bidirectional Alignment ​

Author: Yichen Zhang, Yixiong Xiao, Congxi Xiao, Jingbo Zhou
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05981v1 Announce Type: new Abstract: High-resolution climate data is crucial for meteorological predictions and for informing decision support across diverse domains. However, the acquisition of such high-resolution climate information is often prohibitively costly, necessitating the deve...

📖 Read original article


69. AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning ​

Author: Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao, Jinyang Wu, Jie Wu, Zhengzhou Cai, Yueqing Sun, Ziang Ye, Linji Hao, Qi Gu, Xunliang Cai, Yongliang Shen, Yujiu Yang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05987v1 Announce Type: new Abstract: Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privil...

📖 Read original article


70. OPERA: Operator-residual feedback for reliable autonomous optical experiments with language-model agents ​

Author: Ning Xu, Xiang Zheng, Fuqiang Zhong, Huadong Wang, Xiaolong Wu, Zhiyuan Liu, Hui Ning
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, physics.app-ph, physics.optics

arXiv:2608.05990v1 Announce Type: new Abstract: Autonomous agents choose actions using scores that may not reflect experimental success. We developed OPERA, an operator-residual framework for optical experiments. It represents experimental actions as optical operators and evaluates their outcomes us...

📖 Read original article


71. Hybrid Machine Learning Framework for Herd-Level Cattle Growth Pattern and Weight Gain Forecasting in Grazing-Based Production Systems ​

Author: Muhammad Riaz Hasib Hossain, Rafiqul Islam, Shawn R. McGrath, Md Zahidul Islam, David W. Lamb
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06001v1 Announce Type: new Abstract: Commercial grazing systems yield irregular livestock observations, which challenge cattle growth forecasting. This study developed a hybrid machine learning framework for herd level cattle weight forecasting using automated sensing observations collect...

📖 Read original article


72. HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards ​

Author: Zhuowen Liu, Bohan Cui, YinShang Guo, Yuting Wang, Hao Li
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06012v1 Announce Type: new Abstract: Search-agent rewards mix answer quality, citation grounding, tool cost, and anti-hacking terms; a high score therefore need not imply that cited evidence was retrieved, and added penalties can cancel. We introduce HERALD, an offline audit that applies ...

📖 Read original article


73. From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models ​

Author: Jiale Han, Xiang Li, Jing Qian, Wenyuan Gu, Pin Gao, Ye Luo, Hongyuan Zha, Dacheng Tao, Benyou Wang, Lin William Cong
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.06020v1 Announce Type: new Abstract: Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through which their interactions produce...

📖 Read original article


74. Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis ​

Author: Rafa{\l} Buler (Gda'nsk University of Technology), Jakub Buler (Gda'nsk University of Technology), Maciej Bobowicz (Medical University of Gda'nsk), Micha{\l} Grochowski (Gda'nsk University of Technology)
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG, eess.IV

arXiv:2608.06037v1 Announce Type: new Abstract: Relational inductive biases are essential for capturing structural dependencies among data. This study investigates a dual-level relational framework for image classification, bridging the gap between implicit representation learning and explicit struc...

📖 Read original article


75. When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories ​

Author: Xiaoqing Wu, Xingyu Fan, Feifei Li, Wenhui Que
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06057v1 Announce Type: new Abstract: Tool-calling agents infer task state from accumulated dialogue and tool traces. In persistent interactions, however, historical traces may remain structurally valid and semantically plausible after they cease to be authoritative for the current request...

📖 Read original article


76. Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference ​

Author: Yifan Lyu, Xinran Li, Jiaqi Qiao, Xiujuan Xu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06085v1 Announce Type: new Abstract: Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label's uniform, record-indepe...

📖 Read original article


77. Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents ​

Author: Yuanhong Jiang, Jingjie Zou, Zhenghong Lin, Xusheng Yu, Qiqi Huang, Shuai Jia, Shijie Dai
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06108v1 Announce Type: new Abstract: Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering...

📖 Read original article


78. ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment ​

Author: Abdulkadir K"ul\c{c}e, Alihan Esen, Ca\u{g}la Fikir, Berke Kurt, Kuzey Arar, G"okhan Ercan, Faik Boray Tek
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.06110v1 Announce Type: new Abstract: This paper presents ECHO (Enhanced Care & Health Observer), a locally-deployable conversational health assistant for long-term chronic care management. ECHO integrates three complementary software modules developed under shared supervision as a unifie...

📖 Read original article


79. From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems ​

Author: Manideep Dhar, Ritwik Singh, Sharat Chandra Kumar Manikonda
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA

arXiv:2608.06112v1 Announce Type: new Abstract: Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc., yet most deployments remain isolated point solutions locked inside departmental silos, resulting in duplicated effort, hidden risks, and unrealized enterprise ...

📖 Read original article


80. Mind the Gaps: Mixture-of-Minds for Human Simulation ​

Author: Pranav Dahiya
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06115v1 Announce Type: new Abstract: Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual. Large language model simulators inherit this gap. They recover a population'...

📖 Read original article


81. Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts ​

Author: Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio, Holger Boche
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.06123v1 Announce Type: new Abstract: Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Pol...

📖 Read original article


82. Contextual Information Policy Optimization for Search Agents ​

Author: Xingyu Guo, Wei Chen, Linlin Yang, Baochang Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06128v1 Announce Type: new Abstract: Search agents extend large language models beyond static parametric memory by enabling them to acquire and use ex ternal evidence during multi-step reasoning. For knowledge intensive tasks involving complex or evolving information, their reliability de...

📖 Read original article


83. FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows ​

Author: Bo Deng (Beihang University, Qwen DianJin Team, Alibaba Cloud Computing), Kang Zhou (Qwen DianJin Team, Alibaba Cloud Computing), Lifan Guo (Qwen DianJin Team, Alibaba Cloud Computing), Chongyang Tao (Beihang University), Xuanren Chen (Beihang University), Chenggang Xie (Beihang University), Renzhao Liang (Beihang University), Feng Chen (Qwen DianJin Team, Alibaba Cloud Computing), Chi Zhang (Qwen DianJin Team, Alibaba Cloud Computing)
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06144v1 Announce Type: new Abstract: Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution benchmarks do not jointly cover professional workflows, open-ended deliverables, and multi-aspect eva...

📖 Read original article


84. PaDoc: Layout-Grounded Parallel Decoding for Document Parsing ​

Author: Hao Yu, Jiabo Zhan, Kang Liu, Linnan Zhao, Dongxu Yue, Rui Chen, Jinglin Wang, Chong Sun, Chen Li, Jing Lyu, Chun Yuan
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06146v1 Announce Type: new Abstract: End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, wher...

📖 Read original article


85. CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query? ​

Author: Zijie Wang, Chen Zhong, Wei He
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.06150v1 Announce Type: new Abstract: Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories. Open-Vocabulary Change Detection (OVCD) addresses this need. However, existing methods often entangle temporal perception, semantic discrimi...

📖 Read original article


86. iARCS: Iterative Agentic RL for Controllable 3D Scene Generation ​

Author: Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06161v1 Announce Type: new Abstract: Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits th...

📖 Read original article


87. Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI ​

Author: Modhurita Mitra, Jan-Willem Versteeg, Maarten D. Schermer, Shiva Nadi Najafabadi, Marie L. De Bruin, Lourens T. Bloem
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.06167v1 Announce Type: new Abstract: We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard. The schema, serving...

📖 Read original article


88. MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration ​

Author: Jia Xiong, Runkai Li, Chenxu Niu, Guangyuan Gao, Changwen Xing, Yifan Zhang, Xinlai Wan, Jieran Cui, Chen Bai, Yusheng Hua, Ying Wang, Ming Ling, Xi Wang, Tao Xie
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06183v1 Announce Type: new Abstract: Microarchitecture design space exploration suffers from expansive search spaces and expensive PPA evaluation, leaving only a small simulation budget for design decision-making. Existing methods perform blind search without considering microarchitectura...

📖 Read original article


89. Comparative Approaches to Agent Retrieval over Large Skill Libraries ​

Author: Indivara Kolluru, Nathan Sportsman
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06196v1 Announce Type: new Abstract: Agents backed by large skill libraries must decide which skills to load and in what order. Loading the entire library into context is expensive and provides no structure for autonomous sequencing. We study two systems for this problem over a corpus of ...

📖 Read original article


90. EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning ​

Author: Zishan Xu, Zhiyuan Yao, Yuxin Chen, Yifu Guo, Zhengxi Lu, Yuquan Lu, Jinyang Huang, Yan Xu, Yasheng Wang, Weinan Zhang, Xingshan Zeng, Weiwen Liu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06197v1 Announce Type: new Abstract: Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We...

📖 Read original article


91. TS-RAG: Retrieval Augmented Generation for Time Series Forecasting ​

Author: Yixiong Xiao, Congxi Xiao, Jingbo Zhou
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.06223v1 Announce Type: new Abstract: While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the application of retrieval-augmented generation (RAG) in this domain remains limited. Since RAG has proven effecti...

📖 Read original article


92. DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models ​

Author: ZhiYan Hou, Xinyu Tang, Hongyan An, Jianjin Zhang, Weizhen Wang, Yunyun Han, Gengsheng Li, Xiangzhao Hao, Haiyun Guo, Wenbin Hu, Jinqiao Wang, Yafeng Deng
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06243v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distilla...

📖 Read original article


93. Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints ​

Author: Omid Bazgir, Md Nasir, Jacob Hoffman, Yang Yang, Manu Agrawal, Anusua Trivedi, Vinay Rao Dandin, Chris Gibbons, Christine Swisher
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.DB, cs.LG

arXiv:2608.06265v1 Announce Type: new Abstract: Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve su...

📖 Read original article


94. The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images ​

Author: Zhiheng Wang, Bo Peng, Lai Wei, Chaochao Lu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06270v1 Announce Type: new Abstract: The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cos...

📖 Read original article


95. QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction ​

Author: Mutasim Fuad Sarker, Adiba Rahman Namira, Wafa Binte Alam, Md Adnan Arefeen, Mahzabeen Emu, Sumaiya Tabassum Nimi
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.ET

arXiv:2608.06294v1 Announce Type: new Abstract: Cardiac arrest remains one of the most lethal conditions encountered in intensive care units. Despite the growing availability of electronic health record data, existing mortality prediction studies in this population largely depend on static summaries...

📖 Read original article


96. Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors ​

Author: Arya Labroo, Mengjie Qian, Kate Knill
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06300v1 Announce Type: new Abstract: Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker at...

📖 Read original article


97. HarnessOpt-Bench: Evaluating LLMs at Harness Optimization ​

Author: Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S. Kalmath, Veronica Chatrath, Yuan Xue
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.06301v1 Announce Type: new Abstract: As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness o...

📖 Read original article


98. Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations ​

Author: Sagar Tamang, Ayush Vyas, Tabarakul Hazarika
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR

arXiv:2608.06305v1 Announce Type: new Abstract: Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents -- financial statements, audit rep...

📖 Read original article


99. TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories ​

Author: Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu, Xiaozhi Wang, Lei Hou, Bin Xu, Juanzi Li
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06346v1 Announce Type: new Abstract: LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsib...

📖 Read original article


100. Challenges in Evaluating Explanation Methods for Static and Evolving Data ​

Author: Jerzy Stefanowski
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06351v1 Announce Type: new Abstract: This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through the DetoxAI image recognition system for bias detection and concept unlearning. Then, an example of ...

📖 Read original article


101. The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping ​

Author: Sarvesh Baskar, Zikui Cai, Shayan Shabihi, Anirudh Satheesh, Muhammad R. Islam, Udari Madhushani Sehwag, Tom Goldstein, Furong Huang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06361v1 Announce Type: new Abstract: Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only th...

📖 Read original article


102. Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering ​

Author: Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi, Daniel Kang, Roy Ka-Wei Lee, Koustuv Saha, Christian Poellabauer, Christopher Lee, Sajeev Singh, Piyum Zonooz, Navin Kumar, Zeeshan Ahmed, Priyadarshini Kachroo
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.06366v1 Announce Type: new Abstract: Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults...

📖 Read original article


103. HoloCount: A Holistic Visual Counting Benchmark for MLLMs ​

Author: Jinhong Deng, Limeng Qiao, Guanglu Wan
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.06420v1 Announce Type: cross Abstract: Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. While Multimodal Large Language Models (MLLMs) have achieved remarkable success in qualitative scene...

📖 Read original article


104. Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision Support ​

Author: Gary Simethy, Daniel Ortiz Arroyo, Petar Durdevic
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05151v1 Announce Type: cross Abstract: Wastewater operators need answers grounded in how their plant's variables interact and how fast effects propagate, not in generic pretraining text, when asking causal questions such as "why is N2O rising?" or "what happens if I cut aeration by 20%?"....

📖 Read original article


105. Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models ​

Author: Hao Ai
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05152v1 Announce Type: cross Abstract: Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior may help deepen our understanding and guide model optimization. In this study, we introduce a framew...

📖 Read original article


106. Universal Pathologies, Conditional Consequences: A Triple-Robustness Analysis of RAG for Multi-Hop Traceability ​

Author: Meftun Akarsu, Burak Ozdemir
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05153v1 Announce Type: cross Abstract: GraphRAG underperforms vector RAG on citation precision in many reports, but where and why have remained corpus-bound. We present a triple-robustness analysis that holds the retrieval architecture fixed and varies three orthogonal axes embedder (loca...

📖 Read original article


107. Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for Political News Evaluation ​

Author: Maryam Fooladi, Federico Bottino
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05155v1 Announce Type: cross Abstract: Traditional sentiment analysis (SA) models, while effective for polarity classification, provide limited insight into the rhetorical, ideological, and framing dimensions of political discourse -- dimensions that are central to research in the social ...

📖 Read original article


108. Large Language Models Threaten Double-blind Review ​

Author: Bulambo Mwendelwa Gloire, Prasenjit Mitra
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05157v1 Announce Type: cross Abstract: Double blind peer review serves as the scientific community primary defense against status and affiliation bias. Its effectiveness rests on the assumption that anonymized manuscripts convey scientific merit without revealing their authors. While auth...

📖 Read original article


109. A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper ​

Author: Ali Shendabadi, Parnia Izadirad, Mostafa Salehi
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.SD

arXiv:2608.05165v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we study the use of Whisper for Persian SER with a particular focus on representation dimensionality reduction and lan...

📖 Read original article


110. DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory Graph ​

Author: Zhihao Xiao, Mengting Li, Xintao Wang, Linfeng Li, Limin Shui, Mengqi Ji, Borui Cai
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05170v1 Announce Type: cross Abstract: Role-playing agents (RPAs) have emerged as a key application of large language models, enabling immersive and high-fidelity character simulation. Accurate role-playing of established characters requires not only stylistic imitation but also temporall...

📖 Read original article


111. Beyond Information Retrieval: Generative AI as an Epistemic Arbiter to Enhance Collaborative Problem-Solving ​

Author: Jiaxin Zou, Xiaoming Zhai, Chunlei Gao
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.05171v1 Announce Type: cross Abstract: Generative AI (GAI) creates new opportunities for collaborative problem-solving (CPS), yet its role in shaping student interaction remains unclear. To address this gap, we conducted a six-week quasi-experimental study with 201 fifth-grade students in...

📖 Read original article


112. Estimating time spent on work tasks ​

Author: Stephane Hatgis-Kessell, Tom'as Aguirre, Alexander Wan, Rishi Bommasani
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.05172v1 Announce Type: cross Abstract: The task-based framework in economics models occupations as bundles of tasks. It is the standard lens for understanding how technology affects work: a new technology changes the cost or time each task requires and these task-level effects aggregate t...

📖 Read original article


113. The Closing Window: How Governments Could Lose Their Ability to Restrain Advanced AI ​

Author: Peter Barnett
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.05173v1 Announce Type: cross Abstract: As AI capabilities advance, AI systems will pose greater risks to national security and potentially humanity as a whole. Governments may eventually conclude that these risks warrant restraining AI development. This motivates the question: will govern...

📖 Read original article


114. Challenges for Musical Education in the Age of AI and Digital Transformation ​

Author: Jean-Pierre Briot
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG

arXiv:2608.05176v1 Announce Type: cross Abstract: Music education has never been a static discipline. Each major technological shift has forced educators and institutions to reconsider what they teach, how they teach it, and why. We now stand at what may be the most consequential of such turning poi...

📖 Read original article


115. Who Gets Access? Global Region and Academic Status Bias in AI-Generated Academic Gatekeeping Scenarios ​

Author: Nouar AlDahoul, Hezerul Abdul Karim, Myles Joshua Toledo Tan
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.05178v1 Announce Type: cross Abstract: Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, datasets, or professional materials such as curriculum vitae (CV) must be shared selectively. We introdu...

📖 Read original article


116. Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap ​

Author: Tianyu Ding, Aditya Nannapaneni, Bingfan Liu, Ling Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.05179v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review. End-to-end AI scientist systems can now produce...

📖 Read original article


117. Automatic Detection of Deaths from Social Networking Sites ​

Author: Nuhu Ibrahim, Riza Batista-Navarro
Published: 8/7/2026, 4:00:00 AM
Categories: cs.SI, cs.AI

arXiv:2608.05183v1 Announce Type: cross Abstract: This dissertation analysed and discussed the differences in linguistic characteristics between pre-mortem and post-mortem social media content, and reported machine learning (ML) classifiers that achieved high performance in automatically detecting d...

📖 Read original article


118. Position: It's Time to Optimize LLMs for Self-Consistency ​

Author: Itamar Pres, Belinda Z. Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, Jacob Andreas
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05188v1 Announce Type: cross Abstract: Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and produce confident but...

📖 Read original article


119. Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents ​

Author: Zhenpeng Li
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.05199v1 Announce Type: cross Abstract: Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific technique. Split conformal prediction gives each stage finite-sample coverage, but deployment requires a trajectory...

📖 Read original article


120. ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design, Evaluation, and an OpenClaw Case Study ​

Author: Siyuan Li, Peng Shu, Churan Yu, Peilong Wang, Ruidong Zhang, Bowen Guo, Xinliang Li, Ruiyu Yan, Arif Hassan Zidan, Yi Pan, Wei Ruan, Lifeng Chen, Junhao Chen, Zhaojun Ding, Yiwei Li, Zhengliang Liu, Haixing Dai, Lin Zhao, Yu Bao, Xiang Li, Wei Zhang, Tianming Liu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.05201v1 Announce Type: cross Abstract: Autonomous AI agent platforms differ substantially in architecture, security, tool integration, execution, autonomy, and deployment, yet the field lacks a common classification scheme for comparing these design choices. We propose ASTELD, an operatio...

📖 Read original article


121. Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation ​

Author: Ye Leng, Junjie Chu, Yiting Qu, Mingjie Li, Yun Shen, Yang Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CR

arXiv:2608.05210v1 Announce Type: cross Abstract: Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by the notorious Nazi propaganda picture book \emph{Der Giftpilz}. Recently, frontier text-to-image (T2...

📖 Read original article


122. Quality Diversity for Reliable Data Driven Time-Use Optimization ​

Author: Aneta Neumann, Ty Stanford, Dorothea Dumuid, Frank Neumann
Published: 8/7/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, cs.NE

arXiv:2608.05230v1 Announce Type: cross Abstract: The daily allocation of the finite 24-hour time budget is strongly associated with physical, mental, and cognitive health. While predictive models can estimate the relationship between time-use compositions and health outcomes such as body mass index...

📖 Read original article


123. In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion ​

Author: Lingxiao Yang, Liu Liu, Moran Li, Han Feng, Wenjian Cao, Jiangning Zhang, Ye Shi
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.05237v1 Announce Type: cross Abstract: Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame. However, these clean frames leak excessive local details, which causes the model to take sh...

📖 Read original article


124. One Qubit Can Beat One Bit: Quantum Advantage for Post-Training Quantization ​

Author: Yuma Ichikawa, Moeto Mishima
Published: 8/7/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG

arXiv:2608.05240v1 Announce Type: cross Abstract: One-bit post-training quantization represents each weight using only its sign, requiring all deployment contexts to share the same binary weight matrix even when their activation statistics favor different sign patterns. We study this shared-sign con...

📖 Read original article


125. Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models ​

Author: Duong Bach, Hai Nguyen Hong, Cuong Do
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.05243v1 Announce Type: cross Abstract: Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prior and interpret this as evidence that the style representation is independent of class information. We show tha...

📖 Read original article


126. PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis ​

Author: Xiaomin He, Dongling Xiao, Jiahao Xie, Ruiqi Lu, Qianle Wang, Zhongbin Guo, Wanxuan Sun
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.05249v1 Announce Type: cross Abstract: Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering one self-contained question. We study this gap through \textbf{rubric com...

📖 Read original article


127. An Emerging Retail Portfolio Management Application: Personalized, Tax-Aware Reinforcement Learning with Natural Language Goals ​

Author: Ramin Pishehvar
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2608.05255v1 Announce Type: cross Abstract: Retail investors lack access to the kind of personalized, tax-aware portfolio management that institutional clients take for granted -- existing robo-advisors use static, rule-based allocation, and institutional-grade systems require account minimums...

📖 Read original article


128. IMMENSE: Inductive Multi-perspective User Classification in Social Networks ​

Author: Francesco Benedetti, Antonio Pellicani, Gianvito Pio, Michelangelo Ceci
Published: 8/7/2026, 4:00:00 AM
Categories: cs.SI, cs.AI

arXiv:2608.05259v1 Announce Type: cross Abstract: Online social networks increasingly expose people to users who propagate discriminatory, hateful, and violent content. Young users, in particular, are vulnerable to exposure to such content, which can have harmful psychological and social repercussio...

📖 Read original article


129. Failing Gracefully: Mitigating Impact of Inevitable Robot Failures ​

Author: Duc M. Nguyen, Saad A. Ghani, Andrew Marshall, Allison Andreyev, Gregory J. Stein, Xuesu Xiao
Published: 8/7/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.HC, cs.LG

arXiv:2608.05313v1 Announce Type: cross Abstract: Service robots operate in household environments shared with humans, pets, and everyday objects, where they are highly susceptible to failures such as software crashes, hardware degradation, or unpredictable interactions. While roboticists strive to ...

📖 Read original article


130. Hierarchical Server Architecture for Agentic Science ​

Author: Vanessa Sochat, Daniel Milroy
Published: 8/7/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2608.05332v1 Announce Type: cross Abstract: Agentic science is transforming the landscape of computational work, extending to scientific pipelines and workload managers. The workloads require specialized hardware within and across institutions. If assessing workload needs against environments ...

📖 Read original article


131. Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN Networks ​

Author: Marcos Carvalho, Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci, Daniel F. Macedo
Published: 8/7/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2608.05340v1 Announce Type: cross Abstract: Time-Sensitive Networking (TSN) and Mobile Edge Computing (MEC) hold strong potential for enabling ultra-reliable low-latency communication for time-sensitive applications, such as eXtended Reality (XR). However, the widespread adoption of XR introdu...

📖 Read original article


132. Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application ​

Author: Marcos Carvalho, Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci, Daniel F. Macedo
Published: 8/7/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2608.05346v1 Announce Type: cross Abstract: Time-sensitive networking (TSN) is increasingly integrated into mobile edge computing (MEC) to support applications with stringent latency requirements, such as extended reality (XR). However, existing TSN scheduling solutions predominantly rely on s...

📖 Read original article


133. Perturbation Sensitivity at Convergence: A Simple Signal for Identifying Spuriously Correlated Samples ​

Author: Nilesh Kumar
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.05419v1 Announce Type: cross Abstract: Models trained by empirical risk minimization on data containing spurious correlations achieve high average accuracy while failing on subpopulations where the correlation does not hold. Existing methods for identifying the affected samples without gr...

📖 Read original article


134. Why the Third Axis Is Freedom ​

Author: Michael Timothy Bennett
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.05423v1 Announce Type: cross Abstract: In generative training, a model produces an output and is penalised for its difference from an example. With one output per comparison, a model that produces one common answer can outperform a model retaining a broader repertoire. Explorative Modelin...

📖 Read original article


135. The ethics of artificial intelligence in the life sciences: Universality, cultural diversity and an architecture of care ​

Author: Jean-Pierre Changeux, Gustavo Deco, Morten L. Kringelbach
Published: 8/7/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI

arXiv:2608.05436v1 Announce Type: cross Abstract: The life sciences and health research have started to benefit from artificial intelligence, which raises ethical concerns that are real but, we argue, not special. Any science should be governed by values that rest on how the human brain is built and...

📖 Read original article


136. Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers ​

Author: Zhen Zhang, Amr Alanwar
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.05472v1 Announce Type: cross Abstract: Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping aggregated values to outputs is the same for every input set. We study the consequences of this asy...

📖 Read original article


137. Learning Context-Free Grammars for Grammar-Constrained Decoding via Declarative Agentic Programming with Guarantees ​

Author: Kevin Cheang, Geoff Hulette, Rahul Kumar, Felipe R. Monteiro, Federico Mora, Robin Salkeld, Lin Tan, Serdar Tasiran
Published: 8/7/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.CL, cs.SE

arXiv:2608.05493v1 Announce Type: cross Abstract: Language models (LMs) are increasingly used to interact with external services via programs written in domain-specific languages (DSLs). Unfortunately, since DSLs are often low-resource and esoteric, LMs frequently produce syntactically invalid progr...

📖 Read original article


138. APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning ​

Author: Sadegh Jafari, Mohiuddin Bilwal, Fan Zhou, Brian Gelder, Ali Jannesari
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.05499v1 Announce Type: cross Abstract: Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge devices. Pruning and quantization address this, but rely on manual, expert choices and on algorithms that are ...

📖 Read original article


139. Vibe Compiler: A Research-Logic Synthesis Tool That Runs without Prompt Engineering -Toward Enhancing Metacognition for Sustaining Agency in the Age of Generative AI- ​

Author: Riichiro Mizoguchi, Tomoki Aburatani, Kento Koike, Machi Shimmei
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2608.05545v1 Announce Type: cross Abstract: Generative AI used as a capable servant has greatly accelerated intellectual work, but it also risks eroding human epistemic agency by encouraging uncritical acceptance of AI-generated reasoning. This creates a need for mechanisms that preserve human...

📖 Read original article


140. Turing's Frist Imitation Game: Design Concepts and a Human-Approximates-Machine Reading ​

Author: Sharon Temtsin, Christoph Bartneck
Published: 8/7/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.05558v1 Announce Type: cross Abstract: This paper examines Turing's 1948 report, "Intelligent Machinery", as an important conceptual source for the later imitation games. Its first contribution is to identify and integrate the design concepts underlying the 1948 chess-based imitation game...

📖 Read original article


141. When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems ​

Author: Jialuo Chen, Lingqi Jiang, Xinhao Deng, Xiaohu Du, Jianan Ma, Yunhao Feng, Yuqi Qing, Zhihao Yuan, Linkang Du, Jingyi Wang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.05563v1 Announce Type: cross Abstract: Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We introduce PoisonedEvolution, a trajectory-poisoning attack on this promotion process. Our skill-visib...

📖 Read original article


142. The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions ​

Author: Hadi Hosseini, Samarth Khanna, Leona Pierce
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG

arXiv:2608.05583v1 Announce Type: cross Abstract: As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce medical resources often hinge on judgments of responsibility, particularly when patients' own ...

📖 Read original article


143. Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations ​

Author: He Jiang, Jingtian Yan, Yulun Zhang, Yimin Tang, Tanishq Duhan, Rishi Veerapaneni, Guillaume Sartoretti, Jiaoyang Li
Published: 8/7/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.MA

arXiv:2608.05588v1 Announce Type: cross Abstract: Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that continuously receive new goals upon reaching their current ones. While many learning-based planners have been proposed for LMAPF, most rely on...

📖 Read original article


144. LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction ​

Author: Yingqing Guo, Hui Yuan, Zijian He, Mengdi Wang, Zheng Ding
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2608.05600v1 Announce Type: cross Abstract: Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learning requires stochastic rollouts for policy exploration and optimization. Existing GRPO methods for ...

📖 Read original article


145. SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries ​

Author: Xingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu, Xin Yuan, Liming Zhu, Wenjie Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05604v1 Announce Type: cross Abstract: Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at inference time. As skill libraries grow, a central challenge is to expose the smallest sufficient executable context...

📖 Read original article


146. GAUGE: Granularity-Adaptive Counterfactual Gating of Evidence for Incomplete Multimodal Classification ​

Author: Yunping Shi, En Yu, Kairui Guo, Jie Lu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.05608v1 Announce Type: cross Abstract: Multimodal classification typically assumes all modalities are available, yet real-world inputs are often incomplete. Imputation and dynamic fusion can mitigate such incompleteness, but existing methods operate at a coarse modality level and thus can...

📖 Read original article


147. Relay, Don't Route: Adaptive Population Handoff for Cost-Efficient LLM-Driven Evolution ​

Author: Sichun Luo, Yi Huang, Guanzhi Deng, Haibo Wang, Haochen Luo, Lei Li, Zefa Hu, Junlan Feng, Qi Liu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.NE

arXiv:2608.05651v1 Announce Type: cross Abstract: Large language model (LLM)-driven evolution has shown promise for program search and algorithm discovery, but relying on strong models throughout long evolutionary runs is costly. A natural alternative is to combine cheap and strong models under a fi...

📖 Read original article


148. Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics ​

Author: Jessica Y. Bo, Paula Akemi Aoyagui, Shalaleh Rismani, Dipto Das, Syed Ishtiaque Ahmed, Ashton Anderson
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2608.05656v1 Announce Type: cross Abstract: Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks favor technical methods, such as model benchmarks and LLM simulations, often sidelining empirical rese...

📖 Read original article


149. F$^2$Agent: Financial Fusion of Agentic Intelligence for Multimodal Trading ​

Author: Changshuo Liu, Yanzheng Jin, Shangfeng Cai, Peng Fang, Xiaokui Xiao, Beng Chin Ooi
Published: 8/7/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.MM

arXiv:2608.05668v1 Announce Type: cross Abstract: With increasingly diverse and heterogeneous information sources, effectively leveraging multimodal data is becoming pivotal for high-quality financial trading. Although recent advancements in Large Language Model (LLM)-based agents have enabled the i...

📖 Read original article


150. SafeDivertor: Faithful Divertor Heat Flux Reconstruction from Macroscopic Plasma State Signals via Time-Frequency Prior Exploitation ​

Author: Hao Si, Zehua Chen, Qingquan Yang, Xiao Wang, Dengdi Sun, Wanli Lyu, Gaoting Chen, Guosheng Xu, Hang Su, Jin Tang, Jun Zhu
Published: 8/7/2026, 4:00:00 AM
Categories: physics.plasm-ph, cs.AI, cs.CV

arXiv:2608.05669v1 Announce Type: cross Abstract: Divertor heat-flux analysis is essential for understanding plasma-wall interactions and protecting plasma-facing components in magnetic-confinement fusion devices, while conventional infrared-based inversion is usually performed after discharge and r...

📖 Read original article


151. DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation ​

Author: Jiaxuan Li, Qing Xu, Xiangjian He, Yue Li, Daokun Zhang, Fiseha B. Tesema, Rong Qu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM

arXiv:2608.05683v1 Announce Type: cross Abstract: Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered by uncertainty in both modalities under real-world clinical conditions. Existing vision-language segmentation m...

📖 Read original article


152. Nonvisual Classification of Ground-Condition by Artificial Proprioception in an Amoeba-Inspired Autonomous Walking Robot ​

Author: Hyoto Yamaguchi, Zenji Yatabe, Seiya Kasai
Published: 8/7/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG, cs.SY, eess.SY

arXiv:2608.05684v1 Announce Type: cross Abstract: Nonvisual classification of ground condition based on a multimodal sensing approach was investigated for an amoeba-inspired autonomous walking robot. To classify ground condition without image sensing and processing, we implemented artificial proprio...

📖 Read original article


153. Answer First, Reason Later: Commitment Order in Diffusion LLMs ​

Author: Jewon Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park, Hwiyeong Lee, Taesup Kim
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05687v1 Announce Type: cross Abstract: Masked diffusion language models (dLLMs) can commit tokens in any order -- a freedom marketed as their core advantage over autoregressive decoding. We show that on reasoning tasks this freedom is instead the axis of failure. Logging every commitment ...

📖 Read original article


154. Spectral Aliasing Pretext: A novel task for Self-Supervised fault diagnosis in rotating machinery ​

Author: Victor Gialis, Maxime Metz, David Esteve, Abdenour Soualhi
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.05705v1 Announce Type: cross Abstract: Deep learning is a new way for machinery fault diagnosis but requires extensive labeled data, a scarce resource in industrial settings. We propose Spectral Aliasing Pretext (SAP), a self-supervised learning method that pretrains models on unlabeled v...

📖 Read original article


155. Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots ​

Author: S. M . Bhagya P. Samarakoon, M. A. Viraj J. Muthugala, W. K. R. Sachinthana, Mohan Rajesh Elara
Published: 8/7/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.05715v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding. This tight coupling between perception and instruct...

📖 Read original article


156. ABC: Numerical Data Collection under Local Differential Privacy without Prior Knowledge ​

Author: Incheol Baek, Hyungbin Kim, Yon Dohn Chung
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2608.05737v1 Announce Type: cross Abstract: Local Differential Privacy (LDP) provides strong privacy guarantees for collecting numerical data. A fundamental challenge, however, is that existing LDP mechanisms require a predefined data domain, which is often unknown in practice. This lack of pr...

📖 Read original article


157. Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration ​

Author: Hongrui Bao, Yubing Ren, Yanan Cao, Jinhan You, Fang Fang, Shi Wang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05741v1 Announce Type: cross Abstract: Large language models (LLMs) can generate fluent and convincing text at scale, creating growing risks for misinformation dissemination, educational misuse, and platform governance. These concerns make robust detection of machine-generated text increa...

📖 Read original article


158. Multivariate Time Series Forecasting needs Cross Variable Loss ​

Author: Kuiye Ding, Yifan Hu, Hanchen Wang, Hao Xue
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.05742v1 Announce Type: cross Abstract: Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics. While existing studies mainly focus on cross-variable dependencies in historical observations, dependencies among f...

📖 Read original article


159. UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on ​

Author: Yushe Cao, Shikun Feng, Fei Shen, Haikuo Peng, Jianqiang Xia, Yiheng Zhu, Dianxi Shi, Chun Yu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.05745v1 Announce Type: cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for human parsin...

📖 Read original article


160. HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection ​

Author: Aohua Li, Jin Kuang, Yubing Lu, Pingping Liu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.05771v1 Announce Type: cross Abstract: Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance often degrades markedly when generalizing to unseen infrared domains. Existing methods primarily improve detection ...

📖 Read original article


161. GROM: Gradient-Free Rapid One-Shot Machine Unlearning ​

Author: Pawe{\l} Batorski, Przemys{\l}aw Spurek, Paul Swoboda
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.05783v1 Announce Type: cross Abstract: Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). Current state-of-the-art approaches primarily rely on iterative, training-time unlearning via fine-tuning. Howeve...

📖 Read original article


162. Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation ​

Author: Tirth Bhatt, Naren Kumar S, Mayank Singh
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05785v1 Announce Type: cross Abstract: Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (TCFM), a m...

📖 Read original article


163. A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems ​

Author: Zihan Xu, Haolin Tian, Hai Jiang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2608.05791v1 Announce Type: cross Abstract: Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and their execution strategies directly affect system accuracy, latency, and computational cost. Parallel ex...

📖 Read original article


164. Hierarchical Latent Prediction for Language Models ​

Author: Chang Shi, Tim Pearce, Manan Tomar, Siddhartha Sen, John Langford
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05806v1 Announce Type: cross Abstract: While standard Next-Token Prediction (NTP) lays the foundation of language model pre- training, its teacher-forced training paradigm may not be optimal for long-horizon reasoning and planning. Recent works such as Multi-Token Prediction (MTP) and Nex...

📖 Read original article


165. MameLoshnLM: Yiddish Language Model and Evaluation Benchmark ​

Author: Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05850v1 Announce Type: cross Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progre...

📖 Read original article


166. Evidential Rule Learning for Interpretable Classification with Abstention ​

Author: Javier Fumanal-Idocin, Javier Andreu-Perez
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.05859v1 Announce Type: cross Abstract: Interpretable classification often requires more than accurate predictions for real-life deployment: models should be transparent about the evidence behind their decisions and abstain when they cannot decide reliably. We introduce Fast Evidential Rul...

📖 Read original article


167. MACRO: Markov Chain Routing of Transformer Layers ​

Author: Pawe{\l} Batorski, Abtin Pourhadi, Akylgali Aitaza, Przemys{\l}aw Spurek, Paul Swoboda
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.05872v1 Announce Type: cross Abstract: Standard Large Language Models (LLMs) execute layers sequentially. Dynamic layer routing, i.e. search for a different execution path through layers involving layer repetitions, skips and other moves, can improve performance. Existing routing approach...

📖 Read original article


168. D-CLOT: Double Closed Loop Optimal Transport for Unsupervised Action Segmentation ​

Author: Elena Bueno-Benito, Mariella Dimiccoli
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.05877v1 Announce Type: cross Abstract: Optimal transport (OT) has emerged as an effective framework for unsupervised action segmentation. Yet, in existing OT-based methods, the latent action prototypes that define the OT costs are not re-estimated from the refined frame geometry. Instead,...

📖 Read original article


169. Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation ​

Author: Benjamin Connor, Anna Jurek-Loughrey, Lu Bai, Muhammad Fahim
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.05880v1 Announce Type: cross Abstract: Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data. While numerous explainability techniques exist, they...

📖 Read original article


170. CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents ​

Author: Wuya Chen, Yihao yang, Yang Cao, Yue Lin
Published: 8/7/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.05886v1 Announce Type: cross Abstract: Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent averages 23 rounds and 631K...

📖 Read original article


171. The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025 ​

Author: Przemys{\l}aw Czuma (Polish Association for Artificial Intelligence in Medicine)
Published: 8/7/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.CL, cs.CY

arXiv:2608.05889v1 Announce Type: cross Abstract: Large language models (LLMs) can leave small stylistic traces in text written with their help. The most discussed is the em-dash (U+2014), especially the unspaced form word---word, which is normal in typeset English prose but unusual in U.S. press wr...

📖 Read original article


172. BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks ​

Author: Guanqiao Qu, Shuo Chen, Qian Chen, Kin K. Leung, Xianhao Chen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2608.05926v1 Announce Type: cross Abstract: Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decoding (AD) generates output tokens sequentially, resu...

📖 Read original article


173. Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models ​

Author: Maulik Chevli, Johannes Brandt, Rickmer Braren, Daniel Rueckert, Philip M"uller
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.05960v1 Announce Type: cross Abstract: Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume. 3D CT foundation models could assist this process by providing generalizable representations of anatomy and pathology. To evaluate the...

📖 Read original article


174. SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation ​

Author: Changyuan Wang, Chubin Zhang, Zhenyu Wu, Runhao Li, Angyuan Ma, Ke Chao, Yinan Liang, Xiuwei Xu, Ziwei Wang, Yansong Tang, Jiwen Lu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.05970v1 Announce Type: cross Abstract: Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. However, their potential remains fundamentally constrained by the scarcit...

📖 Read original article


175. TRACE: Learned Proprioceptive Odometry for Legged Robots under Unreliable Contact Conditions ​

Author: Taehyeon Kong, Woojin Kim, Jemin Hwangbo
Published: 8/7/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.05975v1 Announce Type: cross Abstract: In this paper, we present TRACE (Tokenized Robust Attention for Contact-Aware Estimation), an end-to-end learned proprioceptive odometry estimator for legged robots under unreliable contact conditions. The proposed estimator directly predicts relativ...

📖 Read original article


176. ProDVI: Programmatic Dynamics Priors for Value Network Initialization ​

Author: Xinwei Liu, Junyuan Liang, Jianting Zhang, Wuhui Chen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06015v1 Announce Type: cross Abstract: Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch, forcing them to acquire task-relevant knowledge through online interaction. Existing approaches obta...

📖 Read original article


177. FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India ​

Author: Aman Dalmia, Sanskriti Midha, Jigar Doshi
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2608.06027v1 Announce Type: cross Abstract: In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write. Reaching them requires a spoken conversation. Today that work falls to frontline health workers who enroll benefi...

📖 Read original article


178. Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case ​

Author: Shilin Hu, Jingyi Xu, Dimitris Samaras, Hieu Le
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.06075v1 Announce Type: cross Abstract: Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural question: do they reduce the need for classic, physics-informed low-level vision? We study this thr...

📖 Read original article


179. Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping ​

Author: Vaishnav Vaidheeswaran, Dilith Jayakody, Biruk Ambaw, Jaswanth Kumar, Md Mahbub Alam, Gabriel Spadon
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA

arXiv:2608.06105v1 Announce Type: cross Abstract: Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deployment requires reward models that are interpretable and robust to changing environments. Inverse reinforcement l...

📖 Read original article


180. Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers ​

Author: Haris Riaz, Hyungji Kim, Mihai Surdeanu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.06111v1 Announce Type: cross Abstract: Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}. We introduce \textbf{S}yntax-\textbf{i}nformed \textbf{P}ositional \textbf{E}mbeddings (\textbf{SiPE}), which learns ...

📖 Read original article


181. Is Self-Pretraining really useful to improve diagnosis in medical Time Series? ​

Author: Omar Coser, Antonio Orvieto, Paolo Soda, Loredana Zollo
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06122v1 Announce Type: cross Abstract: Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series. Our ob...

📖 Read original article


182. Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture ​

Author: Leo Sambrook, Sampo Sovio
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2608.06130v1 Announce Type: cross Abstract: AI agents performing cryptographic operations (signing Git commits, authenticating API calls, issuing certificates) currently store private keys in software-accessible locations: plaintext files, environment variables, or container memory. Any proces...

📖 Read original article


183. Reducing belief in conspiracy theories as they unfold using large language models ​

Author: Thomas H. Costello, Nathaniel Rabb, Michael Nicholas Stagnaro, Gordon Pennycook, David Rand
Published: 8/7/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.06151v1 Announce Type: cross Abstract: The emergence of conspiracy theories in the wake of major events is a significant societal challenge. Here we test whether conversational dialogues with a large language model (LLM) can reduce belief in immediately unfolding conspiracies. In experime...

📖 Read original article


184. Learning Globally Reusable Skills for Coding Agents ​

Author: Chen Yang, Jiashuo Tian, Ziqi Wang, Xinyin Liu, Meiru Ye, Junjie Chen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.06153v1 Announce Type: cross Abstract: Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing approaches typically treat skill evolution as a sequence of local updates, overlooking relationships among ski...

📖 Read original article


185. Visual Grounding in Zero-Shot Vision-Language Control ​

Author: J. de Curt`o, Dayani Plasencia, Diego S'anchez, I. de Zarz`a
Published: 8/7/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2608.06154v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable sc...

📖 Read original article


186. Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset ​

Author: Eoin Cummins, Zhongyi Huang, Alexandre D'Hooge, Zhuoro Mo, Yaolong Ju
Published: 8/7/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM

arXiv:2608.06165v1 Announce Type: cross Abstract: Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first presents the new SheetSage-A2S Dataset, which includes 61 hours of audio with \texttt{**kern} score...

📖 Read original article


187. What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) ​

Author: Ro Encarnaci'on, Tina Behzad, Emma Lurie, Dana'e Metaxa
Published: 8/7/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform a single run per prompt, and ...

📖 Read original article


188. Continual Learning in Transition ​

Author: Zhiyan Hou, Dan Zhang, Tao Feng, Liyuan Wang, Wei Li, Xiangzhao Hao, Hongyan An, Junfeng Fang, Haokai Ma, Zhaohui Xu, Haiyun Guo, Jinqiao Wang, Tat-Seng Chua
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06216v1 Announce Type: cross Abstract: Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architectural designs, and weight adaptation. However, emerging paradigms are r...

📖 Read original article


189. From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks ​

Author: Christo Kurisummoottil Thomas, Omar Hashash, Walid Saad
Published: 8/7/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.IT, cs.SY, eess.SY, math.IT

arXiv:2608.06227v1 Announce Type: cross Abstract: Despite advances in artificial intelligence (AI) across multiple sectors, today's AI tools, including deep learning and generative AI, still fail when embedded into physical systems, such as robots and vehicles operating under real-world physical law...

📖 Read original article


190. Depth-Guided Video Object Counting in Crowded Scenes ​

Author: Yuanjing Xu, Xinyan Liu, Weidong Chen, Zixuan Zou, Linhao Zhang, Zhuangzhe Meng, Antoni B. Chan, Weigang Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.06236v1 Announce Type: cross Abstract: Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Existing methods rely on RGB information, limiting their discriminative ab...

📖 Read original article


191. PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation ​

Author: Elad Yoshai, Natan T. Shaked
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.06240v1 Announce Type: cross Abstract: Unpaired image-to-image translation must decide, per image, what to change and what to preserve without paired supervision. Many diffusion-based unpaired translators control preservation through a single global noise or guidance value applied across ...

📖 Read original article


192. Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model ​

Author: Saad Ahmed, Md Khalid Syfullaha
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.06252v1 Announce Type: cross Abstract: Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on personal devices could widen access to education and services. Existing systems use controlled-setting datasets withou...

📖 Read original article


193. BaKron: Efficient Quantization with Kronecker-Factored Hessians ​

Author: Johann Birnick, Rayan Saab
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06291v1 Announce Type: cross Abstract: We accelerate a family of algorithms for neural network quantization whose geometry is informed by any Kronecker-factored approximation of the Hessian. GPTQ-style adaptive rounding typically uses one-sided information derived from input activations. ...

📖 Read original article


194. Does FLAIR super-resolution erase or hallucinate small white-matter lesions? ​

Author: Zahra Khodakarami, Yue Li, Pulkit Khandelwal, John Detre, Sandhitsu Das, Christopher Brown, David Wolk, Paul Yushkevich
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.06311v1 Announce Type: cross Abstract: White matter hyperintensities (WMH), bright regions on Fluid-attenuated Inversion Recovery (FLAIR) scans are associated with cerebrovascular pathology and neurodegeneration. FLAIR is usually acquired with thick slices in clinical settings, giving it ...

📖 Read original article


195. Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents ​

Author: Noam Koren, Roy Bar-Haim, Abigail Goldsteen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.06329v1 Announce Type: cross Abstract: Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading...

📖 Read original article


196. Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data ​

Author: Donna Hooshmand, Shubham Shahi, Cameron Barrie, Abhratanu Dutta, Marko Sterbentz, Harper Pack, Kristian J. Hammond
Published: 8/7/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2608.06331v1 Announce Type: cross Abstract: From natural-language query interfaces to automated report generation, data analysis tools need a description of the data: the real-world entities it contains, which columns function as measures or identifiers, and how tables connect into units of an...

📖 Read original article


197. Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents ​

Author: Praphul Chandra, Sujit Gujar, Ganesh Ghalme
Published: 8/7/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.MA

arXiv:2608.06353v1 Announce Type: cross Abstract: We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle that governance should control an AI agent through resource allocation so as to make authorization sel...

📖 Read original article


198. AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games ​

Author: Boning Li, Yu Chen, Longbo Huang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.CL, cs.LG, cs.MA

arXiv:2608.06362v1 Announce Type: cross Abstract: Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the r...

📖 Read original article


199. An Optimal Agnostic PAC Algorithm ​

Author: Markus Engelund Mathiasen, Jian Qian, Nikita Zhivotovskiy
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DS, math.ST, stat.TH

arXiv:2608.06363v1 Announce Type: cross Abstract: Let $H\subseteq{-1,+1}^X$ be a class of finite VC dimension $d\ge1$. Writing $L$ for the binary risk and $L^*=\min_{h\in H}L(h)$, we construct a learner achieving the statistically optimal risk bound: from an i.i.d.\ sample of size $n$, for every $...

📖 Read original article


200. Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria ​

Author: George Grispos, Sajda Qureshi
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.06364v1 Announce Type: cross Abstract: The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including fraud and reduced user control over digital technologies, raising concerns about digital sovereignty. This research examines how Artif...

📖 Read original article


201. Learning When to Trust via Selective Context Preference Optimization ​

Author: Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.06377v1 Announce Type: cross Abstract: Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context ...

📖 Read original article


202. Analogy as Nonparametric Bayesian Inference over Relational Systems ​

Author: Ruairidh M. Battleday, Thomas L. Griffiths
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2006.04156v2 Announce Type: replace Abstract: Our inferences in the real world are rarely na"ive - we acquire experiences through our lifetime that can help us more quickly understand the structure of something new. A fundamental question in cognitive science is how we make such generalizatio...

📖 Read original article


203. Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet ​

Author: James Xu Zhao, Bryan Hooi, See-Kiong Ng
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2509.06861v3 Announce Type: replace Abstract: Test-time scaling increases inference-time computation through longer reasoning chains and has shown strong performance gains across many domains. However, frontier models still suffer from factuality hallucinations, raising the question of whether...

📖 Read original article


204. AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations ​

Author: Berdymyrat Ovezmyradov
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent trends in AI benchmarking is performance of Large Language Models (LLMs) over longer time horizons. Wh...

📖 Read original article


205. Symbol Grounding in Neuro-Symbolic AI: A Gentle Introduction to Reasoning Shortcuts ​

Author: Emanuele Marconato, Samuele Bortolotti, Emile van Krieken, Paolo Morettin, Elena Umili, Antonio Vergari, Efthymia Tsamoura, Andrea Passerini, Stefano Teso
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2510.14538v3 Announce Type: replace Abstract: Neuro-symbolic (NeSy) AI aims to develop deep neural networks whose predictions comply with prior knowledge encoding, e.g. safety or structural constraints. As such, it represents one of the most promising avenues for reliable and trustworthy AI. T...

📖 Read original article


206. BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics ​

Author: Dionizije Fa, Marko Culjak, Bruno Pandza, Mateo Cupic
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.21800v4 Announce Type: replace Abstract: We introduce BioAgent Bench, an evaluation suite designed for measuring the performance and robustness of AI agents in common bioinformatics tasks. The suite consists of manually curated end-to-end tasks (e.g., RNA-seq, variant calling, metagenomic...

📖 Read original article


207. When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation ​

Author: Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, Vil'em Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek \v{S}uppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo, Eliya Habba, Usman Gohar, Siddhesh Pawar, Robert Scholz, Arjun Subramonian, Jingwei Ni, Mykel Kochenderfer, Sanmi Koyejo, Mrinmaya Sachan, Stella Biderman, Zeerak Talat, Avijit Ghosh, Irene Solaiman
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2602.16763v4 Announce Type: replace Abstract: Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly "saturate", making it difficult to differentiate models and diminishing their long-term value. In th...

📖 Read original article


208. CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation ​

Author: Haodong Li, Chunmei Qing, Huanyu Zhang, Dongzhi Jiang, Yihang Zou, Hongbo Peng, Dingming Li, Yuhong Dai, ZePeng Lin, Juanxi Tian, Yi Zhou, Siqi Dai, Jingwei Wu, Pheng-Ann Heng
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.08652v2 Announce Type: replace Abstract: Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the integration of Chain-of-Thought (CoT) reasoning. However, existing CoT-based T2I methods largely rely on ab...

📖 Read original article


209. SimMOF: AI agent for Automated MOF Simulations ​

Author: Jaewoong Lee, Taeun Bae, Jihan Kim
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.29152v3 Announce Type: replace Abstract: Metal-organic frameworks (MOFs) offer a vast design space, and as such, computational simulations play a critical role in predicting their structural and physicochemical properties. However, MOF simulations remain difficult to access because reliab...

📖 Read original article


210. TRU: Targeted Reverse Update for Efficient Multimodal Recommendation Unlearning ​

Author: Zhanting Zhou, KaHou Tam, Ziqiang Zheng, Zeyu Ma, Yang Yang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.02183v3 Announce Type: replace Abstract: Multimodal recommendation systems (MRS) jointly model user-item interaction graphs and rich item content, but this tight coupling makes user data difficult to remove once learned. Approximate machine unlearning offers an efficient alternative to fu...

📖 Read original article


211. GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning ​

Author: Ornith Team, Xiaoya Li, Guoyin Wang, Songqiao Su, Chris Shum, Jiwei Li
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.02721v3 Announce Type: replace Abstract: Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms the best humans competitive programming: the most recent best result, Google's Gemini~3 Deep Think, attained ...

📖 Read original article


212. An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics ​

Author: Miri Liu, ChengXiang Zhai
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.DL

arXiv:2604.15145v2 Announce Type: replace Abstract: The rigorous evaluation of the novelty of a scientific paper is, even for human scientists, a challenging task. With the increasing interest in AI scientists, it is becoming more and more important that this task be automatable and reliable, lest a...

📖 Read original article


213. CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction ​

Author: Jianyou Wang, Youze Zheng, Longtian Bao, Hanyuan Zhang, Qirui Zheng, Yuhan Chen, Yang Zhang, Matthew Feng, Maxim Khan, Aditya K. Sehgal, Christopher D. Rosin, Ramamohan Paturi, Umber Dube, Leon Bergen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2604.16742v2 Announce Type: replace Abstract: Scientists have long sought to accurately predict outcomes of real-world events before they happen. Can AI systems do so more reliably? We study this question through clinical trial outcome prediction, a high-stakes open challenge even for domain e...

📖 Read original article


214. To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling ​

Author: Qinyuan Wu, Soumi Das, Mahsa Amani, Arijit Nag, Seungeon Lee, Krishna P. Gummadi, Abhilasha Ravichander, Muhammad Bilal Zafar
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.00737v3 Announce Type: replace Abstract: Agentic AI architectures augment LLMs with external tools, unlocking strong capabilities but potentially incurring substantial costs. Moreover, tool use is not always beneficial: redundant or low-utility calls can even harm task performance. Effect...

📖 Read original article


215. Behind EvoMap: Characterizing a Self-Evolving Agent-to-Agent Collaboration Network ​

Author: Qiming Ye, Peixian Zhang, Yupeng He, Zifan Peng, Gareth Tyson
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2605.25815v4 Announce Type: replace Abstract: Agent-to-Agent (A2A) networks enable autonomous AI agents to collaborate by sharing reusable problem-solving instructions. However, how these decentralized ecosystems operate in practice remains largely unexplored. We present the first large-scale ...

📖 Read original article


216. SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis ​

Author: Yucheng Yuan, Yuanfeng Ji, Zhongxiao Li, Ruijiang Li
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.24235v2 Announce Type: replace Abstract: Spatial proteomics enables single-cell-resolution characterization of protein expression within tissue architecture, playing a critical role in understanding tumor microenvironments and guiding precision medicine. However, current analysis workflow...

📖 Read original article


217. Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics ​

Author: Arshia Soltani Moakhar, Iman Gholami, Max Springer, Mahdi JafariRaviz, MohammadTaghi Hajiaghayi
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.31134v3 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle errors that evade human detection. Formal mathematical languages like Lean 4 offer mechanical proof checking, st...

📖 Read original article


218. Berkeley and Heiserman as an Unexhausted Architecture for Embodied Machine Intelligence ​

Author: Christopher A. Tucker
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.16465v2 Announce Type: replace Abstract: Edmund C. Berkeley is usually remembered as a mediator between symbolic logic and early computing, yet that standard description understates the scope of his work. This paper argues for a stronger reading: Berkeley should also be understood as an e...

📖 Read original article


219. TRW: TRACE-RealWorld---An Auditable Consistency Contract for World Models as Materialized Views ​

Author: Edward Y. Chang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.DB

arXiv:2607.21910v2 Announce Type: replace Abstract: World models let agents plan against predicted physical state, but that state drifts; re-observation is costly and delayed, and repair can fail. We present TRACE-RealWorld (TRW), to our knowledge the first commitment-level consistency contract for ...

📖 Read original article


220. Localized Anomaly Detection via Differentiable D-vine Copulas ​

Author: Nicholas Andrea Pearson, Francesca Zanello, Davide Russo, Luca Bortolussi, Francesca Cairoli
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.25020v2 Announce Type: replace Abstract: Vine copulas provide a flexible framework for modeling complex multivariate distributions through a hierarchical decomposition into bivariate pair-copulas. Fitting a D-vine requires selecting a copula family and parameter configuration for each pai...

📖 Read original article


221. Property-driven Causal Abstractions for Markov Decision Processes ​

Author: Jule Schmidt, Maximilian Weininger, Clemens Dubslaff, David Parker, Nils Jansen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2607.26787v2 Announce Type: replace Abstract: Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and their valuations. The exponential blowup in the number of states renders many reasoning tasks in M...

📖 Read original article


222. The Geometry of Flow-Matching Uncertainty: A Cost-free Uncertainty Proxy and Its Application in Flow-based VLA Failure Detection ​

Author: Ziyang Rao, Yiren Zhao, Weiyu Guo, Ben Fei, Yandong Guo, Hui Xiong
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.27933v3 Announce Type: replace Abstract: Flow matching (FM) has become a popular action head paradigm for modern embodied models. However, as a conditional generative model, it does not explicitly expose its inherent uncertainty, producing faulty action chunks even when it misinterprets t...

📖 Read original article


223. Shapes from Examples: Foundations of Shape Learning in Recursive SHACL ​

Author: Bente Gortworst, Cem Okulmus, Magdalena Ortiz, Anni-Yasmin Turhan
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2607.27934v2 Announce Type: replace Abstract: SHACL shapes enable data graph validation, making automatic shape learning essential for knowledge graph applications. We investigate the well-known fitting approach to this task: given sets P and N of positive and negative example nodes from an in...

📖 Read original article


224. OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models ​

Author: Qiushi Sun, Kanzhi Cheng, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu, Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV

arXiv:2607.28609v2 Announce Type: replace Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and r...

📖 Read original article


225. AISPA: User-Centric System Prompt Auditing for Large Language Model Applications ​

Author: Xiangning Lin, Shenzhe Zhu, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan, Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi, Robert Mahari, Tobin South, Dazza Greenwood, Zexue He, Rishi Bommasani, Sophia Kazinnik, Andreas Haupt, Samuele Marro, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY, cs.HC

arXiv:2607.28617v2 Announce Type: replace Abstract: System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trus...

📖 Read original article


226. H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases ​

Author: Shusen Zhang, Junyi Hu, Ye Feng, Ziteng Wang, Zhaoyuan Pan, Guosheng Dong, Xiaojun Yuan, Jiangshou Hong, Xiangzhi Wang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.00065v2 Announce Type: replace Abstract: Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constraints, and compositional concepts. However, existing representations lie at two extremes: single-vector retri...

📖 Read original article


227. DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimization ​

Author: Changquan Zhao, Yuxiang Sun, Ruihao Zhu, Cheng Hua, Yulian He
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.00641v2 Announce Type: replace Abstract: Bayesian optimization (BO) relies on a surrogate model and an acquisition function, yet the most suitable choices vary across tasks and optimization stages. Automated Bayesian optimization (AutoBO) addresses this variability by adapting BO componen...

📖 Read original article


228. When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning ​

Author: Andrew Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.02940v2 Announce Type: replace Abstract: A stable compression score can still select the worse model. In our dense study, a split-half reliable path-quadratic score predicted a 16.1% gain, while the selected endpoints were 6.0--7.7% worse than two controls. We ask what a compression stat...

📖 Read original article


229. Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study ​

Author: Qiming Bao, Sherry J. H. Feng, Kim Chester Eugenio, Meng Fon
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.03172v2 Announce Type: replace Abstract: Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes "Maria S.", not [NAME] -- so that clinical text stays fluent and downstream tools keep working. But this onl...

📖 Read original article


230. Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents ​

Author: Siqi Fan, Minghao Li, Xiaoqian Ma, Wenhui Tan, Xiusheng Huang, Juntong Wu, Liujie Zhang, Shuo Shang, Weihang Chen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.03327v2 Announce Type: replace Abstract: Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way the effect goes. Under one identical GUI-MCP harness on the OSWorld-MCP benchmark (309 tasks), the same MCP to...

📖 Read original article


231. EviGraph: Evidence-Guided Autonomous Research Agents ​

Author: Zhenjiang Ren, Ruiji Li, Xujing Zhang, Ziliang Pang, Shuo Ren, Jiajun Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.04738v2 Announce Type: replace Abstract: Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported claims and inconsistencies between research questions, experiments, results, and conclusions. We argue that ...

📖 Read original article


232. Path Planning of Cleaning Robot with Reinforcement Learning ​

Author: Woohyeon Moon, Bumgeun Park, Sarvar Hussain Nengroo, Taeyoung Kim, Dongsoo Har
Published: 8/7/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2208.08211v2 Announce Type: replace-cross Abstract: Recently, as the demand for cleaning robots has steadily increased, therefore household electricity consumption is also increasing. To solve this electricity consumption issue, the problem of efficient path planning for cleaning robot has bec...

📖 Read original article


233. Revisiting Black-Box Model Ownership Verification through Information Theory ​

Author: Aoting Hu, Yanzhi Chen, Renjie Xie, Xinwei Zhang, Wei Xu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2409.06130v2 Announce Type: replace-cross Abstract: Modern machine learning models require substantial computational resources and data to train, making them valuable intellectual property. Model watermarking has emerged as a practical solution for black-box ownership verification, but existin...

📖 Read original article


234. Explanations of Large Language Models Explain Language Representations in the Brain ​

Author: Maryam Rahimi, Mohammad Reza Daliri, Yadollah Yaghoobzadeh
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, q-bio.NC

arXiv:2502.14671v4 Announce Type: replace-cross Abstract: Large Language Model (LLM) representations are known to align with brain activity during language processing, but it remains unclear what drives this alignment. We test whether explainable AI (XAI) can help answer this: using attribution meth...

📖 Read original article


235. ASAT: Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection ​

Author: Daisuke Yamada, Harit Vishwakarma, Ramya Korlakai Vinayak
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2505.02299v2 Announce Type: replace-cross Abstract: Machine Learning (ML) models are trained on in-distribution (ID) data but often encounter out-of-distribution (OOD) inputs during deployment---posing serious risks in safety-critical domains. Recent works have focused on designing scoring fun...

📖 Read original article


Author: Xiaoya Li, Albert Wang, Guoyin Wang, Chris Shum, Jiwei Li
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.DB

arXiv:2508.02091v4 Announce Type: replace-cross Abstract: Approximate nearest-neighbor search (ANNS) algorithms have become increasingly critical for recent AI applications, particularly in retrieval-augmented generation (RAG) and agent-based LLM applications. In this paper, we present CRINN, a new ...

📖 Read original article


237. Uncertainty-aware Predict-Then-Optimize Framework for Equitable Post-Disaster Power Restoration ​

Author: Lin Jiang, Dahai Yu, Rongchao Xu, Tian Tang, Guang Wang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SI

arXiv:2508.04780v3 Announce Type: replace-cross Abstract: The increasing frequency of extreme weather events, such as hurricanes, highlights the urgent need for efficient and equitable power system restoration. Many electricity providers make restoration decisions primarily based on the volume of po...

📖 Read original article


238. Autonomous Learning From Success and Failure: Goal-Conditioned Supervised Learning with Negative Feedback ​

Author: Zeqiang Zhang, Fabian Wurzberger, Gerrit Schmid, Sebastian Gottwald, Daniel A. Braun
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2509.03206v2 Announce Type: replace-cross Abstract: Learning from reward functions and imitation learning of demonstrations are the two principal approaches for training autonomous systems that interact with an environment through action and observation. Both, however, require human specificat...

📖 Read original article


239. AegisShield: Democratizing Cyber Threat Modeling with Generative AI ​

Author: Matthew Grofsky
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2509.10482v2 Announce Type: replace-cross Abstract: The increasing sophistication of technology systems makes traditional threat modeling hard to scale, especially for small organizations with limited resources. This paper develops and evaluates AegisShield, a generative AI enhanced threat mod...

📖 Read original article


240. Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime ​

Author: Leonardo Defilippis, Yizhou Xu, Julius Girardin, Emanuele Troiani, Vittorio Erba, Lenka Zdeborov'a, Bruno Loureiro, Florent Krzakala
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.dis-nn, cs.AI, stat.ML

arXiv:2509.24882v3 Announce Type: replace-cross Abstract: Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models. In this work, we present a systematic analysis of scaling laws for quadratic and diagona...

📖 Read original article


241. Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centric Proxy Denoising ​

Author: Kangjia Yan, Chenxi Liu, Hao Miao, Xinle Wu, Yan Zhao, Chenjuan Guo, Bin Yang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2510.05589v3 Announce Type: replace-cross Abstract: Effective time series forecasting enables various real-world applications, benefiting from the proliferation of mobile devices. However, the volume of time series data may vary significantly across domains due to high data acquisition costs a...

📖 Read original article


242. When Large Language Models Know the Table: A Framework for Assessing Data Contamination in Tabular Datasets ​

Author: Matteo Silvestri, Fabiano Veglianti, Flavio Giorgi, Fabrizio Silvestri, Gabriele Tolomei
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2510.20351v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly exposed to data contamination, i.e., performance gains driven by prior exposure of test datasets rather than generalization. However, in the context of tabular data, this problem is largely unexpl...

📖 Read original article


243. DeepForgeSeal: Latent Space-Driven Semi-Fragile Watermarking for Deepfake Detection Using Adversarial Reinforcement Learning ​

Author: Tharindu Fernando, Clinton Fookes, Sridha Sridharan
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2511.04949v2 Announce Type: replace-cross Abstract: Rapid advances in generative AI have led to increasingly realistic deepfakes, posing growing challenges for law enforcement and public trust. Existing passive deepfake detectors struggle to keep pace, largely due to their dependence on specif...

📖 Read original article


244. A Lexical Analysis of online Reviews on Human-AI Interactions ​

Author: Parisa Arbab, Xiaowen Fang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2511.13480v2 Announce Type: replace-cross Abstract: This study focuses on understanding the complex dynamics between humans and AI systems by analyzing user reviews. While previous research has explored various aspects of human-AI interaction, such as user perceptions and ethical consideration...

📖 Read original article


245. MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation ​

Author: Basel Shbita, Farhan Ahmed, Chad DeLuca
Published: 8/7/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2511.14967v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown great promise in generating structured diagrams from natural language descriptions, particularly Mermaid sequence diagrams for software engineering. However, the lack of existing benchmarks to assess th...

📖 Read original article


246. Trajectory-guided discharge stratification for heart failure using short-context electronic health record sequence modeling ​

Author: Falk Dippel, Yinan Yu, Annika Rosengren, Martin Lindgren, Christina E. Lundberg, Erik Aerts, Martin Adiels, Helen Sj"oland
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2511.16839v4 Announce Type: replace-cross Abstract: Purpose: Heart failure (HF) discharge planning depends on identifying patients at risk of deterioration or death, yet accurate prediction from routinely collected electronic health records (EHRs) remains challenging. Methods: We develop traje...

📖 Read original article


247. CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning ​

Author: Songqiao Su, Xiaoya Li, Albert Wang, Guoyin Wang, Jiwei Li, Chris Shum
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2512.02551v4 Announce Type: replace-cross Abstract: In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels. Using CUDA execution speed as the RL...

📖 Read original article


248. A note on conditional PAC-efficient reasoning in large language model routing ​

Author: Hao Zeng, Bingyi Jing
Published: 8/7/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, math.ST, stat.TH

arXiv:2512.03057v2 Announce Type: replace-cross Abstract: We study distribution-free risk control for model routing, motivated by large language model reasoning. We formalize pointwise conditional efficiency under a probably approximately correct guarantee and show that it forces a nearly impossible...

📖 Read original article


249. One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs ​

Author: Yixin Tan, Zhe Yu, Rui Wen, Jun Sakuma
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2512.14751v3 Announce Type: replace-cross Abstract: Finetuning pretrained large language models (LLMs) has become the standard paradigm for developing downstream applications. However, its security implications remain unclear, particularly regarding whether finetuned LLMs inherit jailbreak vul...

📖 Read original article


250. Agentic Software Issue Resolution with Large Language Models: A Survey ​

Author: Zhonghao Jiang, David Lo, Zhongxin Liu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2512.22256v2 Announce Type: replace-cross Abstract: Software issue resolution aims to address real-world issues in software repositories based on natural language descriptions provided by users, and represents a key aspect of software maintenance. With the rapid development of large language m...

📖 Read original article


251. All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Generalizable Training ​

Author: Chi Liu, Xin Chen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2601.03895v2 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs). However, GRPO inherits PPO's token-level clipping while replacing token-level advantages with a single ...

📖 Read original article


252. Layer-wise Positional Bias in Short-Context Language Modeling ​

Author: Maryam Rahimi, Mahdi Nouri, Yadollah Yaghoobzadeh
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.04098v2 Announce Type: replace-cross Abstract: Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known as positional bias. Prior work characterizes this bias in model behavior through performance drops in l...

📖 Read original article


253. d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation ​

Author: Yu-Yang Qian, Junda Su, Lanxiang Hu, Peiyuan Zhang, Zhijie Deng, Peng Zhao, Hao Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.07568v3 Announce Type: replace-cross Abstract: Diffusion large language models (dLLMs) offer capabilities beyond those of autoregressive (AR) LLMs, such as parallel decoding and random-order generation. However, realizing these benefits in practice is non-trivial, as dLLMs inherently face...

📖 Read original article


254. FI-TW: An Open Train-Weather Dataset for Railway Delay Analysis in Finland ​

Author: Vinicius Pozzobon Borin, Jean Michel de Souza Sant'Ana, Usama Raheel, Nurul Huda Mahmood
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DB

arXiv:2601.16592v2 Announce Type: replace-cross Abstract: Train delays result from complex interactions between operational, technical, and environmental factors. While weather impacts railway reliability, particularly in Nordic regions, existing datasets rarely integrate meteorological information ...

📖 Read original article


255. PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs ​

Author: Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer, Anurag Kumar, Vivek Kumar
Published: 8/7/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, eess.AS

arXiv:2601.21124v2 Announce Type: replace-cross Abstract: Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained to fixed microphone geometries, preventing deployment across d...

📖 Read original article


256. On the Limits of Layer Pruning for Generative Reasoning in Large Language Models ​

Author: Safal Shrestha, Anubhav Shrestha, Aadim Nepal, Minwu Kim, Keith Ross
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.01997v4 Announce Type: replace-cross Abstract: Recent work has shown that layer pruning can effectively compress large language models (LLMs) while retaining strong performance on classification benchmarks, often with little or no finetuning. In contrast, generative reasoning tasks, such ...

📖 Read original article


257. Trust-Based Incentive Mechanisms in Semi-Decentralized Federated Learning Systems ​

Author: Ajay Kumar Shrestha
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.ET

arXiv:2602.08290v2 Announce Type: replace-cross Abstract: In federated learning (FL), decentralized model training allows multi-ple participants to collaboratively improve a shared machine learning model without exchanging raw data. However, ensuring the integrity and reliability of the system is ch...

📖 Read original article


258. MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling ​

Author: Payel Bhattacharjee, Osvaldo Simeone, Ravi Tandon
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, math.IT

arXiv:2602.17658v4 Announce Type: replace-cross Abstract: Reward modeling is central to RLHF, RLAIF, and PPO-based alignment, but its reliability is often limited by scarce and heterogeneous human preference data. In this paper, we introduce MARS (Margin and Semantic-Aware Data Augmentation for Rewa...

📖 Read original article


259. Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data ​

Author: Jason Miklian, Kristian Hoelscher, John E. Katsos
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2603.00059v3 Announce Type: replace-cross Abstract: How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article p...

📖 Read original article


260. MM-ISTS: Cooperating Irregularly Sampled Time Series Forecasting with Multimodal Vision-Text LLMs ​

Author: Zhi Lei, Chenxi Liu, Hao Miao, Wanghui Qiu, Bin Yang, Chenjuan Guo
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.05997v2 Announce Type: replace-cross Abstract: Irregularly sampled time series (ISTS) are widespread in real-world scenarios, exhibiting asynchronous observations on uneven time intervals across diverse variables. Existing ISTS forecasting methods often solely utilize historical observati...

📖 Read original article


261. When Drafts Evolve: Speculative Decoding Meets Online Learning ​

Author: Yu-Yang Qian, Hao-Cong Wu, Yichao Fu, Hao Zhang, Peng Zhao
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.12617v2 Announce Type: replace-cross Abstract: Speculative decoding has emerged as a widely adopted paradigm for accelerating large language model inference, where a lightweight draft model rapidly generates candidate tokens that are then verified in parallel by a larger target model. How...

📖 Read original article


262. NavTrust: Benchmarking Trustworthiness for Embodied Navigation ​

Author: Huaide Jiang, Yash Chaudhary, Yuping Wang, Zehao Wang, Raghav Sharma, Manan Mehta, Yang Zhou, Lichao Sun, Zhiwen Fan, Zhengzhong Tu, Jiachen Li
Published: 8/7/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.LG, cs.SY, eess.SY

arXiv:2603.19229v2 Announce Type: replace-cross Abstract: There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural language instructions; and Object-Goal Navigation (OGN), where agents navigate to a specified target object. H...

📖 Read original article


263. {\lambda}Split: Self-Supervised Content-Aware Spectral Unmixing for Fluorescence Microscopy ​

Author: Federico Carrara, Talley Lambert, Mehdi Seifi, Florian Jug
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2603.23647v3 Announce Type: replace-cross Abstract: In fluorescence microscopy, spectral unmixing aims to recover individual fluorophore concentrations from spectral images that capture mixed fluorophore emissions. Since classical methods operate pixel-wise and rely on least-squares fitting, t...

📖 Read original article


264. Gender-Based Heterogeneity in Youth Privacy-Protective Behavior for Smart Voice Assistants: Evidence from Multigroup PLS-SEM ​

Author: Molly Campbell, Yulia Bobkova, Ajay Kumar Shrestha
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CY

arXiv:2603.27117v2 Announce Type: replace-cross Abstract: This paper investigates how gender shapes privacy decision-making in youth smart voice assistant (SVA) ecosystems. Using survey data from 469 Canadian youths aged 16-24, we apply multigroup Partial Least Squares Structural Equation Modeling t...

📖 Read original article


265. Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning ​

Author: Cai Zhou, Zekai Wang, Menghua Wu, Qianyu Julie Zhu, Flora C. Shi, Chenyu Wang, Ashia Wilson, Tommi Jaakkola, Stephen Bates
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, stat.AP, stat.ML

arXiv:2604.01170v2 Announce Type: replace-cross Abstract: While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant compute costs. These inefficiencies can be attributed to the miscalibration of post-trained language models...

📖 Read original article


266. Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answering ​

Author: Marco Morini, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2604.01280v2 Announce Type: replace-cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires Multimodal Large Language Models (MLLMs) to identify and combine fine-grained visual cues with retrieved textual evidence. However, retrieval often introduces noisy and partially rel...

📖 Read original article


267. CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering ​

Author: Baicheng Chen, Yu Wang, Ziheng Zhou, Xiangru Liu, Juanru Li, Yilei Chen, Tianxing He
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL

arXiv:2604.03750v2 Announce Type: replace-cross Abstract: Reverse engineering (RE) is central to software security, particularly for cryptographic programs that handle sensitive data and are highly prone to vulnerabilities. It supports critical tasks such as vulnerability discovery and malware analy...

📖 Read original article


268. Ge$^\text{2}$mS-T: Multi-Dimensional Grouping for Ultra-High Energy Efficiency in Spiking Transformer ​

Author: Zecheng Hao, Shenghao Xie, Kang Chen, Wenxuan Liu, Zhaofei Yu, Tiejun Huang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CV

arXiv:2604.08894v2 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs) offer superior energy efficiency over Artificial Neural Networks (ANNs). However, they encounter significant deficiencies in training and inference metrics when applied to Spiking Vision Transformers (S-ViTs). E...

📖 Read original article


269. SkillMOO: Multi-Objective Optimization of Agent Skills for Software Engineering ​

Author: Jingzhi Gong, Ruizhen Gu, Zhiwei Fei, Yazhuo Cao, Lukas Twist, Alina Geiger, Shuo Han, Dominik Sobania, Federica Sarro, Jie M. Zhang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2604.09297v3 Announce Type: replace-cross Abstract: Agent skills are increasingly used to configure coding agents for software engineering (SE) tasks, yet current practice treats them as static, hand-crafted assets, or evolved on pass rate alone. This is insufficient: a skill can improve task ...

📖 Read original article


270. Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations ​

Author: Patrick Keough
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2604.17359v2 Announce Type: replace-cross Abstract: Language models asked to simulate psychiatric patients produce cases that survive inspection one at a time and populations that match no real one. We gave GPT-4o-mini, Gemini-3-Flash, DeepSeek-V3 and GLM-4.7 each of 120 demographic cohorts un...

📖 Read original article


271. Text Steganography with Dynamic Codebook and Multimodal Large Language Model ​

Author: Jianxin Gao, Ruohan Lei, Wanli Peng
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2604.20269v2 Announce Type: replace-cross Abstract: With the popularity of the large language models (LLMs), text steganography has achieved remarkable performance. However, existing methods still have some issues: (1) For the white-box paradigm, this steganography behavior is prone to exposur...

📖 Read original article


272. Supervised Learning Has a Geometric Blind Spot ​

Author: Vishal Rajput
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2604.21395v3 Announce Type: replace-cross Abstract: Ordinary supervised training minimises the task loss and then stops. It never pays for how far the representation moves when the input is nudged along directions that helped fit training labels---including directions that are nuisance at depl...

📖 Read original article


273. Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination ​

Author: Jonathan Spieler, Sven Behnke
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO

arXiv:2605.04568v3 Announce Type: replace-cross Abstract: State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learned policy networks, or a combination of policy networks and planning. Hybrid approaches that combine Mod...

📖 Read original article


274. Skill Neologisms: Towards Skill-based Continual Learning ​

Author: Antonin Berthon, Nicolas Astorga, Mihaela van der Schaar
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.04970v3 Announce Type: replace-cross Abstract: Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. However, extending model capabilities to new skills in a scalable manner is an open problem: fine-tuning and parameter-efficient v...

📖 Read original article


275. The Impossibility Triangle of Long-Context Modeling ​

Author: Yan Zhou
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2605.05066v2 Announce Type: replace-cross Abstract: We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactnes...

📖 Read original article


276. Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models ​

Author: Kyuyoung Kim, Kevin Wang, Yunfei Xie, Peiyang Xu, Peiyao Sheng, Chen Wei, Zhangyang Wang, Jinwoo Shin, Pramod Viswanath, Sewoong Oh
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2605.12519v2 Announce Type: replace-cross Abstract: Training language models to produce both correct answers and sound reasoning remains an open challenge. Reinforcement learning with verifiable rewards typically optimizes only final outcomes, which can improve task accuracy at the expense of ...

📖 Read original article


277. Fast Rates for Inverse Reinforcement Learning ​

Author: Andreas Schlaginhaufen, Maryam Kamgarpour
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2605.14599v2 Announce Type: replace-cross Abstract: We establish novel structural and statistical results for entropy-regularized min-max inverse reinforcement learning (Min-Max-IRL) in finite-horizon MDPs with Borel state and action spaces. We show that maximum likelihood estimation (MLE) and...

📖 Read original article


278. Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift ​

Author: Qinwu Xu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.DB, cs.LG

arXiv:2605.16411v2 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization u...

📖 Read original article


279. CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning ​

Author: Yang Liu, Toan Nguyen, Flora D. Salim
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CV

arXiv:2605.20247v2 Announce Type: replace-cross Abstract: Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs) and vision--language models (VLMs). Although Mixture-of-Experts (MoE) architectures offer an efficient path to scaling, existing LoRA-based...

📖 Read original article


280. Domain-Gated Latent Diffusion: Generative Inverse Design of HMX-Class Energetic Materials with First-Principles Validation ​

Author: Yehudit Aperstein, Alexander Apartsin
Published: 8/7/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.AI

arXiv:2605.26540v2 Announce Type: replace-cross Abstract: Energetic materials power mining, demolition, propulsion and airbags, yet today's compounds were designed decades ago. A successor must combine high energy release, low sensitivity to accidental initiation and a practical synthesis route, fou...

📖 Read original article


281. BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali ​

Author: Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham, Ajwad Abrar, Ishmam Tashdeed, Md Taukir Azam Chowdhury
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2605.31483v4 Announce Type: replace-cross Abstract: Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucination evaluation framew...

📖 Read original article


282. PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say ​

Author: Mingxuan Zhang, Jiahui Han, Dadi Guo, Songze Li, Guanchu Wang, Na Zou, Dongrui Liu, Xia Hu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2606.00152v2 Announce Type: replace-cross Abstract: LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often acquire more sensitive information than the task requires. Existing privacy benchmarks audit what the a...

📖 Read original article


283. PhysScene: A Scene Graph Dataset for Scientific Visual Reasoning in Physics Experiments ​

Author: Minghao Zou, Qingtian Zeng, Shangkun Liu, Yanda Meng, Guanghui Yue, Baoquan Zhao, Abdulmotaleb El Saddik, Wei Zhou
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.09368v2 Announce Type: replace-cross Abstract: Scene Graphs (SGs) provide structured representations of visual scenes by modeling objects and their pairwise relationships. Despite recent progress, existing datasets primarily focus on generic natural contexts, leaving domain-specific and f...

📖 Read original article


284. Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech ​

Author: Adarsh Arigala, Arjun Gangwar, S Umesh, Yova Kementchedjhieva
Published: 8/7/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CV, cs.SD

arXiv:2606.14750v2 Announce Type: replace-cross Abstract: Recent advances in pixel-based text modeling show that representing text as images enables models to exploit visual cues for language understanding. Grounding text in its visual form allows structurally similar characters with different Unico...

📖 Read original article


285. RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos ​

Author: Zhenhua Wu, Yun Pang, Mingkun Chang, Yuwei Ning, Liangzhi Wang, Yi Xiao, Guanbin Li
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.16278v3 Announce Type: replace-cross Abstract: Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect at scale. Editable 3D Gaussian Splatting (3DGS) simulation offers a scalable alternative through real-scene reconstruction a...

📖 Read original article


286. Beyond Weights and Gradients: A Taxonomy of Federated Learning Messages ​

Author: Alvaro Javier Vargas Guerrero, Xinguang Wang, Quang Manh Doan, Guy Nagels
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.16891v2 Announce Type: replace-cross Abstract: Federated Learning is rapidly evolving beyond the exchange of traditional model weights and gradients, yet existing definitions fail to capture the full scope of modern payloads like synthetic data and federated analytics. This paper addresse...

📖 Read original article


287. As You Wish: Mission Planning with Formal Verification using LLMs in Precision Agriculture ​

Author: Marcos Abel Zuzu'arregui, Stefano Carpin
Published: 8/7/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2606.18519v2 Announce Type: replace-cross Abstract: Though robotic systems are now being commercialized and deployed in various industries, many of these systems are highly specialized and often require an advanced skill set to operate and ensure they perform as instructed. To mitigate this pr...

📖 Read original article


288. Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents ​

Author: Han Chi, Jiaxin Qi, Yan Cui, Baisheng Lai, Jianqiang Huang
Published: 8/7/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2606.21140v2 Announce Type: replace-cross Abstract: Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLIs determine how models invoke tools, maintain interaction history, and recover from failures. Consequen...

📖 Read original article


289. Accelerating Q-learning through Efficient Value-Sharing across Actions ​

Author: Prabhat Nagarajan, Brett Daley, Martha White, Marlos C. Machado
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.29806v2 Announce Type: replace-cross Abstract: Action values are foundational to many control algorithms such as Q-learning. Therefore, efficient action-value learning is central to reinforcement learning (RL). However, learning them can be slow, requiring many updates to move values from...

📖 Read original article


Author: Jesse Yusuf Chan (Zexi Chen), Mengyao Chen, Yang Hong, Ziyun Song, Haoming Wang, Xianlong Xu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2607.05412v2 Announce Type: replace-cross Abstract: STEM education faces challenges in personalization and interdisciplinary integration. AI technology has brought new possibilities, but the mechanisms by which AI reshapes the STEM education ecosystem require systematic investigation. This stu...

📖 Read original article


291. Automated Numerical Stability Analysis of Deep Learning Operators ​

Author: Xinye Chen
Published: 8/7/2026, 4:00:00 AM
Categories: math.NA, cs.AI, cs.NA

arXiv:2607.25494v2 Announce Type: replace-cross Abstract: Finite-precision arithmetic unavoidably introduces numerical approximation errors. Numerical computations may use insufficient precision or an improper formulation, which leads to numerical instability. In this paper, we introduce the first u...

📖 Read original article


292. Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility ​

Author: Yansen Zhang, Yilu Liu, Tianyu Liu, Jiamin Chen, Xiaokun Zhang, Kai Xie, Xue Liu, Yiyan Qi, Chen Ma
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.26828v3 Announce Type: replace-cross Abstract: Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt le...

📖 Read original article


293. WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization ​

Author: Fanzhe Wei (Metask Lab), Li Liu (Metask Lab), Ziyang Wang (Metask Lab), Chenyu Wang (Metask Lab)
Published: 8/7/2026, 4:00:00 AM
Categories: cs.AR, cs.AI

arXiv:2607.28699v2 Announce Type: replace-cross Abstract: KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the request it is serving right now. We give it a provably sound runtime meter -- a "DTrace for KV quantizat...

📖 Read original article


294. InferQ: A Database-Oriented Benchmark for Quantum Circuits Simulation ​

Author: Andrei Ilinescu, Aadi Patwardhan, Rihan Hai
Published: 8/7/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.DB

arXiv:2607.29134v2 Announce Type: replace-cross Abstract: Recent work suggests that relational database management systems (RDBMSs) can execute quantum circuit simulation by compiling the simulation into SQL workloads (primarily join-and-aggregate tensor contractions). While early results are promis...

📖 Read original article


295. Role Steering of Language Models for Social Simulations ​

Author: Isaac Song, Mohammed Rehan Parwani, Glenn Matlin, Emile Anand, Akhil Theerthala, Arjun Chatterjee, Anthony Wen-Ming Zang, Maria Kostylew, Yonadav G. Shavit, Sebastien Krier, Mark Riedl
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.00023v2 Announce Type: replace-cross Abstract: Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated population. We introduce an activation-steering screening workflow for role-conditioned agents: d...

📖 Read original article


296. Rapid Embodiment Adaptation for Quadrupedal Locomotion ​

Author: Dichen Li, Bo Ai, Nico Bohlinger, Jan Peters, Hao Su, Henrik I. Christensen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2608.01506v2 Announce Type: replace-cross Abstract: Humans readily adapt their movements as their bodies change through aging, injury, or load carrying, but learning-based robot policies often break when hardware properties shift. We introduce an online embodiment adaptation framework for quad...

📖 Read original article


297. PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation ​

Author: Xiaohan Jiang, Zeyu Li, Wei Zhang, Jiang Xu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.ET, cs.AI

arXiv:2608.01791v3 Announce Type: replace-cross Abstract: The rapid development of photonic integrated circuits (PICs) is shifting the design flow from traditional graphical user interface (GUI)-based methods to script-based methods for higher flexibility, portability, and maintainability. However, ...

📖 Read original article


298. OpenAI Privacy Filter: A Cross-Lingual, Cross-Domain PII Evaluation Across 32 Benchmarks ​

Author: Rohith Uppala
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.02616v2 Announce Type: replace-cross Abstract: We present what is, to our knowledge, the first systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter model that converts an autoregressive language model into a bidirectional PII detector, across 32 benchmarks spanning 14 ...

📖 Read original article


299. Output-Aware Rotation for INT2 KV-Cache Quantization ​

Author: Vincent-Daniel Yun, Woosang Lim, Minsoo Cheong, Sunwoo Lee, Murali Annavaram, Sai Praneeth Karimireddy, Sungjoo Yoo
Published: 8/7/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.02691v2 Announce Type: replace-cross Abstract: The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important. However, existing rotation-based INT2 methods optimize cache...

📖 Read original article


300. AI Security Leaderboard: Methodology, Results and Minimal Standard ​

Author: Jasper Timm, Lukas Struppek, Ziwei Xu, Grace Cheong, Oscar Mata, Dan Zhao, Mick Yang, Isadora De Andrade, Xiaojun Jia, Yiming Li, Samuel Bauer, Heather McIntyre, Adam Gleave, Edward Yee, Kellin Pelrine
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL

arXiv:2608.03070v2 Announce Type: replace-cross Abstract: The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests models against the FAR$.$AI Minimal Standard for Safeguards, which represents a minimum bar for securi...

📖 Read original article


301. EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks ​

Author: Pau Arnal, Khaled Denfir, Danylo Smahliuk, Amrut Avhad, Marcus A. Castro
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.04549v2 Announce Type: replace-cross Abstract: Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated on. We dedicate more than 4,000 human expert hours to evaluate a selection of six frontier LLMs on a mem...

📖 Read original article


302. Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders ​

Author: Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Chengpeng Fu, Yu Wang, Ming Liu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.04586v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from the curse o...

📖 Read original article


303. DisMix: Order-Aware Mixup for Medical Imaging via Disentangling Ordinal and Non-Ordinal Features ​

Author: Dileepa Pitawela, Gustavo Carneiro, Hsiang-Ting Chen
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.04652v2 Announce Type: replace-cross Abstract: Image mixup is a widely adopted data augmentation strategy, yet it is ill-suited for ordinal classification tasks such as medical disease grading, where labels encode a progression of severity. By indiscriminately blending disease-severity cu...

📖 Read original article


304. InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval ​

Author: Tsz Ting Chung, Jiangnan Li, Jie Zhou, Mo Yu
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.04761v2 Announce Type: replace-cross Abstract: Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated experience into actionable guidance. At each decision step, retrieving the right insight can help the ...

📖 Read original article


305. RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists ​

Author: Yuexi Yang, Alyssa Wu, Ji Luo, Richeng Xuan, Zhichao Hu, Yuhong Liu, Zhen Qin
Published: 8/7/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.04783v2 Announce Type: replace-cross Abstract: The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repository-scale assistance. However, existing benchmarks largely rely on bug reports from GitHub Issues, which ...

📖 Read original article


306. A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination ​

Author: Wenxiao Zhao, Dong Liu, Kaiyi Xu, Feng Liu, Zhen Zhao, Fei Ben, Shu Wang, Wenhao Li, Ying Nian Wu, Fenghua Ling, Haobo Li, Lei Bai
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.04872v2 Announce Type: replace-cross Abstract: Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We propose A-SR...

📖 Read original article


307. OPD-V: Visual On-Policy Self-Distillation with Modality Balance ​

Author: Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
Published: 8/7/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.05131v2 Announce Type: replace-cross Abstract: On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-...

📖 Read original article