arXiv cs.AI - 2026-08-31 ​
302 items collected.
1. Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model ​
Author: David Noever, Forrest McKee
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27459v1 Announce Type: new Abstract: In 2011, IBM's Watson was something like a sealed capsule of its era's queryable knowledge. Its DeepQA system defeated the strongest human Jeopardy! champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a ...
2. Rating the Raters: Rasch Measurement Theory for LLM Evaluation ​
Author: Pratik S. Sachdeva, Nathan Boudol
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27463v1 Announce Type: new Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed wi...
3. Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI ​
Author: Andrea Beretta, Salvatore Rinzivillo
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2608.27464v1 Announce Type: new Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking. Drawing on Sharot and Sunstein's framework of information-seeking motives, we propose that people evaluate whe...
4. Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis ​
Author: Deborah Dore, Greta Damo, Elena Cabrio, Serena Villata
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.27471v1 Announce Type: new Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political debates, where public opinion is shaped. Spotting a fallacious argument requires contextual knowledge b...
5. LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation ​
Author: Neville K. Kitson, Anthony Constantinou
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27472v1 Announce Type: new Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge. We propose combining these complementary sources throug...
6. Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields ​
Author: YuJie Huang, WenWu He, ZhuoEr Lin, Congcong Liu, Dong Liang, Zhuo-Xu Cui
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.27475v1 Announce Type: new Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it. These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexib...
7. Class-Based Heuristic Selection for Solving the Flying Block Puzzle ​
Author: Sanyar Ahmadi, Pedram Asadzadeh, Amanj Khorramian
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27476v1 Announce Type: new Abstract: Heuristic search underlies planning in autonomous systems ranging from warehouse logistics to robotic navigation, yet generic heuristics fail to exploit the structural constraints that govern constrained spatial domains, causing search performance to d...
8. Benchmarking General Mobile Assistants in Challenging Real-World Scenarios ​
Author: Yiqi Zhu, Feiyu Gao, Jiaxing Fan, Jiahui Zeng, Minggang Wu, Chenliang Li, Haiyang Xu, Peng Li, Ming Yan, Yang Liu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.27477v1 Announce Type: new Abstract: Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks. Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but...
9. Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations ​
Author: D. K. C. Senevirathna, A. A. E. Nanayakkara, H. M. C. K. Kulathunga, J. K. D. P. Nadula, R. M. Mapatuna, Malithi Nawarathne, Jaliya L. Wijayaraja, P. D. Senanayake, Samitha Vidhanaarachchi, Kalpani Manathunga
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SD
arXiv:2608.27480v1 Announce Type: new Abstract: Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial damage when infestations remain undetected. This study proposes an IoT-enabled acoustic monitoring fram...
10. Context Localization for Generalized Level-Based Evaluation in Knowledge-Based Systems ​
Author: Ondrej Hutn'{i}k, Nat'{a}lia Pu\v{s}k'{a}rov'{a}
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27482v1 Announce Type: new Abstract: We study context localization for generalized level-based evaluation in knowledge-based systems. The framework models situations where a structured nonnegative score, defined on facts, rules, cases, criteria or evidence units, is evaluated through cond...
11. CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence ​
Author: Pratik Ghawate, Tanvi Patil
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.IR, cs.MA
arXiv:2608.27484v1 Announce Type: new Abstract: Artificial intelligence is transforming personalized healthcare, yet fragmented clinical, self reported, and wearable evidence remains difficult to interpret and trace. We present CareGraph, an auditable hybrid AI framework that converts heterogeneous ...
12. Thinking Costs Tokens: When More Structure is Worth the Price ​
Author: Thomas Nolasque, John Grey, Calista Pham, Ankit Vani
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27506v1 Announce Type: new Abstract: Adding inference structure to a language model lets it search, verify, and revise, but these actions consume the very budget they are supposed to use well. In this paper, we investigate whether there exists a token-budget threshold, below which the ove...
13. WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning ​
Author: Yu Han, Tianwen Qian
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27508v1 Announce Type: new Abstract: GUI agents trained with reinforcement learning (RL) have showcased strong environment learning capabilities on mobile platforms. However, RL typically demands extensive real-environment interactions, leading to high resource costs and instability, espe...
14. SETU: An Agentic Ecosystem for Multilingual, Persona-Aware Communication Coaching ​
Author: Jonnalagadda Maruthi Tejas, Uponika Barman Roy, Tilottama Goswami, Samir Goswami, Mousita Dhar
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27524v1 Announce Type: new Abstract: Corporate training teams need scalable and explainable tools to improve workforce communication in multilingual settings. Existing systems often score text, audio, or video in isolation, or produce black-box outputs that are difficult to audit for coac...
15. Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety Moderator ​
Author: Varun Singh, Anuj Doshi, Makesh Narsimhan Sreedhar, Shaona Ghosh, Katherine Luna
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27548v1 Announce Type: new Abstract: Safety moderation for deployed AI applications is moving beyond text-only prompts: systems increasingly need to judge images, documents, screenshots, and generated responses under policies that vary across domains. Existing guardrails usually cover onl...
16. LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails ​
Author: Ziyang Chen, Xing Wu, Songlin Hu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27580v1 Announce Type: new Abstract: Safety guardrails serve as the last line of defense against harmful inputs and outputs of large language models (LLMs), yet they are trained and evaluated almost exclusively on short text. We present LongGuard, a framework that evaluates, mechanistical...
17. Generative AI Expands the Intellectual Reach of Course Based Undergraduate Research Experiences (CUREs) ​
Author: Aditi Babar, Kristin J. Davin, Alex Dornburg
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2608.27638v1 Announce Type: new Abstract: Course-based undergraduate research experiences (CUREs) broaden access to authentic scientific inquiry through responsive instructor support as research problems become increasingly complex. Generative artificial intelligence (GenAI) may extend this su...
18. If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary ​
Author: Marc Millstone, Tyler Akidau, Johannes Br"uderl, Marat Pekker
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27646v1 Announce Type: new Abstract: Give an agent a human's credential and it inherits the person's reach without the judgment that limits its use. It can sweep every reachable record into model context, where hidden instructions steer its next call, and every request stays credential-va...
19. A Framework for Object-Centric Predictive Monitoring of Collaborative Processes ​
Author: Daniel Calegari, Andrea Delgado, Leonel Pe~na, Mart'in Rubio
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27671v1 Announce Type: new Abstract: Predictive Process Monitoring (PPM) of collaborative, inter-organizational processes requires reasoning over multiple interdependent entities, including participants, messages, local executions, and the global collaboration case. Existing approaches ex...
20. Agents for Everyone: A Workshop Framework for Building Agentic AI Capabilities in a Distributed Curation Community ​
Author: Seth Carbon, Sierra Moxon, Kimberly Van Auken, Pascale Gaudet, Christopher J. Mungall
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27675v1 Announce Type: new Abstract: Agentic AI has the potential to accelerate curation of biological databases and knowledge bases. However, uptake has been hindered by a number of challenges and obstacles, including access to agents and appropriate training. Here we describe how we hav...
21. PCFBench: A Diagnostic Benchmark for Product Carbon Footprint Estimation ​
Author: Krishna Rao, Andrew Dumit, Shaena Ulissi, Jacob Feintzeig, P. James Joyce, Daniel Frank, Steven Watson, Jonathan Glidden, Gizem Ilayda Dinc, Travis M. Kwee
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27716v1 Announce Type: new Abstract: AI systems are being deployed on high-stakes, domain-specific workflows that demand correctness not just in the final output, but at every intermediate step. One such workflow is estimating a product carbon footprint (PCF), the greenhouse-gas emissions...
22. Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative Controls ​
Author: Manuel Cherep, Pattie Maes, Nikhil Singh
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27727v1 Announce Type: new Abstract: A model's behavior on a task is jointly determined by the input it receives and the prior it brings in, i.e. the distribution over stimuli it implicitly expects. Interpretability research has traditionally studied models by holding inputs fixed and exa...
23. Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models ​
Author: Justin Bronder
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.27768v1 Announce Type: new Abstract: A language model with access to tools can commit to a final claim unsupported by the evidence it has seen, even when a single available tool call would resolve the uncertainty and its instructions explicitly forbid assumptions and guesses. We separate ...
24. Credo: Reusable Declarative Primitives for Agentic Workflows ​
Author: Duo Lu, Andrew Crotty, U\u{g}ur \c{C}etintemel
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.DB
arXiv:2608.27790v1 Announce Type: new Abstract: An LLM application depends on both a model and a harness: the program that determines what each call sees, how many calls to make, and which answers to trust. Coding agents can now discover strong harnesses by searching over candidate programs, but the...
25. ReToolSQL: Agentic Reinforcement Learning for Robust Text-to-SQL ​
Author: Pratik Kakkar, Chandra Dhir, Ravi Shankar, Pareekshit Reddy Gaddam, Anup Shirgaonkar
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27796v1 Announce Type: new Abstract: Recent work has shown that reinforcement learning from execution feedback can substantially improve text-to-SQL performance, often enabling smaller models to match or exceed much larger systems. However, most existing approaches treat SQL generation as...
26. CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action ​
Author: Lekai Chen, Alvaro Velasquez, Ashutosh Trivedi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.FL
arXiv:2608.27797v1 Announce Type: new Abstract: Natural-language tasking of embodied agents is rarely just goal specification: users also impose constraints that must persist while the world changes. Code-generating LLM agents can produce plausible behaviors for such instructions, but their free-for...
27. CURA: Certified Runtime Alarms for Computer-Use Agents ​
Author: Divake Kumar, Sina Tayebati, Devashri Naik, Amanda Sofie Rios, Nilesh Ahuja, Omesh Tickoo, Ranganath Krishnan, Amit Ranjan Trivedi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2608.27808v1 Announce Type: new Abstract: Self-report is the cheapest oversight channel a deployer has, and on capable computer-use agents (CUAs) it fails precisely where oversight matters. On 361 OSWorld tasks our pipeline, a read-only feasibility gate, a planner, and a GUI executor, reaches ...
28. AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics ​
Author: Tejas Srinivasan, Shikib Mehri, Nandita Shankar Naik, Anirban Das, William M. Campbell, Jesse Thomason
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27818v1 Announce Type: new Abstract: User preferences in user-agent collaboration are rarely static and fully-specified upfront: preferences are formed, revealed, adjusted, and relaxed during interaction. Existing benchmarks for evaluating user-agent collaboration focus almost exclusively...
29. Evidential-Based Higher-Order Set Argumentation Framework ​
Author: Shuai Tang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, math.LO
arXiv:2608.27824v1 Announce Type: new Abstract: Evidential argumentation extends Dung's abstract argumentation by requiring arguments and interactions to be backed by chains of evidence rooted in prima-facie elements. However, existing formalisms lack a unified treatment of evidential support, highe...
30. RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests ​
Author: Gyuhyeong Kim, Hyojung Gwon, Jeonghyeon Kim, Kyuhong Shim, Sunjae Lee
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.27831v1 Announce Type: new Abstract: Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues--long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To cha...
31. KLOD: Locality-Preserving Knowledge Editing via Non-Target Distribution Preservation ​
Author: Hojun Jeong, Gyunyeop Kim, Sangwoo Kang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27839v1 Announce Type: new Abstract: Fine-tuning-based knowledge editing is simple and architecture-agnostic, but standard cross-entropy increases the edited target probability without explicitly constraining changes in the non-target output distribution. In sequential editing, such uncon...
32. An Empirical Evaluation of Cross-City POI Recommendation on a Large-Scale Benchmark ​
Author: Peibo Li, Yang Song, Hao Xue, Maarten de Rijke, Flora D. Salim
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.IR
arXiv:2608.27840v1 Announce Type: new Abstract: Cross-city point-of-interest (POI) recommendation is crucial for navigating unfamiliar urban environments, yet its progress has historically been constrained by data limitations. Using the recently proposed large-scale benchmark Trip World, we empirica...
33. From Uncertainty to Clinical Risk: Severity-Aware Conformal Planning for Interactive Medical Diagnosis ​
Author: Yue Zhou, Haiyang Zhou, Jin Zhang, Kong Wang, Yongxin Ni, Youhua Li, Hanwen Du
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27847v1 Announce Type: new Abstract: Interactive medical diagnosis dynamically acquires patient information through multiple rounds of questioning, supporting accurate, efficient, and safe clinical decisions under incomplete evidence. Existing methods commonly guide information acquisitio...
34. SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models ​
Author: Enqiao Lu, Xingrui Yu, Yiwei Fu, Zhenglin Wan, Pengfei Zhou, Wangbo Zhao, Muqing Jian, Xueyi Zhang, Yang You, Ivor Tsang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27857v1 Announce Type: new Abstract: Spiking neural networks (SNNs) offer a path to energy-efficient language modeling through sparse encoding and event-driven computation, but training capable spiking language models from scratch remains difficult. A practical alternative is ANN-to-SNN m...
35. CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning ​
Author: Runze Liu, Naibin Gu, Mingxu Ai, Yuqing Li, Peng Fu, Zheng Lin, Weiping Wang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27867v1 Announce Type: new Abstract: Continual multimodal instruction tuning requires multimodal large language models to acquire new task abilities sequentially while preserving previously learned knowledge. LoRA-MoE provides a promising solution by introducing expert-based capacity, but...
36. See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEs ​
Author: Sarang Manoj Pekhale, Amartya Roy, Rajat Sarkar, Souvik Chakraborty
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27869v1 Announce Type: new Abstract: Discovering governing partial differential equations (PDEs) from observational data remains a core challenge across the sciences. Existing sparse-regression, symbolic-regression, and LLM-based approaches can be constrained by predefined libraries, nois...
37. HyQuant: Hybrid-Precision Quantization for LLM Attention ​
Author: Jiatong Ding, Bingxin Xing, Yu Zhang, Dian Ding, Xiaodong Yi, Xianbin Ouyang, Feihu Zhou, Kun Zhang, Zhenyu Guo, Hao Pan, Guangtao Xue, Yiming Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27875v1 Announce Type: new Abstract: Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low bit-widths, causing performance degradation...
38. Resource Constraints and Performance in Agentic AI Systems ​
Author: Amaz Salman, Malka Halgamuge, Teo Susnjak
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27886v1 Announce Type: new Abstract: Progress toward more autonomous AI increasingly depends on agentic systems that combine a language model with tools, memory, state management, and multi-step execution. These mechanisms shape both task capability and operational burden. We compare Open...
39. Rubric-to-Code Credit Assignment for Reinforcement Learning ​
Author: Rui Jin, Jikai Chen, Yihan Chen, Hao Zhou, Demin Zhu, Kaichen Yang, Dong Wang, Chenyi Zhuang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27906v1 Announce Type: new Abstract: Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from natural language requests. Unlike conventional code generation, application quality depends on multiple user-facing functional requirem...
40. AI Alignment through a Game-theoretic Lens: A Survey ​
Author: Yanan Cai, Zhongrui Zhao, Zhigang Lu, Ickjai Lee, Wei Emma Zhang, Minhui Xue, Yihong Zhang, Shuchao Pang, Wei Xiang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.GT
arXiv:2608.27910v1 Announce Type: new Abstract: As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in improving helpfulness, harmlessness, ...
41. From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning ​
Author: Lokendra Birla, Milind Savagaonkar, Visnu Srinivasan, Sowmya Rasipuram, Shubhashis Sengupta
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27919v1 Announce Type: new Abstract: Financial question answering (QA) has emerged as a key benchmark for evaluating the performance of Large Language Models (LLMs) on domain-specific tasks involving complex data formats such as tables, charts, and rich textual narratives. While recent ad...
42. A Deep Learning-Based Stacking Ensemble Framework for Turbofan Engine Remaining Useful Life Prediction ​
Author: Limon Bin Hossain, Md. Salehin Seyam, Md Rashedul Islam, Abdur Rahman, Md Sharifuzzaman
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27940v1 Announce Type: new Abstract: This study proposes a two-level stacking ensemble framework for Remaining Useful Life (RUL) prediction of turbofan engines, evaluated on the NASA C-MAPSS benchmark using the FD001 and FD003 subsets. The framework integrates four heterogeneous deep lear...
43. CASTANET: Causality-Aware Spatio-Temporal Adversarial Network Using Traffic Incident Effects ​
Author: Toshiya Kitahara, Ryu Shirakami, Koh Takeuchi, Hisashi Kashima
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27942v1 Announce Type: new Abstract: Predicting non-periodic traffic congestion caused by sudden incidents (e.g., accidents and road damage) is crucial for advanced intelligent transportation systems. However, incident-driven congestion is difficult to forecast because incidents are extre...
44. Cross-Session Decomposition Attacks: Scaling Risk and Intent-Aligned Retrieval Defense ​
Author: Disen Liao, Yihan Wang, Freda Shi, Yaoliang Yu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27945v1 Announce Type: new Abstract: Scaling laws are usually read as a capability story: lower language-modeling loss yields more useful models. We study a safety consequence of this mechanism in \emph{cross-session decomposition attacks}, where benign-looking subqueries are asked across...
45. The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs ​
Author: Yucheng Wang, Yuetian Du, Zhengyi Liu, Rongyu Zhang, Bing Zhao, Boyu Yang, Ming Kong, Lin Qu, Hu Wei, Jie Liu, Qiang Zhu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27953v1 Announce Type: new Abstract: Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences. Existing benchmarks largely target bounded settings with fixed variables or single gold outcomes,...
46. When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation ​
Author: Siyuan Gan, Yuhan Li, Xiran Wang, Linjian Meng, Boyan Wang, Zhen Zhao, Jing Huo, Yang Gao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27960v1 Announce Type: new Abstract: On-policy distillation (OPD) has recently emerged as a popular post-training paradigm for large language models (LLMs), providing an efficient way to transfer the knowledge and capabilities of teacher models into student models. However, teacher guidan...
47. SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch Probing ​
Author: Wanli Cheng, Haiya Xiang, Juntao Li, Hongling Wang, Wenliang Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27963v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong reasoning capabilities, yet long-chain reasoning becomes inefficient once the intermediate answer stabilizes across reasoning steps: additional reasoning yields little marginal benefit while incurring substa...
48. AERA: Adaptive Evidence Residual Allocation for Efficient Test-Time Reasoning ​
Author: Ziming Wang, Ivor Tsang, Hangwei Qian
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27964v1 Announce Type: new Abstract: Test-time scaling improves language-model reasoning by generating additional candidate solutions, but allocating the same inference budget to every problem is computationally wasteful. Existing adaptive stopping methods commonly rely on confidence, agr...
49. openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents ​
Author: openJiuwen Team, Tao Yu, Xinyu Zhang, Qianqian Chen, Xiaoneng Xiang, Chia Kwangyang, Xingchen Huang, Ran Chen, Yangkai Ding, Zheng Wang, Yeo Boon Hong, Bingzheng Gan, Enrui Hu, Shuo Cheng, Deyang Li, Ruifeng Shi, Hongbo Wang, Qi Ye, Xuefeng Jin, Zhangchun Zhao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27969v1 Announce Type: new Abstract: Long-horizon coding agents operate over evolving repository states while increasingly relying on heterogeneous capabilities, delegated agents, and multi-agent coordination. These trends pose two complementary challenges for the agent harness. First, de...
50. Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic Sampling ​
Author: Siyuan Gan, Yuhan Li, Xiran Wang, Linjian Meng, Boyan Wang, Zhen Zhao, Jing Huo, Lei Bai, Yang Gao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27982v1 Announce Type: new Abstract: Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) is a prominent variant of Group Relative Policy Optimization (GRPO). DAPO introduces several improvements over GRPO. Among these, Dynamic Sampling contributes the most to DAPO's accuracy ga...
51. When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems ​
Author: Yangxiao Jiang, Jiarun Fan, Mingcong Xu, Yanxi Guo, Jiwen Feng, Shanqing Xu, Mengchen Qian, Wei Chen, Xiaojin Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27984v1 Announce Type: new Abstract: Multi-Agent Systems (MAS) have recently moved from static workflows toward dynamically generated collaboration topologies. However, existing topology generation methods rely primarily on the parametric knowledge of large language models, with external ...
52. GOD: Govern, Observe, and Direct - A Real-Time Control Room for Agent Societies ​
Author: Yige Luo, Ran Guan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.27992v1 Announce Type: new Abstract: Generative-agent systems are easier to start than to inspect. A run can contain many agents, locations, messages, commands, and model calls, yet the operator often gets either a finished replay or raw logs. That makes it hard to ask why an agent moved,...
53. Should I Use This Synthetic Dataset for Training? How to Test with Minimal Real Data ​
Author: Zhenyu Tao, Wei Xu, Xiaohu You, Petar Popovski, Osvaldo Simeone
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27996v1 Announce Type: new Abstract: Digital twins (DTs) and learned world models are increasingly used to generate synthetic data that augment the scarce real datasets available for training artificial intelligence (AI) models in engineering systems. Owing to the inevitable simulation-to...
54. Automated Analysis Framework for Multilingual Climate-Health Literature Based on Multi-Agent Large Language Model ​
Author: Yuze Sun, Shihui Zhang, Jiancheng Pan, Yunjia Ye, Wentao Luo, Jiahao Li, Quan Zhang, Wenjia Cai, Xiaomeng Huang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27998v1 Announce Type: new Abstract: The rapid proliferation of interdisciplinary and multilingual scientific literature has left traditional manual analysis and single-algorithm methods plagued by low efficiency, poor scalability, and insufficient domain adaptability. Targeting the liter...
55. PhenoIntel: A Lifecycle-Aligned Multi-Agent Web Application for Verified, Accessible Plant Phenotype Analysis ​
Author: Narendren S V, Soumyashree Kar
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27999v1 Announce Type: new Abstract: Existing conversational plant-phenotyping platforms are difficult for plant scientists to use and lack the reliability scientific research demands: failed analyses are reported as valid measurements rather than flagged as missing, statistical tests run...
56. Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM Agents ​
Author: Yuxu Ge
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28011v1 Announce Type: new Abstract: Trajectory-level credit assignment can localize which module of a tool-using LLM agent causes failures using only verifiable signals. We ask whether such failure credit should route a fixed zeroth-order/evolution-strategies (ZO/ES) perturbation budget....
57. String: An Agentic OS Where Every App Is a Markdown File ​
Author: Jookyung Song, Nojun Kwak, Simyung Chang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28027v1 Announce Type: new Abstract: LLM agents have become a new class of software user, but every surface they work through was designed for someone else. Pages are built for human eyes, which can skim and ignore; tool schemas for programs, which pay nothing to carry definitions they ne...
58. WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents ​
Author: Zongkai Liu, Hui Zhang, Liqiang Niu, Zhen Cao, Han Li, Juntao Liu, Wenchao Chen, Chengduo Zhao, Chao Yu, Fandong Meng
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28062v1 Announce Type: new Abstract: Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrieved evidence only as text and omit tool-returned images from subsequent ...
59. Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning ​
Author: Zilin Zhao, Han Yang, Tianpei Yang, Fangsheng Huang, Yanfei Cui, Kan Peng, Yi Li, Yiming Zong, Hao Zhang, Yinsong Xue
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28065v1 Announce Type: new Abstract: Complete your ad view and grab a 5-cent bonus! In incentivized advertising, a platform promises users a bonus before observing downstream ad revenue, encouraging them to click and complete ads. It must balance the incentive promised in advance against ...
60. SEPO: Evidence-Grounded Prompt Optimization via Structural Editing ​
Author: Xiaoyu Ma, Haoyue Liu, Yiwen Li, Jionghao Zhu, Zhichao Wang, Ye Chen, Xiaoying Tang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28067v1 Announce Type: new Abstract: Existing API-only prompt optimisers are often described as interpretable, but in practice, this usually means only post-hoc inspectability: each iteration still rewrites the prompt as one opaque string, leaving a trace of full-prompt diffs rather than ...
61. Speculative Probing: LLM Monitoring at Speculative-Decoding Cost ​
Author: Collin Zhang, Tingwei Zhang, Vitaly Shmatikov
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.28099v1 Announce Type: new Abstract: Real-time classification during language model inference is valuable for safety filtering, behavioral analysis, and model monitoring, but current approaches force a trade-off between accuracy and efficiency. Hidden-state probes are fast but limited: th...
62. The Shape of Power: A Multilingual Framework for Social Power Reasoning in Dialogues ​
Author: Farah Atif, Sougata Saha, Monojit Choudhury
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28144v1 Announce Type: new Abstract: Social power plays a fundamental role in shaping human interaction, yet computational studies of power remain limited to narrow linguistic and cultural settings. Existing datasets further lack the demographic and relational depth needed for robust cros...
63. Under-Mattress Temporal Sensing for Next-Day Agitation Risk Scoring in Dementia Wards ​
Author: Zhen Liu, Marta Bono, Robbe Decloedt, Ajda Flisar, Maarten Van Den Bossche, Maarten De Vos
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.28152v1 Announce Type: new Abstract: Agitation fluctuates over short time horizons in people living with dementia, yet continuous physiological information for anticipating next-day risk is limited. We assessed whether contactless under-mattress signals from the preceding night inform nex...
64. CrabOS: An Operating System for Human-AI Co-inhabitation ​
Author: Qi Yang, Yun Ma
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.OS
arXiv:2608.28165v1 Announce Type: new Abstract: AI agents are evolving into long-running computational entities that can invoke tools, maintain memory, and complete complex tasks across applications. In real-world settings, completing a task often requires humans and AI to take turns leading its exe...
65. Expert Knowledge & Machine Understanding: Bridging Reactome's Ontology with LLM Semantic Embeddings ​
Author: Susanna Bravi, Riccardo De Luca, Rosa Sicilia, Christine Nardini, Mario Santoro
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28178v1 Announce Type: new Abstract: Biological knowledgebases like Reactome provide high-quality pathways that include biological elements' relationships and textual descriptions (metadata). The quality of such pathways is granted by manual curation, that presents, however, significant s...
66. Generative AI Alignment with Hinduism's Theological Plurality and Sacred Representation ​
Author: Dipto Das, Arpita Kundu, Nusrat Jahan Mim, Shion Guha, Syed Ishtiaque Ahmed
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.HC
arXiv:2608.28228v1 Announce Type: new Abstract: Generative AI systems are increasingly used to answer personal questions and mediate everyday practices, including religion. However, existing discussions around AI alignment and ethics have largely centered secular, Western, and Abrahamic assumptions ...
67. Stay Within Your Bounds: Distance-Guided Decoding for Guaranteed Context-Free Grammar Compliance ​
Author: Vincenzo Collura, Karim Tit, Eleonora Giunchiglia, Mike Papadakis, Maxime Cordy
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.FL, cs.LG
arXiv:2608.28229v1 Announce Type: new Abstract: Grammar-constrained decoding helps large language models produce syntactically valid structured outputs, such as code, JSON, and SQL. For context-free grammars, many practical decoders enforce local prefix feasibility: each token must keep the current ...
68. REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features ​
Author: Kai-Xuan Ding, Hao-Xiang Xu, Ji-Hua Peng, Zi-Qi Chen, Jiaqi Wang, Zhen-Hua Ling
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28233v1 Announce Type: new Abstract: Steering with Sparse Autoencoders (SAEs) offers a lightweight inference-time path for adapting the behavior of large language models without retraining. By exposing sparse and interpretable features, SAE steering provides a promising interface for safe...
69. Beyond Task-Only Matching: Personalized Skill Routing with Counterfactual Evaluation ​
Author: Tianle Wang, Yanghe Zou, Xiang Liu, Ziyao Huang, Chenchen Fu, Weiwei Wu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28241v1 Announce Type: new Abstract: The rapid expansion of reusable skill repositories makes skill routing a critical capability for large language model (LLM) agents. Existing methods treat routing as task-only semantic matching. However, when users with incompatible constraints issue a...
70. Regime-Aware Portfolio Management via Retrieval-Augmented LLM-Guided Expert Switching ​
Author: Ahmad Asadi, Reza Safabakhsh
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28252v1 Announce Type: new Abstract: Financial markets are inherently non-stationary, making the effectiveness of individual portfolio-management strategies highly dependent on changing market conditions. This work proposes a retrieval-augmented expert-switching framework that dynamically...
71. Physics-Guided Flow Matching for CT Image Reconstruction ​
Author: Davide Evangelista
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.28256v1 Announce Type: new Abstract: Deep generative models have recently emerged as powerful priors for solving ill-posed inverse problems in CT, with diffusion-based approaches achieving state-of-the-art reconstruction performance. However, diffusion models typically rely on stochastic ...
72. Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration ​
Author: Xiaoqing Wang, Keman Huang, Bin Liang, Hongyu Li, Xiaoyong Du, Wuqiong Pan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28264v1 Announce Type: new Abstract: Multi-agent systems (MAS) powered by large language models have shown promise for complex tasks but suffer from high failure rates. Current self-reflection methods for MAS require all agents to reflect upon failure, overlooking a critical reality: fail...
73. RECAST: Recent & Context-Aware Sampling for Test-Time Adaptation in Streaming Biosignals ​
Author: Yong-Yeon Jo, Junho Song, Joon-myoung Kwon
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28271v1 Announce Type: new Abstract: Streaming biosignals vary across subjects and drift over time, so population-trained models lose accuracy during long-term monitoring. Test-time adaptation (TTA) enables online personalization by updating the model on incoming samples. But in a stream,...
74. LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering ​
Author: Yi Wang, Haopeng Zhang, Chengxiang Huang, Rui Dai, Kaikui Liu, Piotr Koniusz, Xiangxiang Chu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28281v1 Announce Type: new Abstract: Loop Engineering is emerging as a practice for organizing development work around coding agents. Instead of writing each prompt by hand, practitioners design loops that monitor progress, assign work, run checks, and decide what the agent should do next...
75. Memristive-Friendly Hadamard Reservoir Computing: Structured, Multiplier-Free Recurrences at Scale ​
Author: Andrea Ceni, Gianluca Milano, Carlo Ricciardi, Claudio Gallicchio
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28295v1 Announce Type: new Abstract: Reservoir Computing (RC) designs Recurrent Neural Networks around a fixed, i.e., untrained, recurrent layer, and is a natural candidate for neuromorphic hardware. Memristive-friendly reservoirs derive the neuron dynamics from memristive-device kinetics...
76. MAIL: Memory-driven, Adaptive, Incremental, and Literature-grounded Framework for Hypothesis Generation in Chemistry ​
Author: Mahdi Babaei, Xueshen Li, Yutao Kuang, Jolene P. Reid, Yu Gan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28315v1 Announce Type: new Abstract: The ever-expanding volume of the chemical literature offers unprecedented opportunities to generate novel and impactful hypotheses. However, the bottleneck lies in efficiently navigating this vast knowledge base to formulate high-quality, experimentall...
77. Real-Valued Hyperdimensional Sequence Representations with Hadamard Product Binding and Shift Equivariance ​
Author: Kenny Schlegel, Dmitri A. Rachkovskij, Denis Kleyko, Amy Loutfi, Stefan Streif, Evgeny Osipov
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28334v1 Announce Type: new Abstract: Encoding temporal order is a fundamental requirement for sequence representations in Hyperdimensional Computing. Fractional Power Encoding provides similarity-preserving position vectors whose inner products approximate shift-invariant kernels, and it ...
78. AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents ​
Author: Pengze Li, Cui Tao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28345v1 Announce Type: new Abstract: AGENT-O is a modular ontology framework that defines a semantic Agent Card for representing health-oriented AI agent systems and supports assessment of reporting completeness in scientific publications. AGENT-O was developed as an OWL 2/RDF ontology co...
79. Propagating construction-time knowledge quality into medical question answering: A framework grounded in clinical guidelines ​
Author: Jie Hu, Junjie Wang, Shan Lu, Yifang Hu, Gong Cheng, Yun Liu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28360v1 Announce Type: new Abstract: Large language models have facilitated knowledge graph (KG) construction from clinical guidelines, but extracted triples vary in structural validity and evidential support. Meanwhile, graph-augmented question answering (QA) systems typically optimize q...
80. GRACE:Gradient-guided Coreset Selection for LLM Unlearning ​
Author: Praveen Bushipaka, Andrea D'Angelo, Lucia Passaro, Tommaso Cucinotta
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.28361v1 Announce Type: new Abstract: Machine Unlearning methods for Large Language Models typically assume pre-specified forget and retain sets. In realistic settings, however, requests may provide only a few examples of undesired behavior, requiring forget and retain sets to be inferred ...
81. EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses ​
Author: Tanmay Sah, Dolly Sah, Harshul Jain, Tanya Sah
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28363v1 Announce Type: new Abstract: LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot be safely reversed in sta...
82. MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places ​
Author: Jason Armitage, Ioannis Tsochantaridis, Linda Mazzone, Chuqiao Yan, Srini Narayanan, Sarah Ebling
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28384v1 Announce Type: new Abstract: We introduce MAP, the first benchmark to evaluate multimodal AI systems as assistants for users with accessibility requirements when planning visits to places in the real world. In our evaluation, systems are presented with requests to verify or recomm...
83. Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation ​
Author: Akshay Kekuda, Shreeranjani Srirangamsridharan, Ishan Bhatt, Yanan Cao, Sinduja Subramaniam, Evren Korpeoglu, Kaushiki Nag, Kannan Achan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.28393v1 Announce Type: new Abstract: Repurchase recommenders in e-commerce are commonly framed as a binary question asking "will this customer buy this item within W days", a formulation that requires a separately trained model for every horizon of interest. We replace this stack with sur...
84. RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents ​
Author: Yupeng Zhang, Liuyuan Jiang, Hongyi Huang, Bingheng Li, Lisha Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, q-fin.TR
arXiv:2608.28399v1 Announce Type: new Abstract: In financial markets, a sequential policy that reacts systematically to price movements may become predictable to other market participants. This paper studies whether large language model (LLM) agents exhibit such directional structure through RetailA...
85. VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC Filings ​
Author: Menghan Liu, Elynn Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28402v1 Announce Type: new Abstract: Across audit applications, judgments must be supported by reasonable evidence. However, standard financial language models prioritize fluency over evidence. They are built for general financial reasoning and may produce plausible but ambiguous answers,...
86. Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs ​
Author: Vishvesh Bhat
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28421v1 Announce Type: new Abstract: Post training a language model to reason means updating its weights. Supervised finetuning and reinforcement learning both place the acquired capability inside the model where it cannot be inspected cannot be checked step by step and cannot be moved to...
87. Prove2Me: An Open Collaborative Platform for Scaling Math Formalization ​
Author: Shuze Chen, Kunal Marwaha, Xiaoyang Lu, Henry Yuen, Tianyi Peng
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LO, cs.MA
arXiv:2608.28433v1 Announce Type: new Abstract: Proof assistants such as Lean 4 promise the paradigm of formally verified mathematics, but large-scale formalization projects have faced major barriers to entry, including the need for expertise in formal verification (as well as the underlying mathema...
88. Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning ​
Author: Minghui Xu, Zi Wang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28447v1 Announce Type: new Abstract: Current large language models (LLMs) increasingly benefit from external tool integration, especially for tasks requiring reliable computation and verification. Motivated by this, we study calculator tool calling for improving mathematical reasoning on ...
89. COVER: Identifiable Evaluation of Coalition Routing ​
Author: Raghul Sugumar, Amrit Gopinath
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28475v1 Announce Type: new Abstract: When a multi-agent system changes its team, it also changes the messages and final answer it produces, so an end-to-end accuracy gap does not by itself identify a routing effect. We introduce method, an evaluation contract that fixes a public informati...
90. AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction ​
Author: Yafei Zhang, Nan Wu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.28491v1 Announce Type: new Abstract: Predicting robot videos requires both precise motion reasoning and preservation of high-frequency appearance, yet monolithic pixel models entangle these objectives and often conceal their progress behind a strong last-frame baseline. We present AcrossV...
91. Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration ​
Author: Simeng Sun, Roger Waleffe
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28511v1 Announce Type: new Abstract: When training Mixture-of-Experts (MoE) language models with expert parallelism, all-to-all token dispatch and combine collectives can consume a substantial fraction of end-to-end training time. In this work, we study communication-efficient MoE models ...
92. When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI ​
Author: Sihan Jia, Oliver Lemon
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.RO
arXiv:2608.28518v1 Announce Type: new Abstract: We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodied AI (EAI) models. We find that ASR errors can lead to harmful instructions being accepted and executed by EAI models, thereby reducin...
93. InstructMesh: Selective Refinement of Generative 3D Models for Fabrication ​
Author: Faraz Faruqi, Ahmed Katary, Demircan Tas, Theresa Hradilak, Ning Zhang, Jiaji Li, Fabian Manhardt, Martin Nisser, Vrushank Phadnis, Ruofei Du, Federico Tombari, Megan Hofmann, Stefanie Mueller
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.28534v1 Announce Type: new Abstract: Recent advances in generative AI allow users to create 3D models from text or images. However, these models prioritize visual plausibility over geometric accuracy, often generating results with flaws that compromise their intended use post-fabrication....
94. Logos: An Agent Harness on a Cross-Process Bus ​
Author: Hanzhang Jia, Liheng Zeng, Hao Cheng, Yi Gao, Bo Ma
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.28553v1 Announce Type: new Abstract: Modern agent systems assemble capabilities at runtime, and this dynamic composition has recently received a complete formal treat ment in the spatiotemporal-composability calculus, in which a capability is a component carrying a tracked inverse, and ag...
95. SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction ​
Author: Nilay Yilmaz, Naga Sai Abhiram Kusumba, Stella Wenxing Liu, Yezhou Yang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.27461v1 Announce Type: cross Abstract: Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts. This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capt...
96. Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech ​
Author: Han Wang, Yuhu Cheng, Xuesong Wang, Yi Zhu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27462v1 Announce Type: cross Abstract: Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and contextual hints, making its detection in online content review challenging. While existing PLM- or LLM-bas...
97. The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models ​
Author: Cheolho Shin, Yoojin Han, Donghun Shin, Kunho Lee
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.HC
arXiv:2608.27465v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice according to the user's emotional state has become an important safety problem. We test whether emotional ex...
98. PACE: Publisher-Adaptive Content Extraction via Agentic Automation ​
Author: Zhanlin Liu, Munirathnam Srikanth
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27466v1 Announce Type: cross Abstract: Web content extraction is essential for reliable LLM data pipelines, yet existing methods often struggle to jointly satisfy accuracy, scalability, and adaptability. General-purpose extractors can be applied broadly, but they are often brittle on publ...
99. UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering ​
Author: Mohammad Arvan, Hossein Haeri, Natalie Parde, Rebecca T. Feinstein
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2608.27467v1 Announce Type: cross Abstract: We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment...
100. Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection ​
Author: Fina Polat, Daniel Daza, Pengyu Zhang, Klim Zaporojets, Paul Groth
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB
arXiv:2608.27470v1 Announce Type: cross Abstract: Entity Disambiguation (ED) is a key task for constructing and using knowledge graphs. State-of-the-art neural approaches commonly model ED as a single task, although it consists of two distinct subproblems: retrieving candidate entities and selecting...
101. XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering ​
Author: Iman Barati, Arash Ghafouri, Behrouz Minaei-Bidgoli
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27481v1 Announce Type: cross Abstract: Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually translate an entire example into one language. This hides failures at language boundaries inside the...
102. A Survey on Rubric-Guided Reinforcement Learning for Language Models ​
Author: Zifei Shan, Fangning Shao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27505v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) with human preferences. However, traditional RLHF relies on scalar reward signals that lack interpretability and fail to capt...
103. Marginal Coverage Credit Reduces Redundant Exploration in Parallel State-Entropy Optimization ​
Author: Junhao Cao, Hongyi Xia, Jianian Wu, Xiaopeng Yi, Lixia Huang, Ping Guo
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.27507v1 Announce Type: cross Abstract: Policy Gradient for Parallel State Entropy maximization (PGPSE) expands state-space coverage by training independently parameterized policies in replicated copies of the same environment. However, its pooled team-entropy score measures only collectiv...
104. Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap ​
Author: Jacopo Dardini, Claudio Stanzione, Giordano Col`o, Giuseppe Fenza
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CR
arXiv:2608.27512v1 Announce Type: cross Abstract: Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models. When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluat...
105. DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization ​
Author: Tao Zhang, Jianchao Tan, Pingwei Sun, Yanqi Yu, Zixu Jiang, Yuchen Xie, Xunliang Cai, Ziqian Zeng
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.27513v1 Announce Type: cross Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length. Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the...
106. Trajectory-Level Speculative Decoding for Diffusion Language Models ​
Author: Tianxiang Pan, Baitao Gong, Mo Guang, Hongwei Yong, Tianpeng Jiang, Yaqian Li, Zheng Cao, Kaiwen Long
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27514v1 Announce Type: cross Abstract: Diffusion-based language models (dLLMs) enable parallel token generation through iterative denoising, but existing decoding strategies collapse to single-token generation under low confidence, severely limiting throughput. Unlike autoregressive model...
107. Destroy Me: Automatic Artifact Generation for Histopathology Images ​
Author: Zuzanna Krawczyk-Borysiak, Adam Krawczyk, Mateusz Miller, Gabriela Kaczmarek, S{\l}awomir Paku{\l}o, Ma{\l}gorzata Sok'o{\l}, .Zaneta Swiderska-Chadaj
Published: 8/31/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG
arXiv:2608.27516v1 Announce Type: cross Abstract: Deep learning's diagnostic utility in pathology is constrained by model vulnerability to real-world data imperfections. While current strategies favor "perfect data" by filtering low-quality regions, which can lead to the loss of valuable diagnostic ...
108. FVeinSyn: Synthetic Finger Vein Image Generator ​
Author: Yifan Wang, Jie Gui, Adams Wai Kin Kong, Baosheng Yu, Changsheng Chen, Qi Li, Zhenan Sun, James Tin-Yau Kwok, Alex Kot
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27527v1 Announce Type: cross Abstract: A major challenge in finger vein recognition is the lack of large-scale public datasets. Existing datasets contain few identities and limited samples per finger, restricting the advancement of deep learning-based methods. To address this, we propose ...
109. Self-Explainable Multi-Label Graph Neural Network for Correlated Evidence Attribution ​
Author: Yingqi Feng, Yufei Tang, Min Shi, Xingquan Zhu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.27574v1 Announce Type: cross Abstract: Multi-label graph learning intends to capture the intrinsic complexity of real-world applications, where one sample is often related to multiple groups or consists of multiple objects. To date, a handful of multi-label graph learning methods exist, b...
110. Quanta Perception as Probabilistic Events ​
Author: Varun Sundar, Pavan Thodima, Sacha Jungerman, Mohit Gupta
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27584v1 Announce Type: cross Abstract: Autonomous systems rely on extracting information from light, yet remain brittle in extreme environments, from nighttime navigation to high-speed robotics. Conventional sensors aggregate photons over fixed exposures, imposing trade-offs between sensi...
111. PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models ​
Author: Davood Soleymanzadeh, Kaidi Zhang, Zhiyuan Zhang, Bihao Zhang, Xiao Liang, Yu She, Minghui Zheng
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.27609v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) have shown strong promise for general-purpose robotic manipulation by mapping language instructions and vision observations directly to actions. However, most VLAs primarily condition action prediction on current ...
112. Tensor-Accelerated Eager Multi-Resolution Grids for Evolving Large-Scale Substrates ​
Author: Romain Claret, Michael O'Neill, Paul Cotofrei, Kilian Stoffel
Published: 8/31/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.LG
arXiv:2608.27612v1 Announce Type: cross Abstract: In neuroevolution, indirect encoding generates neural network connectivity from a compact genome rather than specifying each connection. ES-HyperNEAT automatically discovers where to place hidden nodes by examining CPPN output patterns: it recursivel...
113. LitCurate: A Configuration-Driven AI-Assisted Framework for Scientific Database Construction with an Application to Lower-Mantle Equation-of-State Data ​
Author: Abin Shakya, Wilson Samuels, Dominica Wilson, Gioia A. Marchi, Israa Draz, Chenxing Luo, Renata M. Wentzcovitch
Published: 8/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, physics.geo-ph
arXiv:2608.27629v1 Announce Type: cross Abstract: The growing scientific literature contains decades of experimental and computational results that could support data-driven and physics-based modeling, yet much of this infor- mation remains locked in publications and is not readily usable for large-...
114. Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge ​
Author: Md Monjurul Ahsan Prodhan, Md Nour Hossain
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.27633v1 Announce Type: cross Abstract: Pothole detection and its severity measurement is still an important challenges in urban infrastructure management, where late maintenance directly contributes to vehicle damage, road accidents, and escalating repair costs. Existing automated approac...
115. Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification ​
Author: Alexandre L. M. Levada
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, stat.ML
arXiv:2608.27634v1 Announce Type: cross Abstract: Nearest neighbor classification relies fundamentally on how locality is defined, yet conventional $k$-NN imposes the same neighborhood cardinality throughout the feature space. This assumption can be inadequate for data whose local geometry varies su...
116. Knowing Before Answering: Decoding Language Models for Reliable RAG ​
Author: Syed Mahbubul Huq, Christopher Child, Tillman Weyde, Pranava Madhyastha
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27661v1 Announce Type: cross Abstract: In Retrieval-Augmented Generation (RAG), retrieval may provide insufficient or conflicting information needed to answer a question. The system should not only know when to answer but also be able to identify cases in which the documents provided in R...
117. Semantic Watermarking with Order-Robust Detection over Sub-sentence Units ​
Author: Abdulrahman Diaa, Jonathan Petit, Florian Kerschbaum
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.27666v1 Announce Type: cross Abstract: Semantic watermarks tie the mark to sentence meaning rather than token choices, promising robustness to content-preserving edits. However, the detector only observes attacker-supplied text, which can be reworded, reordered, or resegmented to evade de...
118. First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents ​
Author: Syed Mahbubul Huq, Pranava Madhyastha
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27672v1 Announce Type: cross Abstract: We present Qwen-GuidePlay-2B, a 2B-parameter language model for dialogue-game interaction. We fine-tune Qwen3.5-2B using three steps: a) SFT on only successful game trajectories from Playpen, b) weighted turn-level SFT, and c) teacher-guided SFT. The...
119. CARDINAL Predicts Cardiovascular Risk From Non-contrast Cardiac CT ​
Author: Roy Gabriel, Nattakorn Kittisut, Jamshid Hassanpour, Michael Galarnyk, Abanoub Abdelmalak, Marly van Assen, Carlo N. De Cecco, Arshed Quyyumi, Ali Adibi
Published: 8/31/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG, cs.RO
arXiv:2608.27690v1 Announce Type: cross Abstract: Cardiovascular risk prediction remains limited by incomplete clinical data and imaging biomarkers that reduce computed tomography (CT) to a small number of handcrafted features. We developed CARDINAL (Cardiovascular Assessment via Representation lear...
120. Evaluating Loss Functions in Differentiable Out-of-Domain Sound-Matching with Partial Parameter Distance ​
Author: Amir Salimi, Daniel Penner, Kalvin Eng, Abram Hindle, Osmar R. Za"iane
Published: 8/31/2026, 4:00:00 AM
Categories: cs.SD, cs.AI
arXiv:2608.27698v1 Announce Type: cross Abstract: In out-of-domain (OOD) sound-matching, a synthesizer is optimized to mimic a sound it did not generate. OOD evaluation of loss functions is underexplored in part because the standard "parameter loss" metric requires a shared parameter space between t...
121. RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing ​
Author: Madhusudan Srinivasan, Namith Nishal Raphae
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SE
arXiv:2608.27704v1 Announce Type: cross Abstract: When machine learning classifiers are retrained, inputs correctly classified by the previous model version may be misclassified by the updated version, creating regression faults that are costly to detect because verifying predictions against ground ...
122. Efficient Auto-Interpretability of AI Models in Biology ​
Author: Piotr Jedryszek, Oliver M. Crook
Published: 8/31/2026, 4:00:00 AM
Categories: q-bio.QM, cs.AI
arXiv:2608.27754v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs), and other interpretability methods could turn AI models in Biology and other fields into engines of scientific discovery by explaining the superhuman capabilities of those models. However, a latent is only useful if we kno...
123. Beyond Search-Imitation: Prior-Directed Exploration for Searchless Chess ​
Author: Szymon Mi{\l}osz, Piotr Duch, Szymon Grabowski
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.27757v1 Announce Type: cross Abstract: Searchless chess networks reach human master strength from a single forward pass by imitating a stronger teacher: the strongest, Leela Chess Zero's (Lc0) released Chessformer, distills the visit counts of an AlphaZero-style Monte Carlo Tree Search (M...
124. Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict ​
Author: Adarsh Sudheer, David Li, Omar Elbanna, Ishaan Kodarapu, Arjun Bahuguna, Vasu Sharma
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27785v1 Announce Type: cross Abstract: We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2-7B-AV, three alignment ...
125. How Much Can AI Understand? Toward AI-Assisted Sensemaking of Collaborative Discussion in Groups with Shared History ​
Author: Soobin Cho, Mark Zachry, David W. McDonald
Published: 8/31/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.27799v1 Announce Type: cross Abstract: AI tools that support collaborative discussion typically treat the discussion as a standalone task, focusing only on its content and setting aside the social context of the group having it. But it is groups with a shared history, with their own norms...
126. ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools ​
Author: Yuqi Jia, Ruiqi Wang, Patrick Li, Yuepeng Hu, Peinian Li, Neil Gong
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.27800v1 Announce Type: cross Abstract: Exfiltrating an LLM agent's runtime context -- such as the user prompt, execution trajectory, and tool list -- poses severe security and privacy risks to users. Such attacks can be carried out via malicious tools and typically require three condition...
127. Actionable CBFI: Integrating Structural Decomposition and Causal Counterfactual Recourse for Tabular Machine Learning ​
Author: Sejong Oh
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.27821v1 Announce Type: cross Abstract: Explainable artificial intelligence (XAI) increasingly calls for actionable counterfactual recourse, yet current methodologies face challenges related to causal invalidity, excessive cognitive burden, and predictive failure. Exhaustive causal search ...
128. FISGuard: Defending Against Membership Inference via Fixed Input Subspaces ​
Author: Haocheng Jiang, Hua Shen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.DC
arXiv:2608.27836v1 Announce Type: cross Abstract: As large language models are increasingly adopted in federated learning, protecting user privacy while performing parameter-efficient fine-tuning on distributed private data has become an important challenge. Although clients only share gradients ins...
129. FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling ​
Author: Jun Bai, Ruilin Wang, Yue Li
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA
arXiv:2608.27856v1 Announce Type: cross Abstract: Recent advances in large language models are enabling autonomous clinical agents to perform increasingly complex electronic health record (EHR) modeling workflows. However, agents deployed at individual hospitals remain constrained by institution-spe...
130. From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation ​
Author: Rit Gangopadhyay, Alex Wong
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27860v1 Announce Type: cross Abstract: Vision foundation models are capable of generalizing across 3-dimensional (3D) scenes with high-fidelity estimates; their empirical success can be attributed to training on large-scale datasets of perspective images. However, when transferred to wide...
131. SOMTab: Set-Order Mamba for Efficient Tabular In-Context Learning ​
Author: Hao Wang, Siyu Zhang, Wei Ma
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.27882v1 Announce Type: cross Abstract: Tabular foundation models based on in-context learning have recently emerged as strong alternatives to task-specific model fitting. However, the current performance frontier remains dominated by attention-heavy architectures, where attention is used ...
132. OpenStamp: A Watermark for Open-Source Language Models ​
Author: Miroojin Bakshi, Saksham Rastogi, Danish Pruthi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.27899v1 Announce Type: cross Abstract: With the growing prevalence of large language model (LLM) generated content, watermarking is considered a promising approach for attributing text to LLMs and distinguishing it from human-written content. A prominent class of techniques embeds subtle ...
133. LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages ​
Author: Injun Baek, HyeongSeok Lee, Yearim Kim, Junhoo Lee, Nojun Kwak
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.27902v1 Announce Type: cross Abstract: Landing pages are goal-oriented web interfaces that must communicate a target-specific value proposition while organizing information flow, visual hierarchy, and calls to action (CTA). Although large language models can generate plausible webpage cod...
134. Low-Altitude Fluid Antenna Network with Multi-Agent Reinforcement Learning ​
Author: Tong Zhang, Yanfei Su, Shuai Wang, Wanli Ni, Chengzhong Xu, Huseyin Arslan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, math.IT
arXiv:2608.27909v1 Announce Type: cross Abstract: Low-altitude wireless networks (LAWNs) integrate terrestrial and aerial platforms to provide ubiquitous communication, sensing, and localization services for unmanned aerial vehicles (UAVs) and electric vertical takeoff and landing (eVTOL) aircraft. ...
135. PCBnet: A Dataset and Automatic Construction of SPICE Netlists from Schematic Images ​
Author: Zhen Huang, Yuhao Gao, Yuzhi Liu, Daian Cheng, Chengyuan Shao, Yucheng Chen, Yongjian Jia, Futing Zhang, Yichen Shi, Wenhao Wang, Zuyan He, Yangbo Wei, Zhanfei Chen, Jinlong Yan, Yu Zhang, Haoying Wu, Ting-Jung Lin, Lei He
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27923v1 Announce Type: cross Abstract: Printed circuit boards (PCBs) are fundamental to modern electronic systems, yet AI-driven PCB design automation remains constrained by the lack of large-scale paired schematic-netlist datasets. PCB schematics are particularly challenging due to diver...
136. Antipatterns in AI-assisted Qualitative Data Analysis: A Catalog of Temptations and Pitfalls for Software Engineering Researchers ​
Author: Rashina Hoda, Carolyn Seaman, Victoria Gomes, Rodrigo Spinola
Published: 8/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.27927v1 Announce Type: cross Abstract: AI-assisted qualitative data analysis (QDA) offers unprecedented opportunities to streamline software engineering (SE) research, yet uncritical use risks compromising analytical rigor and flooding the field with accelerated production of low-quality ...
137. Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification ​
Author: Cameron Wilding, Mina Shaker, Fatemeh Ganji
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2608.27954v1 Announce Type: cross Abstract: Post-deployment changes to large language models can alter behavior while leaving routine outputs largely unchanged, creating a challenge for AI governance when model weights are proprietary. We present a privacy-preserving zk-SNARK-based audit frame...
138. CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks? ​
Author: Zi Liang, Xiaoyu Xu, Yanyun Wang, Minxin Du, Qingqing Ye, Haibo Hu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.27990v1 Announce Type: cross Abstract: Prompt injection attacks on Large Language Model (LLM) agents seek to introduce malicious instructions or content into external text sources retrieved by agents, forcing the underlying LLMs to execute harmful actions outside their benign scope. While...
139. A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation Constraint ​
Author: Artem Safronov
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28003v1 Announce Type: cross Abstract: This paper proposes a layer bit allocation method for Gemma-3-1B, formulating the problem as performance maximization (latency decrease) given a degradation budget constraint (allowable level of generation quality loss). This approach is different fr...
140. When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood? ​
Author: Yansen Han, Hongxin Sun, Tao Lin
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28010v1 Announce Type: cross Abstract: Flow matching enables likelihood-free training, yet alignment methods increasingly reuse conditional flow matching (CFM) losses as endpoint negative log-likelihoods (NLLs) and their old/new differences as log-likelihood ratios. We characterize when t...
141. Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning ​
Author: Vy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He, Feng Xia, Renqiang Luo, Erik Cambria, Xiuzhen Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.28018v1 Announce Type: cross Abstract: Knowledge-intensive reasoning requires Large Language Models (LLMs) to ground answers in provided evidence. When evidence is insufficient, it is desirable that models abstain rather than confidently generating unsupported answers. Existing abstention...
142. Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code ​
Author: Animesh Shaw
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.MA, cs.SE
arXiv:2608.28021v1 Announce Type: cross Abstract: Large language models are increasingly used to author Infrastructure-as-Code (IaC), where a single insecure default can be deployed directly into production. Prior evaluations report raw vulnerability counts for model-generated IaC, but without a hum...
143. SimpCue: Cue-Based Prompting for Multilingual Text Simplification ​
Author: Mehrzad Tareh, Horacio Saggion, Stefan Bott
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28042v1 Announce Type: cross Abstract: Text simplification aims to make complex texts easier to understand while preserving their original meaning. Recent large language models can perform simplification through prompting, but it remains unclear whether adding explicit linguistic informat...
144. Explainable Uncertainty Estimation for Reliable Medical AI ​
Author: Li Rong Wang, Jamie Duell, Xinran Xu, Thomas C. Henderson, Yu Yue Hew, Pik Wan Erica Chiang, Xiao Wei Alstar Ang, Bingwen Eugene Fan, Xiuyi Fan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28052v1 Announce Type: cross Abstract: Artificial intelligence has strong potential to support clinical decision-making, yet its adoption in healthcare remains limited due to a lack of trust. Uncertainty estimation can signal unreliable predictions, and explainable AI (XAI) can clarify ho...
145. Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models ​
Author: Kairong Yu, Zixin Zhu, Le Yu, Hongwei Wang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28058v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) remain prone to hallucinations, producing responses that are irrelevant or inconsistent with the multimodal input. Existing mitigation methods mainly rely on external supervision, output calibration, or attention ...
146. VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D Gaussians ​
Author: Ruijie Su, Lingxiao Yang, Xiaohua Xie, Jianhuang Lai
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28069v1 Announce Type: cross Abstract: Recent progress has been made in 3D Gaussian representation for reconstruction, generation, and physical simulation. However, current approaches mainly concentrate on physics-based dynamic generation of solid objects and only handle single-phase coll...
147. Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations ​
Author: Naren Akash, Neeraja Ramanan
Published: 8/31/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG
arXiv:2608.28092v1 Announce Type: cross Abstract: Interpreting a CT scan means comparing structures on either side, judging how far apart organs sit, and knowing where each one belongs. Medical vision encoders are evaluated on diagnostic accuracy, or through assembled multimodal systems where a fail...
148. VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning ​
Author: Pengcheng Li, Zhengyang Zhang, Dongxu Zhang, Sui Huang, Shaohua Ma
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28128v1 Announce Type: cross Abstract: Fine-grained credit assignment is a central challenge in reinforcement learning for long horizon LLM agents. Standard objectives often train from programmatically verifiable terminal rewards by broadcasting each sparse outcome to every action in a tr...
149. CheXtriev: Anatomy-Centered Representation for Case-Based Retrieval of Chest Radiographs ​
Author: Naren Akash, Arihanth Tadanki, Jayanthi Sivaswamy
Published: 8/31/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG
arXiv:2608.28137v1 Announce Type: cross Abstract: We present CheXtriev, a graph-based, anatomy-aware framework for chest radiograph retrieval. Unlike prior methods focussed on global features, our method leverages graph transformers to extract informative features from specific anatomical regions. F...
150. Post-Edit Re-Verification in Simulator-Backed Engineering Agents: A Controlled Comparison of Verification-Cadence Guidance ​
Author: Qingchuan Zhu, Shuyue Tong, Pengju Ren
Published: 8/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.28147v1 Announce Type: cross Abstract: Engineering agents that interact with external simulators may need to coordinate design modification with reacquisition of engineering evidence for the modified state. We ask whether first post-edit re-verification changes when explicit verification-...
151. The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension ​
Author: Yuhe Sui, Jianing Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28150v1 Announce Type: cross Abstract: Which geometry controls the rank complexity of normalized softmax attention? We study maximum-row-$\ell_1$ approximation rank, exactly the least unrestricted rank preserving every bounded vector-valued output. Two sharp worst-case laws isolate suppor...
152. Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result ​
Author: Christos Koutsiaris
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2608.28151v1 Announce Type: cross Abstract: A byte-level BPE tokenizer is an ordered list of merge rules, so applying only a prefix yields a vocabulary whose token identifiers are the first rows of the full vocabulary. This prefix nesting allows one language model to operate at several vocabul...
153. Gen-TAS: A Generative AI-Aided Hardware-Software Task Allocation Framework for FPGA-GPP Heterogeneous Systems ​
Author: Mary Kong, Yuqin Zhao, Semih Vazgecen, Cristian Sestito, Themis Prodromakis
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.HC
arXiv:2608.28160v1 Announce Type: cross Abstract: FPGA-GPP heterogeneous systems combine software flexibility with the performance and energy efficiency of reconfigurable hardware. However, determining which application tasks should execute on the GPP or FPGA requires extensive expertise and design-...
154. Text Restoration of Ancient Documents with Language Models ​
Author: Shibingfeng Zhang, Edoardo Caraffa, Annafelicia Zuffrano, Maddalena Modesti, Giovanni Colavizza
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28170v1 Announce Type: cross Abstract: Purpose - This study investigates the feasibility of restoring missing text caused by physical lacunae in damaged ancient manuscripts using language models. Methodology - The study proposes different scenarios to replicate real-world conditions. Lang...
155. Conformal Risk-Averse Decision Making with Optimized Certainty Equivalent Risk Control ​
Author: Amirmohammad Farzaneh, Osvaldo Simeone
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.IT, cs.LG, math.IT
arXiv:2608.28179v1 Announce Type: cross Abstract: We study risk-averse decision making, in which an agent selects actions while being uncertain about the true system state. The risk is measured via optimized certainty equivalent (OCE) metrics, which generalize popular criteria such as mean-variance ...
156. Beyond Flat Netlist: Hierarchical Graph Representation Learning for Scalable Analysis of Sequential Circuits ​
Author: Jingyi Zhou, Zhengyuan Shi, Jiaying Zhu, Ziyang Zheng, Qiang Xu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR
arXiv:2608.28188v1 Announce Type: cross Abstract: Circuit Representation Learning (CRL) offers a powerful paradigm to guide and optimize core Electronic Design Automation (EDA) tasks, but its practical adoption is hindered by the immense scale of industrial netlists and a failure to explicitly model...
157. Performative Privacy: When Differential Privacy Maximizes Utility ​
Author: Uddalak Mukherjee, Edwige Cyffers, Yann Chevaleyre
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2608.28198v1 Announce Type: cross Abstract: Privacy-preserving learning is often motivated by the idea that protecting users' data can preserve trust and thus participation, improving utility in the long term. However, this claim has not been formalized so far. In parallel, performative learni...
158. Training-free Suction Grasp Detection for Deformed Aseptic Cartons Using Vision-Language Models and Geometric Surface Scoring ​
Author: Marin Maletic, Goran Vasiljevic
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.28246v1 Announce Type: cross Abstract: Robotic sorting of recyclable waste is challenging due to the deformable and geometrically inconsistent nature of target objects. We present a training-free suction grasping system for sorting deformed aseptic beverage cartons, decoupling target iden...
159. A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation ​
Author: Tadej Tomani\v{c}, Alice Baudhuin, Jan Soto\v{s}ek, Jure Brence, Pan\v{c}e Panov, Nikola Simidjievski, Dragi Kocev
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28247v1 Announce Type: cross Abstract: Change detection in Earth observation (EO) is critical for monitoring land surface transformations, yet recent research in the field is constrained by inconsistent evaluation protocols and a narrow focus on predictive accuracy without regard for comp...
160. Spatial-Semantic Reasoning using Large Language Models for Efficient UAV Search Operations ​
Author: Marin Maletic, Marijana Peti, Tamara Petrovic, Stjepan Bogdan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.28270v1 Announce Type: cross Abstract: We present a real-time semantic navigation framework for Unmanned Aerial Vehicles (UAVs) focused on improving time efficiency in the Object Goal Navigation (ObjectNav) task. Central to our approach is a Large Language Model (LLM) that interprets user...
161. Embedding Models for Stance-Aware Argument Retrieval ​
Author: Angelo Sparacino, Francesca Toni, Adam Dejl
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28283v1 Announce Type: cross Abstract: In computational argumentation, obtaining arguments that explicitly support or attack given claims is a critical precursor to downstream reasoning tasks. When these supporting and attacking arguments are to be retrieved using semantic search methods,...
162. A Probabilistic Interpretation of KV Cache Eviction ​
Author: Renato Geh, Alex Chen, Daniel Israel, Aditya Grover, Guy Van den Broeck
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28293v1 Announce Type: cross Abstract: The premise and promise of KV (cache) eviction is simple: higher throughput can be achieved by evicting some entries from the KV cache, at a negligible cost to quality. This holds empirically for many existing methods, though most rely on creative he...
163. MaCoPlanner: LLM-Assisted Manual-Compiled Task Planning with Proactive Safety Verification for Robotic Industrial Panel Operation ​
Author: Guipeng Xin, Jiahe Xua, Mohammad Deghat, Chenhui Wan, Jie Liu, Youmin Hu, Zhongxu Hu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.28300v1 Announce Type: cross Abstract: Robotic industrial panel operation requires not only accurate control localization but also compliance with operating procedures, safety rules, and device-state constraints distributed across heterogeneous manuals. This study presents MaCoPlanner, a ...
164. PanelShield: Verifiable Closed-Loop Safe Planning for Robotic Industrial Panel Operation ​
Author: Guipeng Xin, Jiahe Xu, Chenhui Wan, Jie Liu, Youmin Hu, Zhongxu Hu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.28305v1 Announce Type: cross Abstract: Industrial panel operation is knowledge-intensive and safety-critical. Beyond control recognition and action generation, execution must satisfy constraints in operation manuals and safety regulations. While foundation-model-based planners show strong...
165. VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation ​
Author: Zewen Ding, Zezhong Wu, Zhou Tao, Shida Wang, Shizhuo Hou, YongXiang Hua, Haoyu Cao, Linli Xu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.28306v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) improves reasoning by training a problem-only student on its own rollouts using dense token-level supervision from a privileged teacher that also sees a reference solution. However, standard OPSD treats the teacher ...
166. Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss ​
Author: Niccol`o Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan, J"org Franke, Sampo Pyysalo, Jenia Jitsev, Aaron Klein
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28308v1 Announce Type: cross Abstract: We study the scaling behavior of learning rate and batch size in pretraining dense large language models on English-prevalent corpora. Beyond scaling \textit{jointly optimal} learning rates and batch sizes, we investigate their \textit{marginal} evol...
167. Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers ​
Author: Abrar Alotaibi, Muhammad Shahid Jabbar, Sadam Al-Azani, Moataz Ahmed
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.28327v1 Announce Type: cross Abstract: Practitioners defend large language models (LLMs) by stacking defenses, assuming the layers compound. A stack is an ensemble, and ensembles compound only under a condition the LLM security literature recommends but never measures: the members must fa...
168. BanglaMed-QA: A Question Answering System for Healthcare Support in Bangla ​
Author: Rowzatul Zannat, Abdullah Al Shafi, K. M. Azharul Hasan, Atia Shahnaz Ipa
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.28329v1 Announce Type: cross Abstract: Medical question answering (QA) systems have become crucial tools for providing reliable health information. But they remain very unexplored for low-resource languages like Bangla due to limited datasets and systems tailored to these languages. To ad...
169. Cross-Spectral Dense Correspondence for Multimodal Spectral Medical Imaging ​
Author: Eric L. Wisotzky, Jost Triller, Simon W. H"artl, Oliver T. Bruns, Peter Eisert, Anna Hilsmann
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28341v1 Announce Type: cross Abstract: Precise dense correspondence is a fundamental prerequisite for multimodal spectral imaging systems that fuse disparate wavelength ranges for subsequent analysis in medical and scientific imaging. Corresponding image points are often observed with non...
170. Optimal Adversarial Testing: Extracting Honest Test Results from Dishonest Test Takers ​
Author: Owen Cox, April Xu, Weiyu Xu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.GT, eess.SP, stat.AP, stat.ME
arXiv:2608.28362v1 Announce Type: cross Abstract: In applications, it is often required to test objects or people to determine their qualities in terms of certain metrics. However, besides being naturally noisy, the test results can be corrupted by adversarial behaviors of objects or people being te...
171. Real-Time Musculoskeletal Surrogates for Pediatric Cerebral Palsy: a Credibility Pilot ​
Author: Mohammad Arif Ul Alam
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28371v1 Announce Type: cross Abstract: Real-time musculoskeletal (MSK) surrogates could support personalized rehabilitation for children with cerebral palsy (CP), but their credibility depends on subject-wise evaluation, low inference latency, and calibrated uncertainty. We develop a subj...
172. AI as Teammate: Rethinking Task Distribution in Medical Training ​
Author: Fendi Tsim, Alina Gutoreva, Anthony Weiss, Nicole Dubosh
Published: 8/31/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.28373v1 Announce Type: cross Abstract: Integrating Artificial Intelligence (AI), particularly generative AI, into medical training has prompted concerns about learner over-reliance, misuse, and erosion of foundational clinical competencies. We propose a conceptual reframing at the decisio...
173. When Linguistic and Internal Confidence Diverge in Large Language Models ​
Author: Hefan Zhang, Bingquan Zhang, Ming Cheng, Saeed Hassanpour, Weicheng Ma, Soroush Vosoughi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28382v1 Announce Type: cross Abstract: Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study this question across 8 classification tasks, 2 generation tasks and 3...
174. LongPIBench: A Long-Context Benchmark for Prompt Injection ​
Author: Yupei Liu, Yuqi Jia, Neil Zhenqiang Gong, Jinyuan Jia
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.28411v1 Announce Type: cross Abstract: Prompt injection attacks pose a serious security risk to large language models in real-world applications. However, existing prompt injection benchmarks primarily focus on short-context inputs, leaving the attacks and defenses in long-context setting...
175. Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL ​
Author: Jiayan Lin, Yujia Liu, Zijin Hong, Zheng Yuan, Yilin Xiao, Hao Chen, Qinggang Zhang, Xiao Huang, Feiran Huang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB
arXiv:2608.28432v1 Announce Type: cross Abstract: Recent advances in in-context learning (ICL) text-to-SQL have substantially improved execution accuracy on public benchmarks by assembling increasingly elaborate pipelines around the base generator, yet existing studies typically report aggregate end...
176. Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction ​
Author: Qing Ye, Meng-Hsuan Lin
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28439v1 Announce Type: cross Abstract: One model passed our fidelity check without ever opening the datasheet. We found it while qualifying models for an internal extraction service: a structured-output constraint had silently disabled tool use, and the model answered anyway, with fabrica...
177. ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT ​
Author: Huseyin Umut Isik, Mehmet Alp Ozaydin, Sila Kurugol, \c{S}eyda Ertekin
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28455v1 Announce Type: cross Abstract: Contrastive vision-language learning uses paired chest CT volumes and radiology reports to learn abnormality classifiers without manually annotated labels. However, two characteristics of chest CT challenge conventional global contrastive learning. F...
178. Anatomy-Aware Promptable Segmentation with Online Interactive Training for AUTOPET V ​
Author: Pablo Lozano-Jimenez, Sergio Romero-Tapiador, Ruben Tolosana
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28461v1 Announce Type: cross Abstract: We present an anatomy-aware, promptable model for whole-body lesion segmentation in FDG and PSMA PET/CT, developed for the AUTOPET V challenge. The proposed method is built as family of nnU-Net-based models and trained in two stages: i) a pre-trainin...
179. Real-time virtual circuits for plasma shape control via neural network emulators: experimental demonstration on MAST Upgrade ​
Author: Nicola C. Amorisco, Kamran Pentland, Adriano Agnello, George K. Holt, Alasdair Ross, Matthew J. Marshall, Edward Jones, Graham J. McArdle, Charles Vincent, Timothy Nunn, Martin Kochan, Pedro Cavestany, Aran Garrod, Stanislas Pamela, James Buchanan
Published: 8/31/2026, 4:00:00 AM
Categories: physics.plasm-ph, cs.AI
arXiv:2608.28468v1 Announce Type: cross Abstract: Conventional plasma shape control in tokamaks relies on virtual circuits (VCs) that are computed offline from linearisations around a small, tailored number of reference equilibria, and deployed as expertly prepared schedules during the discharge. He...
180. NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry ​
Author: Samuel Xiao, Judy Song, Rory Hu, Ziliang Zong
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.28481v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have demonstrated strong capabilities in natural language understanding and mathematical reasoning. However, their ability to translate informal mathematical problems into formal representations remains...
181. How Proper Scoring Rules Shape LLM Forecasting ​
Author: Benjamin Turtel, Paul Wilczewski, Kris Skotheim, Ville A. Satop"a"a, Philip E. Tetlock
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28482v1 Announce Type: cross Abstract: This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same ...
182. LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment ​
Author: Jingjing Nie, Jiawei Guo, Krishna Meda, Haipeng Cai
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.28490v1 Announce Type: cross Abstract: Software and systems security workflows are typically procedural: analysts inspect heterogeneous artifacts, form hypotheses, invoke tools, interpret outputs, and revise plans. Large language model (LLM)-based agents, which can plan, use tools, retain...
183. On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces ​
Author: Ahmed Hereiz, Yingzhe Lyu, Hao Li, Bram Adams, Ahmed E. Hassan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.28497v1 Announce Type: cross Abstract: AI coding agents, software tools that automate development tasks through reasoning and tool use, are increasingly extended through plugin marketplaces, yet the structure, maintenance, and co-evolution dynamics of these emerging repositories remain em...
184. Conformal Uncertainty Quantification Guarantees for Neural Operators ​
Author: Tom Stent, Nicolas Boull'e
Published: 8/31/2026, 4:00:00 AM
Categories: math.NA, cs.AI, cs.NA, math.PR
arXiv:2608.28515v1 Announce Type: cross Abstract: Neural operators provide fast surrogate models for approximating operators between function spaces, but their predictions often lack uncertainty quantification. We develop a split conformal framework to guarantee that a calibrated pointwise band arou...
185. Texture Image Classification Using DWT AlexNet Feature Fusion and Deep Neural Networks ​
Author: Arun D. Kulkarni
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28524v1 Announce Type: cross Abstract: Texture image classification plays a significant role in computer vision applications, including industrial inspection, medical image analysis, remote sensing, and object recognition. Handcrafted features can capture local texture characteristics but...
186. An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models ​
Author: Javier Aguilar Mart'in
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28541v1 Announce Type: cross Abstract: A code world model accepted by a sampling gate can be exactly right on everything the gate can see and arbitrarily wrong beyond it. We characterize what a certified model can know, and what its errors can cost, when the omission is an annular freeze ...
187. Video Generative Models as Geometry Learner ​
Author: Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.28549v1 Announce Type: cross Abstract: Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for ...
188. Blog: Survey of Optimizers ​
Author: Ruoran Xu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.28557v1 Announce Type: cross Abstract: Neural-network optimization in 2025-2026 is no longer well described as a succession of new Adam variants. The design space has expanded from coordinates to matrices and layers, from fixed training horizons to policies over time, and from mathematica...
189. Learning a Size-Weight Frontier for Synthetic-Augmented Inference ​
Author: Chengpiao Huang, Kaizheng Wang
Published: 8/31/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, cs.LG, stat.ML
arXiv:2608.28576v1 Announce Type: cross Abstract: Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference acro...
190. Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning ​
Author: Nan Wang, Mohit Yadav, Jonathan Wulff, Aidan Rosenbaum, Kezhou Chen, Yuvan Sharma, Xu Dong, Yiwei Tao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2608.28578v1 Announce Type: cross Abstract: Tendon-driven hands are anthropomorphic, and moving the actuators off the joints is what makes a hand of this capability affordable to build. Two effects produce that saving. Routing force through a cable removes the requirement that a motor fit insi...
191. Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning ​
Author: Ye Mo, Kai Ye, Xianwei Mao, Zirui Shao, Gang Huang, Bo Zhang, Hangdi Xing, Kehan Chen, Huan Zhou, Zixu Yan, Jiajun Bu, Sheng Zhou
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2505.18603v3 Announce Type: replace Abstract: Document understanding aims to perform question answering and information extraction over document images, where the visual content is highly information-dense and most queries rely on only a few relevant layout regions. However, existing methods e...
192. BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment Understanding ​
Author: Haofei Hou, Shunyi Zhao, Fanxu Meng, Kairui Yang, Lecheng Ruan, Qining Wang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2601.04524v3 Announce Type: replace Abstract: Understanding biomedical experiments provides a foundation for downstream tasks, e.g., laboratory automation, and facilitates effective cross-disciplinary communication. Two challenges, High Information Density (HID) and Multi-Step Reasoning (MSR),...
193. Multimodal Collaborative Debate for Zero-Shot Time Series Reasoning ​
Author: Patara Trirat, Jin Myung Kwak, Jay Heo, Heejun Lee, Sung Ju Hwang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2601.19151v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as natural-language interfaces to structured data, yet they remain brittle when reasoning over time series. Visual patterns can be misleading, numerical claims can be hallucinated, and textual cont...
194. Real-Time AI Service Economy: A Framework for Agentic Computing Across the Continuum ​
Author: Lauri Lov'en, Alaa Saleh, Reza Farahani, Ilir Murturi, Miguel Bordallo L'opez, Praveen Kumar Donta, Schahram Dustdar
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.05614v2 Announce Type: replace Abstract: Real-time AI services run across the device-edge-cloud continuum, where autonomous AI agents generate latency-sensitive workloads, orchestrate multi-stage pipelines, and compete for shared resources under governance constraints. This article shows ...
195. Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models ​
Author: Massimiliano Pappa, Luca Romani, Valentino Sacco, Alessio Palma, St'ephane Lathuili`ere, Fabio Galasso, Xavier Alameda-Pineda, Indro Spinelli
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.23149v2 Announce Type: replace Abstract: Deploying safety-critical agents requires anticipating the consequences of actions before they are executed. While world models offer a paradigm for this proactive foresight, current approaches relying on visual simulation incur prohibitive latenci...
196. PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization ​
Author: Zelin Tan, Zhouliang Yu, Bohan Lin, Zijie Geng, Hejia Geng, Yudong Zhang, Mulei Zhang, Yang Chen, Shuyue Hu, Zhenfei Yin, Chen Zhang, Lei Bai
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2603.26535v4 Announce Type: replace Abstract: We propose Process-Aware Policy Optimization (PAPO), a method that integrates process-level evaluation into Group Relative Policy Optimization (GRPO) through decoupled advantage normalization, to address two limitations of existing reward designs. ...
197. Prompts Without Evidence: How Neuroimaging Mentions Shift Clinical Vision-Language Model Predictions ​
Author: Doan Nam Long Vu, Simone Balloccu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2603.28387v3 Announce Type: replace Abstract: Trustworthy clinical AI must use real evidence and avoid relying on surface-level artifacts. We evaluate 12 open-weight vision-language models (VLMs) on two clinical neuroimaging cohorts for binary classification of affective disorders and cognitiv...
198. Understanding and Enforcing Weight Disentanglement in Task Arithmetic ​
Author: Shangge Liu, Yuehan Yin, Lei Wang, Qi Fan, Yinghuan Shi, Wenbin Li, Yang Gao, Dacheng Tao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.17078v2 Announce Type: replace Abstract: Task arithmetic provides an efficient, training-free way to edit pre-trained models, yet lacks a fundamental theoretical explanation for its success. The existing concept of ``weight disentanglement" describes the ideal outcome of non-interfering t...
199. D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery ​
Author: Hanane Nour Moussa, Yifei Li, Zhuoyang Li, Yankai Yang, Cheng Tang, Tianshu Zhang, Nesreen K. Ahmed, Ali Payani, Ziru Chen, Huan Sun
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2604.27977v3 Announce Type: replace Abstract: Despite recent progress in language models and agents for scientific data-driven discovery, advancing their capabilities is held back by the absence of verifiable environments representing real-world scientific tasks. To fill this gap, we introduce...
200. Rethinking Vacuity for OOD Detection in Evidential Deep Learning ​
Author: Claire McNamara
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.06382v2 Announce Type: replace Abstract: Vacuity, or Uncertainty Mass (UM), is commonly used as a metric to evaluate Out-of-Distribution (OOD) detection in Evidential Deep Learning (EDL). It generally involves dividing the number of classes ($K$) by the total strength of belief ($S$) of t...
201. Evidence-Based Intelligent Diagnostic and Therapeutic Visualization System with Large Language Models: Multi-Turn Interaction and Multimodal Treatment Plan Generation ​
Author: Yunhan Wang, Yuda Wang, Zhiying Tu, Mingqiang Song, Li Song, Kun Li, Dianhui Chu, Bolin Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.06869v2 Announce Type: replace Abstract: Aim: Existing AI-assisted traditional Chinese medicine diagnostic tools suffer from opaque reasoning processes, passive interaction, and limited treatment plan presentation. This study proposes a knowledge-enhanced visual diagnostic system to impro...
202. ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs ​
Author: Ashutosh Hathidara, Sai Shruthi Sistla, Sebastian Schreiber, Sahil Bansal
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.IR, cs.LG
arXiv:2606.12451v2 Announce Type: replace Abstract: Large language models deployed as agents over large tool catalogs face a critical tool-retrieval bottleneck. As embedding-based retrieval approaches rely on compact encoders that may under-capture specialized tool semantics, parametric tool retriev...
203. AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Properties ​
Author: Yifan Jiang, Meige Yang, Zitong Li, Jay Pujara
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.14240v2 Announce Type: replace Abstract: Affordance reasoning, the inference of an object's action possibilities from its physical properties (e.g., shape and material), is fundamental to human physical understanding and increasingly critical for Large Language Models (LLMs). However, exi...
204. RecourseBench: A Modular Framework for Reproducible Algorithmic Recourse Evaluation ​
Author: Hashir Ahmed, Zahra Khotanlou, Chenghao Tan, Ahmed Abdelaal, Amir-Hossein Karimi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2606.16113v2 Announce Type: replace Abstract: Algorithmic recourse methods provide counterfactual explanations that inform individuals of the actions required to overturn an unfavorable model decision. Despite rapid methodological progress, principled comparison remains elusive; existing frame...
205. Flow Reasoning Models: Turning Discrete Flows Into Efficient Recurrent Reasoners ​
Author: Alec Helbling, Andrey Bryutkin, Mauro Martino, Duen Horng Chau, Nima Dehmamy, Hendrik Strobelt
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.29150v2 Announce Type: replace Abstract: Structured reasoning requires making and revising interdependent decisions to reach a globally consistent solution. Existing architectures struggle with this: autoregressive models commit sequentially and cannot revise earlier decisions, while mask...
206. APeB: Benchmarking Personalization Ability of Large Language Model Agents ​
Author: Garry Yang, Zizhe Chen, Xinru Chen, Yongqiang Chen, Jianxiang Wang, Deyu Zou, Linyi Ding, Jialiang Wu, Yunzhong He, Yu Gong, James Cheng, Huaixiao Tou
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.HC
arXiv:2607.03162v2 Announce Type: replace Abstract: LLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent intent, extract preferences from noisy interaction histories, and select among competing alternatives. Existing...
207. Atomic Units of X: The Compression Layer of Intelligence ​
Author: Sachin Dev Duggal, Pradyumna Swarnalatha Ramanna, Alexandros Vassiliades
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.12634v3 Announce Type: replace Abstract: This paper proposes a theoretical and empirical framework for understanding intelligence as a process of atomic compression and compositional reuse. It argues that scalable cognitive, biological, computational, and organisational systems reduce com...
208. Set-shifting Behavioral Test for Harnessed Agents ​
Author: Ye Ziwei
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.SE
arXiv:2607.13396v2 Announce Type: replace Abstract: What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow the notion of set-shifting from cognitive psychology to study how well agents adapt to hidden reliability shifts. Our cognitive ...
209. SEGRA: A Structured Experience Guided Reasoning Agent for Property Graph Question Answering ​
Author: Saiyue Lyu, Mariam Dundua, Vishaal Kapoor, Sarthak Ahuja, Neda Kordjazi, Evren Yortucboylu, Harsh Amin, Rebecca Steinert
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.22713v2 Announce Type: replace Abstract: Enterprise IT support knowledge graphs capture rich relationships among cases, users, devices, symptoms, taxonomic categories, root causes, and historical resolutions. Yet querying them in Gremlin requires knowledge of graph schemas, traversal sema...
210. HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents ​
Author: Daeyoung Roh, Donghee Han
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.02009v3 Announce Type: replace Abstract: Retrieval-augmented search agents answer multi-hop questions by repeatedly issuing search queries and accumulating evidence. This creates a stopping problem: after the necessary evidence has appeared, further retrieval often adds cost, latency, and...
211. Agentao: A Policy-Governed Runtime Harness for Embeddable Tool-Using LLM Agents ​
Author: Bo Jin, Qiang Jiao, Xin Tong
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.13574v2 Announce Type: replace Abstract: LLM agents increasingly operate as execution systems that invoke tools, modify local state, use persistent memory, and interact with external protocols. These capabilities make agents useful, but they also introduce risks related to over-privileged...
212. When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation ​
Author: Avyay M. Casheekar, Hariganesh Tangirala
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.14940v3 Announce Type: replace Abstract: Agent evaluations commonly score the state observed when a run stops and count the run as one trial. Interpreting that score as a final result from a separate trial requires outcome finality and cross-unit separation. Outcome finality requires that...
213. RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training ​
Author: Yugu Li, Jimmy Cao, Jianglin Qiao, Siyi Hu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.18682v2 Announce Type: replace Abstract: Training multi-turn agentic workflows with reinforcement learning (RL) enables large language models to perform complex reasoning, use external tools, and conduct iterative search beyond single-turn settings. Yet multi-turn RL training remains high...
214. When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation ​
Author: Yearim Kim, Injun Baek, Nojun Kwak
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.19812v2 Announce Type: replace Abstract: To prevent the adoption of aesthetically polished but pedagogically flawed AI content, we study a video authoring pipeline featuring two layers of structured refusal. The first layer empowers educators to iteratively reshape AI scripts based on mul...
215. STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control ​
Author: Mengxi Luo, Changjia Chen, An Cao, Zirong Huang, Wanyi Dai
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.22538v2 Announce Type: replace Abstract: Policy-governed agents must interpret case evidence while reliably following authorized procedures. We present STAGE, an executable-graph framework that confines model judgment to policy-scoped nodes while placing procedural control in deterministi...
216. Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf ​
Author: Davood Wadi, Yu Ma
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, econ.GN, q-fin.EC
arXiv:2608.22697v2 Announce Type: replace Abstract: Search rankings are valuable because human attention is scarce and sequential. Higher-placed alternatives are easier to find, so they are examined and bought more often. Consumers are now delegating search to AI agents that can ingest an entire res...
217. Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors ​
Author: Joshua Penman
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CR, cs.LG
arXiv:2608.23873v2 Announce Type: replace Abstract: Everything a language model sees is tokens. The serving stack knows what each span is -- user input, tool output, instructions -- but the model must keep track of that itself, and can lose track or be confused: text can be written to read like anyt...
218. Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding ​
Author: Hyunho Kook, Junhyuk So, Tianyu Fu, Haizhong Zheng, Beidi Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.24024v2 Announce Type: replace Abstract: Confidence-based voting aggregates parallel LLM rollouts by weighting each with internal signals such as token log probabilities, and has been actively studied for single-turn reasoning. However, modern LLMs increasingly act as multi-turn search ag...
219. SKILL.state: Scalable Long-Horizon Agent Skills ​
Author: Sanket Badhe, Priyanka Tiwari, Jonghyun Chung
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.26263v2 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running procedural skills. Existing agent runtimes maintain execution by continually appending observations, actions, and intermediate reasoning traces to an...
220. AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design ​
Author: Mingquan Liu, Jiangyu Chen, Hanqun Cao, Xujun Zhang, Pengsen Ma, Xiangru Tang, Shuting Jin, Zhuo Yang, Annie Zheng, Tianfan Fu, Fang Wu, Xiangxiang Zeng
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.26747v2 Announce Type: replace Abstract: Scientific LLM agents have shown promise in literature reasoning, tool use, and experiment planning, but it remains unclear whether they can autonomously improve large, tightly coupled scientific machine-learning systems through executable code cha...
221. Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation ​
Author: Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong, Philippe Schwaller
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.27429v2 Announce Type: replace Abstract: Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model them either through de novo generation of product molecules or through heuristic graph edits that operate directly on molecular topol...
222. Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey ​
Author: Juan Zhong, Yuhang Shi, Zukang Xu, Xi Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, cs.RO, cs.SY, eess.SY
arXiv:2304.10891v4 Announce Type: replace-cross Abstract: Transformer-based models are becoming a central paradigm in autonomous driving because they can capture long-range spatial dependencies, multi-agent interactions, and multimodal context across perception, prediction, and planning. At the same...
223. Evaluating the Performance of Large Language Models on GAOKAO Benchmark ​
Author: Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying, Liang He, Xipeng Qiu, Tianxiang Sun, Peng Li, Shiqiao Meng, Yanjun Zheng, Jun Zhan, Zhangyue Yin, Xiannian Hu, Guofeng Quan, Qixiang Wang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2305.12474v4 Announce Type: replace-cross Abstract: Large Language Models(LLMs) have demonstrated remarkable performance across various natural language processing tasks; however, how to comprehensively and accurately assess their performance becomes an urgent issue to be addressed. This paper...
224. Let the Flows Tell: Solving Graph Combinatorial Optimization Problems with GFlowNets ​
Author: Dinghuai Zhang, Hanjun Dai, Esmeralda S. Whitammer, Aaron Courville, Yoshua Bengio, Ling Pan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DM, stat.ML
arXiv:2305.17010v4 Announce Type: replace-cross Abstract: Combinatorial optimization (CO) problems are often NP-hard and thus out of reach for exact algorithms, making them a tempting domain to apply machine learning methods. The highly structured constraints in these problems can hinder either opti...
225. Long Story Short: Story-level Video Understanding from 20K Short Films ​
Author: Ridouane Ghermi, Xi Wang, Vicky Kalogeiton, Ivan Laptev
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2406.10221v3 Announce Type: replace-cross Abstract: Recent developments in vision-language models have significantly advanced video understanding. Existing datasets and tasks, however, have notable limitations. Most datasets are confined to short videos with limited events and narrow narrative...
226. PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection ​
Author: Jinhe Bi, Aniri, Zengjie Jin, Yifan Wang, Danqi Yan, Wenke Huang, Xiaowen Ma, Sikuan Yan, Artur Hecker, Mang Ye, Xun Xiao, Hinrich Schuetze, Volker Tresp, Yunpu Ma
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2502.12119v5 Announce Type: replace-cross Abstract: Visual instruction tuning adapts pre-trained Multimodal Large Language Models (MLLMs) to follow human instructions for real-world applications. However, the rapid growth of these datasets introduces significant redundancy, leading to increase...
227. Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations ​
Author: Eunkyu Park, Wesley Hanwen Deng, Gunhee Kim, Motahhare Eslami, Maarten Sap
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2507.20409v3 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) prompting helps models think step by step. But naive CoT breaks down in visually grounded social tasks, where models must perceive, understand, and judge all at once; bridging perception with norm-grounded reasoning. Re...
228. Attention as Conditioning: What Classical Learning Theory Predicts About Linear Transformers ​
Author: Mu Qiao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.NC
arXiv:2508.08289v3 Announce Type: replace-cross Abstract: Attention is widely understood as an associative memory, but that description alone does not predict how the memory will behave. Predictive theories do exist, but in the literature on animal learning. We show that the state updates of the maj...
229. Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics ​
Author: Carter Blum, Katja Filippova, Ann Yuan, Asma Ghandeharioun, Julian Zimmert, Fred Zhang, Jessica Hoffmann, Tal Linzen, Martin Wattenberg, Lucas Dixon, Mor Geva
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2508.11017v4 Announce Type: replace-cross Abstract: Large language models (LLMs) struggle with cross-lingual knowledge transfer: they sometimes hallucinate when asked in one language about facts expressed in a different language during training. This work introduces a controlled setting to stu...
230. Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning ​
Author: Abdullah Abdelfattah, Mahmoud I. Khalil, Hazem Abbas
Published: 8/31/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL, cs.LG, cs.SD
arXiv:2509.00094v2 Announce Type: replace-cross Abstract: Assessing spoken language is challenging, and quantifying pronunciation metrics for machine learning models is even harder. However, for the Holy Quran, this task is enabled by the rigorous recitation rules (Tajweed) established through the e...
231. Steering Multimodal Large Language Models Decoding for Context-Aware Safety ​
Author: Zheyuan Liu, Zhangchen Xu, Guangyao Dou, Xiangchi Yuan, Zhaoxuan Tan, Radha Poovendran, Meng Jiang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2509.19212v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet their ability to make context-aware safety decisions remains limited. Existing methods often fail to balance oversensitivity (unjustified refus...
232. CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models ​
Author: Jie Cai
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2509.22737v3 Announce Type: replace-cross Abstract: Visual comparison reasoning is a fundamental capability of vision-language models (VLMs), covering judgments of object quantity, geometric dimensions, spatial relations, and temporal order. Yet existing benchmarks rarely isolate comparison as...
233. Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection ​
Author: Sojung An, Kwanyong Park, Yong Jae Lee, Donghyun Kim
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2509.24192v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have advanced multimodal perception, demonstrated by open-vocabulary object detection with simple language queries. State-of-the-art VLMs still struggle to handle complex queries involving descriptive attributes ...
234. OceanGym: A Benchmark Environment for Underwater Embodied Agents ​
Author: Yida Xue, Mingjun Mao, Xiangyuan Ru, Yuqi Zhu, Baochang Ren, Shuofei Qiao, Mengru Wang, Shumin Deng, Xinyu An, Ningyu Zhang, Ying Chen, Huajun Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.LG, cs.RO
arXiv:2509.26536v3 Announce Type: replace-cross Abstract: We introduce OceanGym, the first comprehensive benchmark for ocean underwater embodied agents, designed to advance AI in one of the most demanding real-world environments. Unlike terrestrial or aerial domains, underwater settings present extr...
235. PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering ​
Author: Md Mahadi Hasan Nahid, Davood Rafiei
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR
arXiv:2510.14278v2 Announce Type: replace-cross Abstract: Retrieval plays a central role in multi-hop question answering (QA), where answering complex questions requires gathering multiple pieces of evidence. We propose PRISM, an agentic retrieval framework that leverages large language models (LLMs...
236. Riverbank Erosion Analysis in Bangladesh Using Spatiotemporal Segmentation ​
Author: M. Saifuzzaman Rafat, Akif Islam, Mohd Ruhul Ameen, Momen Khandoker Ope, Abu Saleh Musa Miah, Jungpil Shin
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2510.17198v2 Announce Type: replace-cross Abstract: Riverbank erosion is a serious environmental problem in Bangladesh, causing land loss, damage to infrastructure, and displacement of local communities. Manual analysis of satellite images is often slow and difficult to apply consistently acro...
237. Quantifying Affective Bias in Low-Resource Media: Large-Scale Emotion Profiling of Bengali Headlines ​
Author: Mohd Ruhul Ameen, Akif Islam, Ayesha Siddiqua, Abu Saleh Musa Miah, Jungpil Shin
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2510.17252v2 Announce Type: replace-cross Abstract: News media can influence readers not only through the events they report but also through the emotional tone used to present them. This issue is especially important in digital news environments, where headlines often shape first impressions ...
238. Think-at-Hard: Dynamic Looped Transformers for Improved Reasoning ​
Author: Tianyu Fu, Yichen You, Zekai Chen, Guohao Dai, Huazhong Yang, Yu Wang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.PF
arXiv:2511.08577v4 Announce Type: replace-cross Abstract: Improving the reasoning abilities of Large Language Models (LLMs), especially under parameter constraints, is crucial for real-world applications. Looped transformers address this by performing multiple latent iterations to refine each token ...
239. OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion ​
Author: Sai Koneru, Matthias Huck, Jan Niehues
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2512.00234v3 Announce Type: replace-cross Abstract: There has been significant progress in open-source text-only translation large language models (LLMs) with better language coverage and quality. However, these models can be only used in cascaded pipelines for speech translation (ST), perform...
240. The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior ​
Author: Erik Larsen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2512.12066v3 Announce Type: replace-cross Abstract: Current safety evaluations of large language models rely on single-shot testing, implicitly assuming that model responses are deterministic and representative of the model's safety alignment. We challenge this assumption by investigating the ...
241. FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation ​
Author: Junseok Lee, Chang-Jae Chun
Published: 8/31/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.SD
arXiv:2601.06199v4 Announce Type: replace-cross Abstract: Scaling Multimodal Large Language Models (MLLMs) to long-form speech is bottlenecked by the explosive growth of input tokens. Existing speech-language models project high-frame-rate acoustic features directly into the LLM input space, making ...
242. Aligning Agentic World Models via Knowledgeable Experience Learning ​
Author: Baochang Ren, Yunzhi Yao, Rui Sun, Shuofei Qiao, Ningyu Zhang, Huajun Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.LG, cs.MM
arXiv:2601.13247v2 Announce Type: replace-cross Abstract: Current Large Language Models (LLMs) exhibit a critical modal disconnect: they possess vast semantic knowledge but lack the procedural grounding to respect the immutable laws of the physical world. Consequently, while these agents implicitly ...
243. CoFrGeNet: Continued Fraction Architectures for Language Generation ​
Author: Amit Dhurandhar, Vijil Chenthamarakshan, Dennis Wei, Tejaswini Pedapati, Karthikeyan Natesan Ramamurthy, Rahul Nair
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2601.21766v5 Announce Type: replace-cross Abstract: Transformers are arguably the preferred architecture for language generation. In this paper, inspired by continued fractions, we introduce a new function class for generative modeling. The architecture family implementing this function class ...
244. Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning ​
Author: Yu Xu, Yuxin Zhang, Lin Gao, Oliver Deussen, Tong-Yee Lee, Fan Tang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2602.01335v2 Announce Type: replace-cross Abstract: A visual metaphor constitutes a high-order form of human creativity, employing cross-domain semantic fusion to transform abstract concepts into impactful visual rhetoric. Despite the remarkable progress of generative AI, existing models remai...
245. SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models ​
Author: Hyeonbeom Choi, Daechul Ahn, Youhan Lee, Taewook Kang, Seongwon Cho, Jonghyun Choi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2602.04208v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic control, with test-time scaling (TTS) gaining attention to enhance robustness beyond training. However, existing TTS methods for VLAs require...
246. ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents ​
Author: Youjin Wang, Run Zhou, Yingjie Ma, Rong Fu, Jiani Liang, Shuaishuai Cao, Min Huang, Tao Fang, Liangming Pan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2602.04935v4 Announce Type: replace-cross Abstract: Adapting LLM agents to domain-specific tool calling remains notably brittle under evolving interfaces. Prompt and schema engineering is easy to deploy but often fragile under distribution shift and strict parsers, while continual parameter-ef...
247. FENCE: A Financial and Multimodal Jailbreak Detection Dataset ​
Author: Mirae Kim, Seonghun Jeong, Youngjun Kwak
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB
arXiv:2602.18154v3 Announce Type: replace-cross Abstract: Jailbreaking poses a significant risk to the deployment of Large Language Models (LLMs) and Vision Language Models (VLMs). VLMs are particularly vulnerable because they process both text and images, creating broader attack surfaces. However, ...
248. From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves ​
Author: Haritz Puerto, Haonan Li, Xudong Han, Timothy Baldwin, Iryna Gurevych
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2602.24210v4 Announce Type: replace-cross Abstract: Large reasoning models (LRMs) produce reasoning traces (RTs) that often contain sensitive information. These leaky thoughts are difficult to control and frequently violate explicit privacy directives. Because RTs can be exposed through prompt...
249. Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts ​
Author: Lucas Bandarkar, Alan Ansell, Trevor Cohn
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.17070v2 Announce Type: replace-cross Abstract: In this work, we analyze shortcomings in cross-lingual knowledge transfer in large, modern reasoning LLMs. We demonstrate that the perceived gap in knowledge transfer is primarily a script barrier. First, we conduct an observational data anal...
250. InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model ​
Author: Youjin Wang, Jiaqiao Zhao, Rong Fu, Run Zhou, Ruizhe Zhang, Jiani Liang, Suisuai Cao, Feng Zhou
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2603.18031v2 Announce Type: replace-cross Abstract: Balancing fine-grained local modeling with long-range dependency capture under computational constraints remains a central challenge in sequence modeling. While Transformers provide strong token mixing, they suffer from quadratic complexity, ...
251. The Autonomy Tax: Defense Training Breaks LLM Agents ​
Author: Shawn Li, Yue Zhao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2603.19423v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously complete complex multi-step tasks. Practitioners deploy defense-trained models to protect against prompt...
252. Var-JEPA: A Variational Formulation of the Joint-Embedding Predictive Architecture - Bridging Predictive and Generative Self-Supervised Learning ​
Author: Moritz G"ogl, Christopher Yau
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2603.20111v2 Announce Type: replace-cross Abstract: The Joint-Embedding Predictive Architecture (JEPA) is often seen as a non-generative alternative to likelihood-based self-supervised learning, emphasizing prediction in representation space rather than reconstruction in observation space. We ...
253. Select, Label, Evaluate: Active Testing in NLP ​
Author: Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu, Fabrizio Silvestri, Amin Mantrach
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.21840v2 Announce Type: replace-cross Abstract: Human annotation cost and time remain significant bottlenecks in Natural Language Processing (NLP), with test data annotation being particularly expensive due to the stringent requirement for low-error and high-quality labels necessary for re...
254. Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence ​
Author: Peter Fasogbon, Ugurcan Budak, Patrice Rondao Alface, Hamed Rezazadegan Tavakoli
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2603.21933v3 Announce Type: replace-cross Abstract: The pruning of 3D Gaussian splats is essential for reducing their complexity to enable efficient storage, transmission, and downstream processing. However, most of the existing pruning strategies depend on camera parameters, rendered images, ...
255. Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning ​
Author: Juekai Lin, Yun Zhu, Honglin Lin, Sijing Li, Tianwei Lin, Zheng Liu, Xiaoyang Wang, Wenqiao Zhang, Lijun Wu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2604.06079v2 Announce Type: replace-cross Abstract: Graphics Program Synthesis is pivotal for interpreting and editing visual data, effectively facilitating the reverse-engineering of static visuals into editable TikZ code. While TikZ is the de facto standard for scientific schematics due to i...
256. PolicyLong: Towards On-Policy Context Extension ​
Author: Junlong Jia, Jiang Zhou, Ziyang Chen, Xing Wu, Chaochen Gao, TingHao Yu, Feng Zhang, Songlin Hu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.07809v2 Announce Type: replace-cross Abstract: Extending LLM context windows is hindered by scarce high-quality long-context data. Recent methods synthesize data with genuine long-range dependencies via information-theoretic verification, selecting contexts that reduce a base model's pred...
257. Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks ​
Author: Yuangang Li, Justin Tian Jin Chen, Ethan Yu, David Hong, Iftekhar Ahmed
Published: 8/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2604.12379v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly rely on explicit reasoning to solve coding tasks, yet evaluating the quality of this reasoning remains challenging. Existing reasoning evaluators are not designed for coding, and current benchmarks fo...
258. Benefits of Low-Cost Bio-Inspiration in the Age of Overparametrization ​
Author: Kevin Godin-Dubois, Anil Yaman, Anna V. Kononova
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2604.20365v2 Announce Type: replace-cross Abstract: While Central Pattern Generators (CPGs) and Multi-Layer Perceptrons (MLP) are widely used paradigms in robot control, few systematic studies have been performed on the relative merits of large parameter spaces in highly constrained settings. ...
259. Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs ​
Author: Joseba Fernandez de Landa, Carla Perez-Almendros, Jose Camacho-Collados
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2604.21751v2 Announce Type: replace-cross Abstract: LLMs have limitations when it comes to cultural coverage and competence, and in some cases, show specific cultural biases. Although prior studies have examined the cultural capabilities of LLMs, none have specifically investigated their regio...
260. G-Loss: Graph-Guided Fine-Tuning of Language Models ​
Author: Aditya Sharma, Vinti Agarwal, Rajesh Kumar
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2604.25853v4 Announce Type: replace-cross Abstract: Traditional loss functions, including cross-entropy, contrastive, triplet, and su pervised contrastive losses, used for fine-tuning pre-trained language models such as BERT, operate only within local neighborhoods and fail to account for the ...
261. ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space ​
Author: Gabe Guo, Thanawat Sornwanee, Lutong Hao, Elon Litman, Stefano Ermon, Jose Blanchet
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.27443v3 Announce Type: replace-cross Abstract: Generating continuous-time, continuous-space stochastic processes (e.g., videos, weather forecasts) conditioned on partial observations (e.g., first and last frames) is a fundamental challenge. Existing approaches, (e.g., diffusion models), s...
262. SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces ​
Author: Chang Jin, An Wang, Zeming Wei, Kai Wang, Biaojie Zeng, Qiaosheng Zhang, Chao Yang, Jingjing Qu, Xia Hu, Xingcheng Xu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG, cs.MA
arXiv:2605.12015v3 Announce Type: replace-cross Abstract: Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environments. However, this modularity introduces attack surfaces that...
263. Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control ​
Author: Rohith Uppala
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2605.18414v3 Announce Type: replace-cross Abstract: Large language models increasingly operate as autonomous agents that select and invoke tools from large registries. We identify a critical gap: when unauthorized tools are visible in an agent's context, models select them in 48-68% of adversa...
264. SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals ​
Author: Zihang Lin, Huaiyuan Qin, Muli Yang, Hongyuan Zhu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2605.21919v2 Announce Type: replace-cross Abstract: Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development indicators, where incomplete evidence use and imperfect evidence integration can introduc...
265. More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations ​
Author: Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, Shu Zhong
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2605.26647v2 Announce Type: replace-cross Abstract: Feedforward network (FFN) layers account for a large fraction of parameters and nonlinear expressivity in Transformer-based large language models (LLMs). Despite the evolution from ReLU and GELU to gated variants such as SwiGLU, most FFN desi...
266. Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models ​
Author: Mingze Wang, Shuchen Zhu, Yuxin Fang, Binghui Li, Kai Shen, Shu Zhong
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2605.26895v2 Announce Type: replace-cross Abstract: Normalization layers in modern large language models (LLMs) consist of a deterministic normalization operation and a learnable scale vector. While the normalization operation has been extensively studied, the scale vector remains poorly under...
267. LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis ​
Author: Kewei Xu, Xiaoben Lu, Shuofei Qiao, Zihan Ding, Haoming Xu, Lei Liang, Ningyu Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.MA
arXiv:2605.30434v2 Announce Type: replace-cross Abstract: Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long horizons untested. We introduce LongDS, a be...
268. DiffuSent: Towards a Unified Diffusion Framework for Aspect-Based Sentiment Analysis ​
Author: Shu Long, Yanglei Gan, Xuchuan Zhou
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.01323v2 Announce Type: replace-cross Abstract: Aspect-Based Sentiment Analysis (ABSA) encompasses seven distinct subtasks, each focusing on different extracted elements. Despite the proven success of generative models in unified aspect sentiment analysis, existing approaches often rely on...
269. The Granularity Gap: A Multi-Dimensional Cross-Generational Audit of Sycophancy in Gemini Models ​
Author: Patrick Keough
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2606.05183v3 Announce Type: replace-cross Abstract: Pass/fail safety evaluation reports whether a model refused. It does not report how far a model went to please the user, and we show these are close to different measurements. We audited sycophancy across three Gemini generations, scoring N=8...
270. TokenPilot: Cache-Efficient Context Management for LLM Agents ​
Author: Buqiang Xu, Zirui Xue, Dianmou Chen, Chenyang Fu, Chiyu Wu, Caiying Huang, Chen Jiang, Jizhan Fang, Xinle Deng, Yijun Chen, Yunzhi Yao, Xuehai Wang, Jin Shang, Gong Yu, Ningyu Zhang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.MA
arXiv:2606.17016v2 Announce Type: replace-cross Abstract: As LLM agents are deployed in long-horizon sessions, context accumulation drives up inference costs. Existing approaches utilize text pruning or dynamic memory eviction to minimize token footprints; however, their unconstrained sequence mutat...
271. The Discrete-Log Clock: How a Transformer Learns Modular Multiplication ​
Author: Huu Danh Nguyen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.17399v2 Announce Type: replace-cross Abstract: When small transformers grok modular multiplication, prior work reports that the learned embedding has a "dense" Fourier spectrum requiring all frequencies. This contrasts with modular addition, where only a sparse set of key frequencies suff...
272. CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories ​
Author: Anneliese Brei, Abhisheik Sharma, Nicholas Sanaie, Lu Wang, Snigdha Chaturvedi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2606.22454v2 Announce Type: replace-cross Abstract: As LLM-generated text is increasingly used, especially in fictional domains, we explore how much LLM-generated stories differ from human-written stories. In this work, we focus on characters. We borrow definitions from narratology to analyze ...
273. An LLM-Based Framework for Intent-Driven Network Topology Design ​
Author: Kholoud El-Habbouli, Fen Zhou, Stephane Huet
Published: 8/31/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.CL
arXiv:2607.00292v2 Announce Type: replace-cross Abstract: Designing deployable and resilient network topologies from natural language requirements remains a challenging problem in network automation. This work investigates the ability of Large Language Models (LLMs) to generate structurally valid an...
274. GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning ​
Author: Kaicong Huang, Weiheng Oh, Jack M. Reilly, Thomas Guggisberg, Ruimin Ke
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2607.13569v3 Announce Type: replace-cross Abstract: Transit video understanding can provide valuable fine-grained data that conventional passenger counters and fare systems cannot capture. However, supervised video models require task-specific annotations, while applying vision-language models...
275. On the Depth Scalability of Logic Gate Networks ​
Author: Taegun An, Dohun kim, Haebeom Lee, Changhee Joo
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO
arXiv:2607.21633v3 Announce Type: replace-cross Abstract: Logic Gate Networks (LGNs) compute through compositions of Boolean operations, yet existing LGNs do not reliably benefit from increased depth. We identify two causes: optimization collapse and topology-induced degradation of output-specific c...
276. REPREC: Representation Driven Parameter-Efficient Recommendation System ​
Author: Harshini Kavuru, Dwipam Katariya, Giri Iyengar, Pranab Mohanty, Kalanand Mishra, Raghu Machiraju
Published: 8/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2607.24845v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have been applied to sequential recommendation by incorporating collaborative signals through input conditioning or model adaptation. However, existing approaches often require LLM fine-tuning, additional architec...
277. Where Steering Signals Come From: Activation Source Selection in Activation Steering ​
Author: Jiaran Ye, Lingxu Ran, Zijun Yao, Chenpeng Wang, Yong Jiang, Lei Hou, Juanzi Li, Liangming Pan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.25270v2 Announce Type: replace-cross Abstract: Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail. We study this source choice as activation...
278. Locked Evaluation Surfaces: Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation-Effect Prediction ​
Author: Mehrdad Shoeibi, Niloofar Yousefi
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.00152v2 Announce Type: replace-cross Abstract: Predicting how held-out target genes respond to CRISPRi perturbation, and whether such predictions transfer across biological screens, is hard to evaluate: a representation can be informative within one screen yet fail across screens, while e...
279. Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Search Agents ​
Author: Shuai Wang, Haodong Chen, Yu Yin, Shengyao Zhuang, Bevan Koopman, Guido Zuccon
Published: 8/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2608.02751v3 Announce Type: replace-cross Abstract: Existing deep-search agents use a Search-Visit workflow that retrieves whole webpages without considering the structure they expose through titles, headings, sections, and metadata. This prevents agents from directly constraining retrieval to...
280. ED-CSP: Crystal Structure Prediction from Electron Diffraction ​
Author: Germain Poloudenny, Arnaud Demorti`ere, Ya"el Fr'egier
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06448v3 Announce Type: replace-cross Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. Existing ED-based learning methods mainly predict crystallographic labels, reconstruct struc...
281. BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference ​
Author: Jinlong Yang, Jinke Wu, Lizilin, Yao Zhou
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.07572v2 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their massive computational overhead, temporal feature caching has been proposed to bypass redundant computation...
282. RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation ​
Author: Boxiong Wang, Hui Kang, Geng Sun, Jiahui Li, Chao Yu, Daxin Tian
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.09467v2 Announce Type: replace-cross Abstract: Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-actio...
283. PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models ​
Author: Siddharth Patel
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.14741v2 Announce Type: replace-cross Abstract: We introduce PolyComp, a procedurally generated and verified benchmark that stresses visual recognition and compositional spatial reasoning. In each problem, a model must identify which of four options shows a pair of polycube components that...
284. How Far Should Tokenization Go? Predictive Effectiveness and Relational Losslessness ​
Author: Yi Wang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SD
arXiv:2608.18025v2 Announce Type: replace-cross Abstract: GPT-style models have achieved remarkable success with finite vocabularies of reusable tokens, making the token interface a central component of modern sequence modeling. Symbolic music appears naturally compatible with this paradigm: it cons...
285. JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification ​
Author: Tianxin Zhou, Ruixi Lin
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.20607v2 Announce Type: replace-cross Abstract: Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several reference-free judges agree can create a hidden risk: agreement may reflect shared false-negative blind...
286. Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation ​
Author: Rujin Liang, Zhongpu Chen, Yuhao Lei, Xin Miao
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.20756v2 Announce Type: replace-cross Abstract: While multimodal retrieval-augmented generation (RAG) systems increasingly rely on images as external knowledge sources, the introduction of poisoned visual evidence can severely compromise multimodal large language model (MLLM) generation. U...
287. Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic Constraints ​
Author: Gwen Yidou-Weng, Edward Sun, Tianyi Ma, Metin Alp Dogan, Benjie Wang, Allen Peng, Guy Van den Broeck, Yuchen Cui
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.22149v3 Announce Type: replace-cross Abstract: LLMs generate fluent plans for robots but routinely violate the syntactic and se8mantic constraints they must satisfy to execute, and existing remedies trade formal guarantees against plan quality: soft methods (affordance scoring, grounded d...
288. GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets ​
Author: Saif Ahmed, Asadullah Hil Galib, S. M. Riaz Rahman Antu, Ahmed Faizul Haque Dhrubo, Souvik Pramanik, Mohammad Abdul Qayum, Mohsin Sajjad, Mohammad Ashrafuzzaman Khan
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.22272v2 Announce Type: replace-cross Abstract: Generative adversarial networks (GANs) can provide efficient image generation, while diffusion models offer high-quality image restoration but require iterative sampling. This paper presents a hybrid GAN-guided diffusion framework that uses a...
289. Multi-Winner Voting with Argumentative Ballots ​
Author: Ryuta Arisaka, Hirotaka Ono
Published: 8/31/2026, 4:00:00 AM
Categories: cs.GT, cs.AI
arXiv:2608.23247v2 Announce Type: replace-cross Abstract: We introduce multi-winner voting with argumentative ballots (MVArg) and investigate theoretical properties. As our conceptual contribution, we generalise approval ballots to argumentative ballots, thereby allowing voters to express defeasible...
290. Macro-Operator Generation and Predicate Selection for TAMP Operator Learning ​
Author: Can Emir Bora, Emre Ugur
Published: 8/31/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.23629v2 Announce Type: replace-cross Abstract: Creating symbolic operators by hand is one of the main bottlenecks in deploying Task and Motion Planning systems (TAMP). Recent works show that these operators can instead be learned directly from demonstration data. Existing methods, however...
291. On-policy Distillation with Verifiable Reward ​
Author: Wenze Lin, Jiale Zhao, Xitai Jiang, Songde Rao, Yining Li, Shenzhi Wang, Bingxiang He, Gao Huang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.24696v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides...
292. SpecMine: A Large-Scale Corpus of Spec-Driven Development Artifacts ​
Author: Shyam Agarwal, Bogdan Vasilescu
Published: 8/31/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.25202v2 Announce Type: replace-cross Abstract: Spec-Driven Development (SDD) is a fast-emerging practice in which a structured natural-language specification, written by a developer, or (more often) drafted by an AI tool and then curated by the developer, drives an AI coding agent's imple...
293. MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize ​
Author: Jiaxin Yuan, Connor Martinez Lockhart, Xiaoyu Liu, Jiaqi Wang, Chenghao Deng, Xiayimei Han, Vlassis Mastrantonis, Dmitrii Gudin, Shaopeng Zhu, Abdirisak Mohamed, Bilal Aytekin, Jiewen Lang, Zezheng Song, Furong Huang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LO
arXiv:2608.25449v2 Announce Type: replace-cross Abstract: Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of robustness...
294. When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory ​
Author: Kazuki Nakayashiki
Published: 8/31/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2608.25553v3 Announce Type: replace-cross Abstract: Provenance links keep the evidence behind an inherited belief reachable; an agent with a verification budget must still choose which links to inspect. We study a consolidated memory that states a decision constraint and whose source record ha...
295. TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development ​
Author: Jiarui Yan, Weiwei Sun, Sijie Li, Wenhan Li, Yiming Yang
Published: 8/31/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.26086v2 Announce Type: replace-cross Abstract: Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning development, where an agent must revise data pipelines, models, and validation over hours of feedback, and on most competition...
296. AI Models Can Predict and Collaboratively Modulate Human Memory Search ​
Author: Eric Lacosse, Mariana Duarte, Graham Todd, Peter M. Todd, Daniel C. McNamee
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC
arXiv:2608.26152v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in many language-based tasks, for example routine coding, these artificial intelligence models have reduc...
297. Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation ​
Author: Jesse St. Amand, Callum Canavan, Sohaib Imran, Joseph Hewson, Aaron Lutz, Shi Feng, Puria Radmard, Lennie Wells
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26159v2 Announce Type: replace-cross Abstract: Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors: an LLM may recognize outputs from other copies of the same model and make bia...
298. Comparing Chunking and Embedding Strategies for Turkish RAG Systems ​
Author: Mustafa Serta\c{c} T"urkel, Fatma Nur Korkmaz, Ahmet Tu\u{g}rul Bayrak
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.26192v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation conditions a language model on chunks retrieved from a document collection. Its accuracy is therefore limited by the chunking and embedding stages that determine what can be retrieved. We compare Turkish documen...
299. Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI ​
Author: Architect Labs
Published: 8/31/2026, 4:00:00 AM
Categories: cs.AR, cs.AI
arXiv:2608.26418v2 Announce Type: replace-cross Abstract: Modern AI workloads and the hardware that runs them evolve on different timescales: architectural definition precedes volume silicon by years, while target workloads shift in months. Design decisions are therefore committed under deep uncerta...
300. LiveVVT: High-Fidelity Video Virtual Try-On in Real Time ​
Author: Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.26714v2 Announce Type: replace-cross Abstract: Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overhead in practical continuous deployment. Na...
301. Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents ​
Author: Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang, Hao Zheng, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Jifeng Zhu, Bin Chong
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.27141v2 Announce Type: replace-cross Abstract: Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iteratio...
302. PAWBench: How Far Are We from Probabilistically Aligned World Modeling? ​
Author: Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram {\DJ}or{\dj}evi'c, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jinbo Xing, Xi Chen
Published: 8/31/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.27345v2 Announce Type: replace-cross Abstract: Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of pos...