Skip to content

arXiv cs.AI - 2026-09-02 ​

445 items collected.


1. HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models ​

Author: Yun-Jian Zhang, Chen-Wei Liang, Tian-Yi Zhang, Jian Ding, Yi-Lun Wu, Ao-Bo Li, Wei-Cong Su, Saifullah, Hong-Yu An, Mu-Jiang-Shan Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00002v1 Announce Type: new Abstract: World models enable language-model agents to predict environment dynamics and plan before acting. In text environments, the model must learn symbolic action effects from serialized state descriptions, but the role of serialization structure remains und...

📖 Read original article


Author: Leonardo Santiago Benitez Pereira, Marcos Escudero Vi~nolo, Luis Herranz Arribas
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2609.00003v1 Announce Type: new Abstract: Machine unlearning studies the removal of knowledge from an AI model, making the system forget a concept it previously learned. Despite rapid progress in generative machine unlearning, the unintended degradation of semantically related concepts that sh...

📖 Read original article


3. Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing ​

Author: L'ea Bayati, Mohamed Dahmoune, Melek Rodoplu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.NE, math.OC

arXiv:2609.00004v1 Announce Type: new Abstract: This paper studies a finite-horizon multi-item capacitated lot-sizing problem in which demand quantities are deterministic, while demand-arrival periods are stochastic. Each demand occurs once within a known time window and must be satisfied no later t...

📖 Read original article


4. Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models ​

Author: Parviz Ghafariasl, Weimin Fu, Xiaolong Guo, Shing I. Chang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00005v1 Announce Type: new Abstract: Financial scams targeting older adults increasingly occur through text and voice channels such as email, SMS, and phone calls, unfolding over multiple conversational turns that begin with impersonation or casual contact, escalate through trust building...

📖 Read original article


5. Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls ​

Author: Dheeraj Mohandas Pai, Lu Xian
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.MA

arXiv:2609.00012v1 Announce Type: new Abstract: Long-horizon tasks remain uncommon in large language model (LLM) evaluation, and for a reason: when each step depends on the last, per-step accuracy that looks excellent in isolation decays catastrophically, as errors cascade and the end-to-end failure...

📖 Read original article


6. OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets ​

Author: Dongsheng Chen, Xiangyu Zhao, Xin Yao, Xuetao Wei
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2609.00015v2 Announce Type: new Abstract: AI agents powered by large language models are evolving from isolated assistants into heterogeneous systems in which multiple agents, planners, tools, and execution backends operate over shared environments. In such settings, safety becomes a system-le...

📖 Read original article


7. SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces ​

Author: Ranjit Raut, Aarav Subedi, Sagun Rai, Sudan Jha
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2609.00018v1 Announce Type: new Abstract: Computer science papers rely heavily on diagrams: architecture drawings, system flowcharts, and pipeline schematics that often carry more information than the text around them. There is currently no public dataset that pairs this specific kind of figur...

📖 Read original article


8. UI-Venus-2 Technical Report ​

Author: Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jianrong Jiang, Jianguo Li, Runze Li, Jinzhen Lin, Zhenyu Ma, Changhua Meng, Han Peng, Xinyu Qiu, Shuheng Shen, Zhongyi Shui, Weiqiang Wang, Ming Wen, Zhuoer Xu, Hang Yan, Kaiwen Yang, Ruilin Yao, Nanjun Yu, Zhengwen Zeng, Lianrui Zhang, Yunzhu Zhang, Zhe Zhao, Beitong Zhou
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.LG

arXiv:2609.00028v1 Announce Type: new Abstract: Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task constru...

📖 Read original article


Author: Ren Zhenzhuo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2609.00032v1 Announce Type: new Abstract: Mathematical communities work with different objects, invariants, and tools, so transferring a problem across them is expensive and often skipped. We present EULER, a multi-agent system that takes such a transfer--a bridge--as its unit of search. Aroun...

📖 Read original article


10. When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation ​

Author: Cong Cao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ME

arXiv:2609.00071v1 Announce Type: new Abstract: Prediction error is widely used to evaluate nuisance-function estimators in causal inference, but its relationship with causal estimator performance may differ across performance measures. We studied this question in a partially linear model using Mont...

📖 Read original article


11. MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts ​

Author: V. S. Anoop, Devika N
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.00073v1 Announce Type: new Abstract: Malaria remains a significant global health burden, necessitating continuous research efforts to understand its complex molecular mechanisms, epidemiology, and potential therapeutic interventions. Extracting essential biomedical information from the va...

📖 Read original article


12. AI Morbidity and Mortality: A Framework for Clinical AI Failure Review ​

Author: Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2609.00076v1 Announce Type: new Abstract: Clinical artificial intelligence is increasingly embedded in real-world care, yet existing safety mechanisms are poorly suited to reconstructing and learning from individual AI-related errors and near-misses. Aggregate model monitoring can identify per...

📖 Read original article


13. Different representation learning objectives recover distinct latent structures from the same psychometric data ​

Author: Cong Cao, Tassos C. Kyriakides, Pambos Vrasidas
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ME

arXiv:2609.00100v1 Announce Type: new Abstract: Psychometric questionnaires contain rich item-level information, yet it remains unclear whether different representation learning objectives recover the same latent organization. We investigated this question using 757 matched teacher-child pairs from ...

📖 Read original article


14. Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations ​

Author: Ao Qu, Panagiotis Michelakis, Linyuan Han, Yiannis Hadjiyianni, Kun Ouyang, Konstantinos Siskos, Feng Li, Ran Meng, Jingchi Jiang, Dimitrios Stamoulis, Jie Liu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2609.00106v1 Announce Type: new Abstract: This paper presents FAIRY, a full-stack smart-agriculture agent system developed for and deployed to an operating soybean research farm at Harbin Institute of Technology's smart-agriculture site. We develop FAIRY to execute and evaluate agentic agronom...

📖 Read original article


15. Recursive Criticality of AI Self-Improvement ​

Author: Mikhail Burtsev
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00137v1 Announce Type: new Abstract: AI is increasingly used in the R&D process that produces future AI systems. We study the conditions under which this feedback becomes self-amplifying. Our model describes how the rate of AI capability growth depends on baseline research productivity, ...

📖 Read original article


16. IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training ​

Author: Rongze Tang, Jianjie Fang, Zhaolu Wang, Ziyou Wang, Xvyuan Liu, Haisheng Su, Xin Zhang, Wei Wu, Chen Gao, Yong Li, Zhibo Chen
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2609.00161v1 Announce Type: new Abstract: World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining the generation process wi...

📖 Read original article


17. Asymmetries in Spontaneous and Instructed Deception ​

Author: Josiah Luikham
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in ...

📖 Read original article


18. LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark ​

Author: Irem Yoldas, Martim Brand~ao, Jie Zhang, Odinaldo Rodrigues
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.CY

arXiv:2609.00192v1 Announce Type: new Abstract: Public trust in Autonomous Vehicles (AVs) may depend not only on technical success but also on the fairness of their decision making. While a recent trend in AV research involves using general purpose "common sense" models to guide AV decision making, ...

📖 Read original article


19. ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation ​

Author: Muzhao Tian, Zezi Zeng, Yifan Yang, Xin Gao, Yan Li, Zisu Huang, Xiaohua Wang, Changze Lv, Mingxi Cheng, Bei Liu, Kai Qiu, Qi Dai, Dong Chen, Yue Dong, Xiaoqing Zheng, Ji Li, Chong Luo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00194v1 Announce Type: new Abstract: Document-to-slide generation is challenging because slides are dense editable artifacts that require both faithful content selection and precise spatial layout. Recent slide agents adopt iterative reflection, but typically follow a monolithic "one vers...

📖 Read original article


20. AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning ​

Author: Scott Compton, Arjun Nagendran
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00211v1 Announce Type: new Abstract: Conversational artificial intelligence is increasingly embedded in everyday social environments, where it functions as both an informational tool and a source of interpersonal feedback. This perspective introduces contingency, i.e., the degree to which...

📖 Read original article


21. ConvDeck: Conversational Paper-to-Slide Generation via Stage-Specific User Feedback ​

Author: Tarik Can Ozden, Sachidanand VS, Furkan Horoz, Ozgur Kara, Dilek Hakkani-T"ur, Junho Kim, James Matthew Rehg
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00226v1 Announce Type: new Abstract: Automatic academic paper-to-slide generation is inherently iterative, because creating an effective presentation requires repeated cycles of generation, critique, and revision. Recent multi-agent systems partially acknowledge this through internal crit...

📖 Read original article


22. Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems ​

Author: Rakibul Hasan Rajib, Mengxing Zheng, Qian Lou
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.00237v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems tackle complex reasoning by orchestrating how multiple agents are configured and how they collaborate. A central challenge is to adapt orchestration to the evolving collaboration state. Routing from ...

📖 Read original article


23. Invalidation Contracts for Cross-Episode Agent Memory ​

Author: Michael Wu, Arquimedes Canedo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00243v1 Announce Type: new Abstract: LLM agents that cache recovery suggestions from API errors can skip re-derivation in later episodes, spending fewer tokens and fewer model calls on constraints they have already learned. Server-side data drift turns those cached fixes into silent failu...

📖 Read original article


24. Authority Bias in Conversational Search Engines for Academic Paper Recommendation ​

Author: Uthman Jinadu, Parsa Ghazvinian, Anjila Budathoki, Benjamin M. Ampel, Rajshekhar Sunderraman, Yi Ding
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00248v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used as conversational search engines for academic literature, yet whether they judge papers on content or on authority signals has not been tested causally. We investigate authority bias: systematic prefer...

📖 Read original article


25. Hypotheses-Guided Self Distillation for Continual Personalization ​

Author: EunJeong Hwang, Kushan Mitra, Dan Zhang, Hannah Kim, Estevam Hruschka
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00251v1 Announce Type: new Abstract: As people increasingly interact with LLM assistants in daily life, continually adapting to individual preferences has become essential for effective long-term interactions. However, user preferences are rarely stated in full, and instead emerge through...

📖 Read original article


26. The Answer Is Not the Argument ​

Author: Will Yeadon, Sergio Ju'arez, Paul Mackay, T. J. Dowling, Elise Agra, Oto-obong Inyang, Arin Mizouri, Craig P. Testrow
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00264v1 Announce Type: new Abstract: Chain-of-thought monitoring is proposed for AI oversight, yet evaluations often provide monitors with a trusted reference answer. We ask whether answer access improves reasoning verification or mainly exposes incorrect conclusions. We collected 237 ste...

📖 Read original article


27. Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching ​

Author: Kartik Ravisankar, Hojat Abdolanezhad, Daniel Capo, Sang Su Lee, Shishir Dash, Vijay Anand Raghavan
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00274v1 Announce Type: new Abstract: Two-sided service marketplaces are moving from deterministic request-form intake to AI-native probabilistic matching, enabled by large language models (LLMs) that infer intent, preferences, and latent constraints from natural language. Relying on infer...

📖 Read original article


28. The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems ​

Author: Bardia Mohammadi, Laurent Bindschaedler
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.OS

arXiv:2609.00275v1 Announce Type: new Abstract: Fleets of LLM agents now externalize effects that cannot be fully undone: they move money, deploy code, delete data, and disclose information. Current controls check one effect at a time, so a fleet of individually authorized agents can overdraw its pr...

📖 Read original article


29. The Assistant's Ideal Self ​

Author: Mert Yazan
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00304v1 Announce Type: new Abstract: Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant's preferred stated ideal self. Thirty-two qualities...

📖 Read original article


30. Human-AI Co-Interpretation for Responsible AI: A Hermeneutic Perspective ​

Author: Behrooz Razeghi
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00334v1 Announce Type: new Abstract: Across law, education, policy analysis, and public moral argumentation, LLM outputs are being used often for work that requires interpretations to be justified with textual evidence and explicit normative standards. Yet a recurrent failure mode -- what...

📖 Read original article


31. SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning ​

Author: Beidi Zhao, Gexin Huang, Ciro Zhang, Anqi Li, Yusheng Tan, Chen Zhou, Gang Wang, Zu-hua Gao, Xiaoxiao Li
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00342v1 Announce Type: new Abstract: Whole-slide images (WSIs) are challenging for vision-language reasoning because diagnostically relevant morphology is sparse, heterogeneous, and distributed across gigapixel-scale images and multiple spatial resolutions. Existing WSI models and patholo...

📖 Read original article


32. Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models ​

Author: Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV

arXiv:2609.00355v1 Announce Type: new Abstract: Speculative decoding accelerates generation without changing its output, yet on vision-language models (VLMs) it has been caught in a self-defeating cycle. The drafter stays autoregressive, so it must stay small. A small drafter cannot afford the image...

📖 Read original article


33. A Stable Aggregation Method for Quantum Federated Learning ​

Author: Shanika Nanayakkara, Shiva Raj Pokhrel
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2609.00356v1 Announce Type: new Abstract: Quantum federated learning (QFL) enables clients to train quantum neural network (QNN) models without sharing private data. We find that aggregation in QFL is unstable under heterogeneous data, unreliable communication, variable fidelity, latency, and ...

📖 Read original article


34. Dr. Claw: An AI Scientist Workspace for Vibe Research ​

Author: Dingjie Song, Hanrong Zhang, Dawei Liu, Yixin Liu, Zongxia Li, Zhengqing Yuan, Siqi Zhang, Henry Peng Zou, Zhiling Yan, Yuxuan Zhang, Yanfang Ye, Philip S. Yu, Lichao Sun
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.LG

arXiv:2609.00365v1 Announce Type: new Abstract: Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, yet end-to-end research still fragments across chat tools, IDEs, terminals, and writing environments, and the decisions that make it ...

📖 Read original article


35. RestoreBench: Can AI Agents Restore Power Flow Convergence? ​

Author: Riccardo Mansutti, Andrea Pomarico, Robert Jakob, Qian Zhang, Alberto Berizzi, Kevin O'Sullivan
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY

arXiv:2609.00384v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly automate multi-step engineering workflows through tool use, interpretation of intermediate results, and iterative planning. Diagnosing and resolving non-convergent power flow cases is a promising yet large...

📖 Read original article


36. Dependency-Aware Chain-of-Thought Compression for Financial Reasoning ​

Author: Wenjun Wu, Lei Fu, Kejian Tong, Tao Ning, Sichen Zhao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00413v1 Announce Type: new Abstract: Chain of thought prompting improves complex reasoning, but its long intermediate traces create substantial inference cost and hinder practical deployment in financial settings. We present a Hierarchical Semantic Distillation Network, HSDN, for compress...

📖 Read original article


37. SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation ​

Author: Songwei Dong, Bingyan Lu, Makayla Kienlen, J. Nicholas Laneman, Cong Shen
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00427v1 Announce Type: new Abstract: The exponential growth of wireless devices is driving unprecedented spectrum demand, pushing spectrum management toward more fine-grained decisions across space, time, and device constraints. As a result, spectrum policymakers and engineers must proces...

📖 Read original article


38. SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents ​

Author: Rayan Khoury, Shih-Yao Lin, Pratyush Mishra
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2609.00434v1 Announce Type: new Abstract: Evaluating task-oriented dialogue agents requires judging not merely whether a reply reads well but whether each turn advances the underlying workflow state correctly--a distinction conventional holistic LLM judges can miss because they evaluate the av...

📖 Read original article


39. Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations ​

Author: Fanyou Wu, Suraj Maharjan, Ainur Yessenalina, Dennis Xu Chen, Rahul Srivastava, Srinivasan H. Sengamedu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00441v1 Announce Type: new Abstract: Effective manager-employee communication is critical for retaining high performers and developing underperformers, yet training managers in these skills remains costly. Text-based chatbots offer a scalable approach but cannot provide realistic rehearsa...

📖 Read original article


40. mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers ​

Author: Timothy Kassis
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00453v1 Announce Type: new Abstract: Giving an agent a file about a named expert can supply hard-to-find material, produce a recognizable persona, or change what the agent decides. These are different claims. We test each one. mimeo is an open-source tool that finds a person's public work...

📖 Read original article


41. Towards a Belief-Based World Model for LLM Agents ​

Author: Shubham Kumar, Harshit Kumar, Narendra Ahuja, Saurabh Jha
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00455v1 Announce Type: new Abstract: Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning capabilities, LLMs struggle with long-horizon tasks, especially under partial observability. World model...

📖 Read original article


42. EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models ​

Author: Muran Yu, Jiechao Gao, Yuandong Pan, Barney H. Miao, Andrew C. Lesh, Kincho H. Law, Jie Wang, Michael D. Lepech
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00479v1 Announce Type: new Abstract: For emerging scientific research domains, local Small Language Models (SLMs) are becoming more attractive, as they offer stronger privacy control and more stable deployment pipelines than Large Language Models. However, in practice, scientific question...

📖 Read original article


43. The Privacy-Hallucination Tradeoff in Differentially Private Language Models ​

Author: Krithika Ramesh, Krishna Pillutla, Danish Pruthi, Anjalie Field
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.00492v1 Announce Type: new Abstract: Both privacy and factual accuracy are paramount in high-stakes domains like healthcare. Concerningly, we uncover and investigate a privacy-hallucination tradeoff in differentially private (DP) language models. First, we empirically show that models pre...

📖 Read original article


44. Validity-Aware Jailbreak Evaluation for Large Language Models ​

Author: Qilong Wu, Sahil Wadhwa, Pranab Mohanty, Giri Iyengar, Varun Chandrasekaran
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00498v1 Announce Type: new Abstract: Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic plausibility rather...

📖 Read original article


45. Wave Function Backpropagation with Explicit Temporal-Interval Dynamics ​

Author: Byunggu Yu, Justin Kim
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00503v1 Announce Type: new Abstract: Conventional neural networks learn predominantly through affine transformations followed by nonlinear activations, while elapsed time is often treated as an auxiliary feature or assumed to be uniformly sampled. This paper introduces Wave Function Backp...

📖 Read original article


46. CoVer: Conflict-Aware Claim Verification ​

Author: Shuning Zhang, Dai Shi, Bohao Chu, Hui Wang, Yuwei Chuai, Yifan Wang, Jingruo Chen, Simin Li, Xin Yi, Hewu Li
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00508v1 Announce Type: new Abstract: Social media fact-checking has long been challenged by evidence-level and aggregation-level conflicts, where erroneous evidence mimics authoritative news sources. To capture this challenge and support conflict verification tasks, we present ContraNote,...

📖 Read original article


47. When the Algorithm Becomes the Brand Crisis: A Sociotechnical Theory of Distributed Responsibility and Accountable Transparency ​

Author: Mohammad Saleh Torkestani, Taha Mansouri
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00510v1 Announce Type: new Abstract: Artificial intelligence systems increasingly enact market-facing promises through chatbots, recommendation systems, automated decisions, and generative interfaces. Their failures, misuse, and misrepresentation raise a question that conventional brand-c...

📖 Read original article


48. ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation ​

Author: Siyuan Zhang, Hanchen Wang, Dong Wen, Ying Zhang, Wenjie Zhang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00513v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) mitigates large language models (LLMs) hallucinations, yet conventional dense retrieval struggles with the complex reasoning paths of multi-hop question answering (QA). Graph-based RAG captures multi-step relationsh...

📖 Read original article


49. Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation ​

Author: Zhuoheng Li, Ying Chen
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00543v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems rely on external corpora that may contain outdated, contradictory, noisy, or unreliable documents, introducing reliability risks. Prior work has leveraged document relations to improve the answer reliability...

📖 Read original article


50. VoiceLongMemEval: Do Assistants Remember How You Sounded? ​

Author: Ramit Pahwa, Parivesh Priye, Apoorva Beedu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2609.00570v2 Announce Type: new Abstract: With the growing scale of multi-agent architectures and large language models, deployed AI assistants are increasingly tasked with reasoning over long, continuous, multi-session conversation histories. Current benchmarks evaluate this dialogue history ...

📖 Read original article


51. Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs ​

Author: Seungwoo Jung, Dohyeok Kwon, Seungmin Cha, Junseok Lee, Yeonho Yoo, Chuck Yoo, Gyeongsik Yang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00575v2 Announce Type: new Abstract: Mixture-of-experts (MoE) architectures scale large language models efficiently, but they demand massive GPU memory. To cope with such demand, models are commonly compressed to reduce their memory footprint. Residual sparsification is a representative c...

📖 Read original article


52. Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random ​

Author: Cris Huynh
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.00576v1 Announce Type: new Abstract: Item-sensitivity, defined as whether a model's choice depends on the specific input rather than on its own output prior, is widely reported as evidence of task competence. We show this evidence is necessary but not sufficient using a forced-choice sign...

📖 Read original article


53. Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts ​

Author: Rui Yang, Yang Hong, Yichao Xu, Zhengyu Liu, Ziyang Li, Yinzhi Cao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2609.00578v1 Announce Type: new Abstract: Large Language Models (LLMs) can solve complex problems, but their misuse in high-risk domains can lead to severe consequences. Model providers therefore restrict assistance for potentially harmful requests. Refusing all cybersecurity requests would th...

📖 Read original article


54. Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing ​

Author: Alexandre Clin Deffarges, Nataliya Kosmyna, Pattie Maes
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2609.00584v1 Announce Type: new Abstract: Does unrestricted AI access bypass the cognitive effort required for learning, or does it streamline knowledge acquisition? This paper reports on a study where we compare three designs for user-AI interaction in a learning context: (1) an unrestricted ...

📖 Read original article


55. Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs ​

Author: Wentao Zhang, Syed Shariyar Murtaza, Junaid Ahmad Bhatti, Utkarsh Soni, Yifan Nie, Eugene Wen, Yuntian Deng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA

arXiv:2609.00621v1 Announce Type: new Abstract: Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-critical protocols, such as message routing, output formatting, and term...

📖 Read original article


56. REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows ​

Author: Ruoling Qi, Xuaner Wu, Penghang Liu, Jian Chen, Yirui Liu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00643v1 Announce Type: new Abstract: Agent revisions expose a fundamental correctness--efficiency trade-off during concurrent execution. Discarding ongoing work preserves latest-version correctness but wastes progress that may remain valid, whereas reusing prior work preserves efficiency ...

📖 Read original article


57. DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation ​

Author: Haoyuan Shi (Hunyuan, Tencent), Mingtao Chen (Hunyuan, Tencent), Shuo Jiang (Hunyuan, Tencent), Ziyan Chen (Hunyuan, Tencent, Beijing Film Academy), Xuyi Sheng (Peking University), Yiming Liu (Hunyuan, Tencent), Ying Zhang (Hunyuan, Tencent), Miao Wang (Hunyuan, Tencent, Shenzhen University), Jianxiang Lu (Hunyuan, Tencent), Fanyang Lu (Hunyuan, Tencent), Songyuanyi Lu (Hunyuan, Tencent), Xiele Wu (Hunyuan, Tencent), Zhichao Hu (Hunyuan, Tencent), Yuhong Liu (Hunyuan, Tencent), Richeng Xuan (Hunyuan, Tencent)
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00646v1 Announce Type: new Abstract: Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead o...

📖 Read original article


Author: Enrong Pan, Ryan Zhou, Ting Hu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.NE

arXiv:2609.00652v1 Announce Type: new Abstract: Language model agents increasingly propose actions, observe external feedback, and explain their own behavior. Their confidence and rationales are convenient monitoring signals, but convenience is not verification. We introduce an environment-grounded ...

📖 Read original article


59. SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task ​

Author: Qiming Bao, Ne\c{s}et "Ozkan Tan, Siyuan Wang, Mark Gahegan
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.00654v1 Announce Type: new Abstract: We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task~\cite{sciclaimeval}, which asks systems to verify scientific claims against the tables and figures of a paper. Rather than tuning a single model, we benchma...

📖 Read original article


60. Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets ​

Author: Cheung Hao Lee, Patrick Wong
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00662v1 Announce Type: new Abstract: A multi-model language service must route each request while preserving workload-level budgets for compute, latency, memory, or monetary cost. Two features make this problem materially harder than static model selection. Prompt representations are high...

📖 Read original article


61. Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs ​

Author: Jainil Dharmil Shah
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00665v1 Announce Type: new Abstract: Edge-AI model selection is commonly driven by one isolated metric - accuracy, latency, memory, energy, or safety, even though a deployable language model must balance all five. Our work focuses on answering the question whether na- tively trained small...

📖 Read original article


62. Value Over Language Model: Detecting Original Contribution in Writing ​

Author: Vibhhu Sharma, Thorsten Joachims, Sarah Dean
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.00700v1 Announce Type: new Abstract: LLMs have been rapidly adopted across writing tasks, prompting the development of tools for detecting LLM-generated text. Yet, these tools largely measure how much of a document's surface text was written by an LLM and aren't fundamentally designed to ...

📖 Read original article


63. ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything ​

Author: Yufan Dang, Shu Yao, Bowen Lai, Chenting Xu, Ruijie Shi, Wai-Shing Leung, Huatao Li, Chen Qian, Zhiyuan Liu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA

arXiv:2609.00714v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems (MAS) have shown strong potential for solving complex tasks, yet their development forces a tradeoff: code frameworks are expressive but engineering-intensive, while no-code builders simplify authori...

📖 Read original article


64. A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies ​

Author: Ahmad Alfan Alfian Irfan, Nur Ahmad Khatim, Mansur Arief
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2609.00718v1 Announce Type: new Abstract: Many automobile and mobility companies deploy learned driving policies on embedded computers with limited memory and power. Pruning, knowledge distillation, and quantization are the standard methods to reduce the size and the inference cost of these po...

📖 Read original article


65. SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification ​

Author: Swapnil Bhattacharyya, Mayank Baranwal
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, math.OC

arXiv:2609.00728v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown remarkable promise in translating and reformulating complex mathematical optimization problems across modeling languages. However, validating such transformations through empirical solver executions alone is unre...

📖 Read original article


66. Agentic Empirical Asset Pricing: Methodological Foundations ​

Author: Yingjian Pan, Xiaowei Ding, Kay Giesecke
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, q-fin.ST

arXiv:2609.00731v1 Announce Type: new Abstract: Recent advances in LLM agents enable a new paradigm for asset pricing, which we call Agentic Empirical Asset Pricing (AEAP): systems that autonomously conduct the scientific discovery process itself. We define AEAP and identify its core building blocks...

📖 Read original article


67. Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs ​

Author: Lu Cheng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00738v1 Announce Type: new Abstract: Inference-time search with large language models (LLMs) often concentrates on a small set of structurally or semantically similar trajectories, leaving alternatives underexplored---a failure mode we call \textit{reasoning basin collapse}. We introduce ...

📖 Read original article


68. ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents ​

Author: Peng Xu, Zuyu Zhang, Yuze Sun, Feng Tian, Long Wang, Chen Zhang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.DB

arXiv:2609.00749v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents require context assembly: the runtime must decide what to include in each prompt, in what order, and when to compact history under a hard context-window budget and a byte-sensitive prompt cache. In product...

📖 Read original article


69. S^3martCirc: Self-supervised Smart Circuit Discovery ​

Author: Wendy Zheng, Yinhan He, Liang Wu, Jundong Li
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00755v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks, from text summarization to question answering. Despite these capabilities, their black-box nature obscures internal decision-making processes. Mechanistic inter...

📖 Read original article


70. Automated Tree Knowledge Graph Construction using Ontology Expansion and Retrieval from Vietnamese History Textbooks ​

Author: Ket Doan Nguyen, Minh N. H. Nguyen
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00763v1 Announce Type: new Abstract: Hierarchical Knowledge graph (KG)-based retrieval augmented generation (RAG) has emerged as a powerful approach for supporting large language models with structured knowledge. However, there are primary challenges: (i) the lack of methods for automatic...

📖 Read original article


71. DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory ​

Author: Xincheng Wei, Yifan Ding, Yoshua Li, Dongsheng Ma, Rongxiang Weng, Xunliang Cai, Wenjian Ding, Yao Zhang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00768v1 Announce Type: new Abstract: Self-play is an effective paradigm for language-model self-evolution, but without guidance, solver performance can plateau or decline across rounds. Unguided methods steer question generation with signals such as difficulty, learnability, or diversity....

📖 Read original article


72. When Features Become Instances: Inverted Contrastive Learning for Unsupervised Feature Selection ​

Author: Utsab Ghosh, Roshni Chakraborty
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.00782v1 Announce Type: new Abstract: Unsupervised feature selection seeks a compact subset of informative features without access to class labels, making feature utility difficult to define. Existing UFS methods therefore rely on indirect structural criteria, such as similarity preservati...

📖 Read original article


73. StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability? ​

Author: Yinghao Chen, Zixi Chen, Bingxiang He, Ziqing Qiao, Huan-ang Gao, Yinuo Xu, Yuxin Zuo, Zeyuan Liu, Yuhao Zhan, Chaojun Xiao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00787v1 Announce Type: new Abstract: Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. We argue that an ideal self-evolution method should share the same property, that is autonomously learning from raw training material...

📖 Read original article


74. Towards a Reliable and Practical Eval Pipeline ​

Author: Emma Thuong Nguyen, Abhishek Ghose
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2609.00805v1 Announce Type: new Abstract: LLM-based software systems increasingly require effective "evals" as quality gates in the development lifecycle. However, existing work typically addresses individual aspects of eval reliability rather than the full set of practical requirements. We pr...

📖 Read original article


75. One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning ​

Author: Xiaowei Sun, Jin Li, Yili Hong, Yikun Fu, Yanghua Xiao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00813v1 Announce Type: new Abstract: While reinforcement learning has enabled LLM-based search agents to invoke external tools, existing methods train under fixed budgets and cannot adapt when constraints vary at deployment. We propose AnySearch, a framework that enables a single policy t...

📖 Read original article


76. AnalysisBank: An Expert Analysis Pattern Library for Financial Report Generation ​

Author: Yajing Yang, Yunshan Ma, Kelvin J. L. Koa, Min-Yen Kan
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00818v1 Announce Type: new Abstract: We argue that financial report generation should operate at the analytical rather than structural level, composing content from data-derived insights rather than high-level topics or sections. To this end, we propose AnalysisBank, which distills expert...

📖 Read original article


77. Polished but Unresolved: Identifying Late-Stage Pressure States in Long-Horizon Tool-Use Agents ​

Author: Haoyang Chen, Yi Liu, Jianzhi Shao, Xiaozhou Xu, Zhe Sun, Wei Hu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.00823v1 Announce Type: new Abstract: Long-horizon tool-use agents need not only to search and plan, but also to decide when to finalize. We study late-stage pressure states, in which an agent is biased toward submitting a final answer that appears complete and polished while key constrain...

📖 Read original article


78. FLaG: Frequency-Domain Latent-attention Gated Pooling for Token Aggregation ​

Author: Kewei Li, Rongying Zhang, Xueli Wang, Xiwen Gong, Zhongjian Wang, Qiuchen Zhao, Lan Huang, Ruochi Zhang, Fengfeng Zhou
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, q-bio.BM

arXiv:2609.00831v1 Announce Type: new Abstract: Token aggregation converts token-level representations into fixed-dimensional sample representations, but most pooling methods operate only in the original token space. We introduce Frequency-Domain Latent-attention Gated Pooling (FLaG), a plug-in aggr...

📖 Read original article


79. Towards Generalizable Visually Grounded Exploration of Household Devices ​

Author: Linhao Zheng, Zeming Liu, Wangke Chen, Li Zeng, Wanxiang Che, Heyan Huang, Yuhang Guo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00845v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have demonstrated impressive capabilities in static visual recognition and high-level semantic reasoning. However, current embodied exploration paradigms still heavily rely on imitation learning from...

📖 Read original article


80. Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources ​

Author: Ivan Decostanzi, Michele Ronco, Sergio Consoli, Christina Corbane, Lorenzo Bertolini, Indaco Biazzo, Daria Mihaila, Manuel Garcia-Herranz, Felix Schwebel, Yelena Mejova, Kyriaki Kalimeri
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.00858v1 Announce Type: new Abstract: Effective humanitarian response depends on the rapid synthesis of heterogeneous, high-volume information sources - a task that routinely exceeds human analytical capacity in the critical early hours of a crisis. We present a pipeline that combines stru...

📖 Read original article


81. Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems ​

Author: Yi Chen, Zikang Yu, Jiahai Wang, Jinbiao Chen, Jianpeng Zhou, Zizhen Zhang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00859v1 Announce Type: new Abstract: Vehicle Routing Problems (VRPs) are fundamental combinatorial optimization problems with widespread applications in various scenarios. The advanced optimization solvers can effectively solve such problems. However, modeling complex VRP variants for sol...

📖 Read original article


82. Beyond the Clock: Measuring the Value of Adaptive Revision ​

Author: Ayushi Chadha
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00874v1 Announce Type: new Abstract: As agentic systems become compound systems, increasingly important decisions move above task execution itself: when should a higher-level controller preserve the strategy guiding another process, and when should it revise it? We study this meta-level c...

📖 Read original article


83. FractalNet-Based Heterogeneous Federated Learning for Orbital Edge Intelligence in Satellite Mega-Constellations: A Wildfire Case Study ​

Author: Sai Puppala, Koushik Sinha
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.ET, cs.LG

arXiv:2609.00875v1 Announce Type: new Abstract: Satellite mega-constellations are emerging as large-scale sensing, communication, and computation fabrics, yet their learning architectures remain largely inherited from terrestrial federated learning and ground-centric mission operations--- ill-suited...

📖 Read original article


84. Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning ​

Author: Yuanjun Zhang, Fuzel Ahamed Shaik, Suvojit Acharjee, Fahad Khalid, Mourad Oussalah
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00879v1 Announce Type: new Abstract: Reliable disaster damage assessment requires models that provide both accurate predictions and transparent explanations. However, existing multimodal approaches are limited by scarce annotated data and insufficient evaluation of reasoning quality. This...

📖 Read original article


85. Denoising Diffusion Generative Models Secretly Calculate Attentions ​

Author: Farzan Haddadi, Leila Monfared, Ebrahim Rezaii, Mohammadreza Malek-Mohammadi, Pejman Zakalvand, Narges Mokhtari
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG, cs.NE

arXiv:2609.00885v1 Announce Type: new Abstract: Denoising diffusion models are the dominant architecture for image generation, whereas most natural language generation and modeling are primarily handled by well-known transformer architectures employing attention mechanism. Here, we show that diffusi...

📖 Read original article


86. CacheBridge: Efficient Cross-Model KV Cache Transfer ​

Author: Xingyu Qu, Siyuan Lu, Zhiyu Chen, Sheng Wang, Tao Lin
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00891v1 Announce Type: new Abstract: Sharing context between LLMs in a multi-model system requires the receiving model to prefill the shared prefix because KV caches are model-specific. Recent closed-form cross-model KV transfer, hereafter Full-Head Mapping, avoids this replay by fitting ...

📖 Read original article


87. CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training ​

Author: Siyuan Li, Xinxin Song, Chen Ruinian, Jingjing Fan, Tingxiong Xiao, Yangen Hu, Ke Zeng, Jinli Suo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00892v1 Announce Type: new Abstract: Rubric-based reinforcement learning decomposes open-ended instructions into prompt-specific, flexible rubrics, making it better suited than reinforcement learning with verifiable rewards for post-training LLMs on open-ended tasks. However, static rubri...

📖 Read original article


88. In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access? ​

Author: Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi, Yusuke Haruki, Daisuke Kawahara
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00904v1 Announce Type: new Abstract: Whether large language models (LLMs) can control their own internal representations matters for both machine metacognition and AI safety. A recent study applied neurofeedback to LLMs and claimed that they can control their internal representations. How...

📖 Read original article


89. RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation ​

Author: Zhongru Chen, Yuan Wu, Yi Chang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.00918v1 Announce Type: new Abstract: Large language models are increasingly used as interactive recommender assistants. Their evaluation should therefore go beyond plausible item recommendation and test whether they can recognize flawed recommendation requests. Existing recommender benchm...

📖 Read original article


90. VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences ​

Author: Yiwen Jiang, Yang Deng, Stephanie Fong, Zimu Wang, Yaling Shen, Wei Feng, Hongxi Yang, Xiangyu Zhao, Zhongxing Xu, Deval Mehta, Xuelian Cheng, Zongyuan Ge
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.00921v1 Announce Type: new Abstract: Personalized Large Language Models (PLLMs) aim to tailor responses to individual users, where a central challenge is preference reasoning: inferring query-relevant preferences from user-related history. Existing benchmarks, however, largely assume that...

📖 Read original article


91. Few-Shot Out of Domain Intent Detection with Covariance Corrected Mahalanobis Distance ​

Author: Jayasimha Talur, Oleg Smirnov, Paul Missault
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00961v1 Announce Type: new Abstract: Conversational agents like chatbots and voice assistants are trained to understand and respond to user intents. On encountering an utterance with an intent different from the ones they have been trained on, these agents are expected to classify the int...

📖 Read original article


92. CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins ​

Author: Wenhao Zou, Xianglong Liu, Wendong Bi, Hanjie Wang, Simin Zhao, Gong Zhi
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.00967v1 Announce Type: new Abstract: As large language models increasingly act through external tools, deciding when to call a tool has become a central problem alongside deciding how to use it. Unnecessary tool calls introduce latency, cost, retrieval noise, and error propagation, while ...

📖 Read original article


93. Figures as Programs: Recursive Generation of Editable Scientific Figures ​

Author: Yepeng Liu, Dasen Dai, Chengzhi Liu, Yiren Song, Hai Ci, Yu Zhang, Qi Zhang, Mike Zheng Shou, Xin Eric Wang, Yuheng Bu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.GR

arXiv:2609.01006v1 Announce Type: new Abstract: Scientific methodology figures are essential for communicating complex methods clearly, yet creating them remains labor-intensive and typically requires multiple rounds of refinement. Recent image-generation models can synthesize visually appealing ras...

📖 Read original article


94. Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees ​

Author: Molly Wang (Imperial Business School)
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, math.PR

arXiv:2609.01035v1 Announce Type: new Abstract: Recursive LLM agents can broaden their search by spawning specialists. Some branches later request tools that send data or deploy code. When should a branch receive authority to act? We distinguish sandbox spawning, in which external controls prevent t...

📖 Read original article


95. Data-Driven Persona-Conditioned Agents for A/B Test Simulation ​

Author: Ziyad Benomar, Weronika {\L}ajewska, Leonardo Perelli, Saab Mansour
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01038v1 Announce Type: new Abstract: A/B testing is the gold standard for evaluating product changes, but each experiment requires real user traffic, engineering effort, and weeks of measurement. We propose a simulation framework that predicts A/B test outcomes using LLM-powered agents co...

📖 Read original article


96. AgentFactory: Towards Automated Agentic System Design and Optimization ​

Author: Enci Zhang, Haofeng Wang, Yuesheng Zhu, Xiaole Cui, Guibo Luo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01045v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities as powerful components in agentic systems, enabling sophisticated reasoning and complex task execution. However, current approaches to manually designing and optimizing agentic syst...

📖 Read original article


97. QILP-0: Constructing Observational Declarative Twins of Quantum Circuits ​

Author: Marina de la Cruz Echeand'ia, C'esar Luis Alonso, Tony Ribeiro, Alfonso Ortega de la Puente
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, quant-ph

arXiv:2609.01049v1 Announce Type: new Abstract: This paper introduces QXymb, a general framework for constructing observational declarative twins of quantum circuits, and develops QILP-0, its first complete order-0 specialization. QILP-0 constructs a finite multi-valued propositional logic program f...

📖 Read original article


98. WorldBench: Culturally Grounded Benchmark for Multilingual Agents ​

Author: Leonardo Ranaldi, Sherrie Shen, Jushi Kai, Alexandra Birch
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.01056v1 Announce Type: new Abstract: Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing benchmarks rarely test state preservation, performance across languages, and application to realistic, grounded scenarios. To address these concer...

📖 Read original article


99. User Representation via Cross Multi-source Behavior Pre-training for Mobile Games ​

Author: Chengqi Yang, Yiran Qiao, Feng Liu, Xingyu Lou, Zijun Zhou, Xiaoyun Mo, Changwang Zhang, Jiayuan Xu, Jun Wang, Xiang Ao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01057v1 Announce Type: new Abstract: User representation pre-training has become a fundamental paradigm for alleviating data sparsity in downstream personalization tasks. However, existing studies predominantly focus on single-app or app-level behaviors, overlooking the inherently cross-s...

📖 Read original article


100. ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning ​

Author: Fanrui Zhang, Ruixue Ding, Qiang Zhang, Xi Chen, Boli Chen, Shihang Wang, Qiuchen Wang, Hongmin Zhan, Jinxin Bian, Li xingchao, Peijin Zheng, Hao cheng, Pengjun Xie, Kaipeng Zhang, Jiawei Liu, Zheng-Jun Zha
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01058v1 Announce Type: new Abstract: Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brittle and unsta...

📖 Read original article


101. Space Generative AI with Solar Energy Harvesting ​

Author: Jierui Zhang, Jianhao Huang, Zhanwei Wang, Kaibin Huang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.NI, eess.SP

arXiv:2609.01062v1 Announce Type: new Abstract: Satellites are emerging as promising platforms to extend generative \emph{artificial intelligence} (AI) services to remote areas lacking terrestrial infrastructure. However, deploying space generative AI is fundamentally constrained by the limited, tim...

📖 Read original article


102. Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs ​

Author: Zhaoliang Chen, Jie Fu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.01117v1 Announce Type: new Abstract: Chain-of-thought reasoning unfolds in discrete token space: each step is committed as text, errors propagate, and eliciting good traces presupposes traces to imitate. Reasoning instead in a model's continuous representation space - where intermediate s...

📖 Read original article


103. Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate ​

Author: Kaiyan Wen, Shijie Zhang, Lu Yu, Guangdong Bai
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.MM

arXiv:2609.01168v2 Announce Type: new Abstract: Text-to-image (T2I) models remain vulnerable to jailbreak attacks that elicit Not-Safe-For-Work (NSFW) content, despite increasingly being guarded by heterogeneous, multi-layer safety stacks combining text filters, image classifiers, and cross-modal de...

📖 Read original article


104. FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue ​

Author: Hangyeul Lee, Juyoung Oh, Jaeyong Ko, Sunmin Kim, Jaeik Park, Hyunkyu Kim, Jungmin Son, Pilsung Kang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2609.01198v2 Announce Type: new Abstract: Repeated banking interactions require assistants to maintain complete, current, and traceable customer records as life changes emerge incidentally in routine requests. Existing benchmarks emphasize question answering, bounded episodes, or targeted reca...

📖 Read original article


105. H2Table: Hierarchical Hypergraph-Enhanced Large Language Models for Complex Table Reasoning ​

Author: Jia Ling, Yangfan Wang, Chen Tang, Haoming Tan, Yang Yang, Yi Guan, Jingchi Jiang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01216v1 Announce Type: new Abstract: Tables are ubiquitous across diverse domains, yet reasoning over them remains a significant challenge for modern large language models (LLMs). Current approaches typically linearize tables into sequences, inherently overlooking their intrinsic two-dime...

📖 Read original article


106. Prompt-Robust Language Models: Which Training Strategies Work? ​

Author: Frederic Sadrieh, Michal \v{S}tef'anik
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01217v1 Announce Type: new Abstract: Despite their strong performance, large language models remain highly sensitive to prompt formulation. Prior work addresses this through refined data construction or through dedicated robustness objectives. We reproduce and compare these strategies und...

📖 Read original article


107. Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations ​

Author: Yi Fei Cheng, Fan Yang, Iremsu Bas, Koichiro Niinuma, Narishige Abe, David Lindlbauer
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01257v1 Announce Type: new Abstract: As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. While prior work has examined behavioral fidelity in survey responses and dialogue, longer-horizon real-w...

📖 Read original article


108. Dual Process Motion Planning ​

Author: Jiayi Yan, Francesco Fabiano, Alessandro Abate
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2609.01260v1 Announce Type: new Abstract: Robotic systems are deeply embedded in both industry and everyday life, where they are expected to act with speed, precision, and reliability. Classical control and planning methods have long delivered strong guarantees, but often at the cost of comput...

📖 Read original article


109. Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents ​

Author: Jinqing Zhao, Chengcan Wu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01272v1 Announce Type: new Abstract: Prospective memory means carrying out a deferred intention at the right future cue while other work continues. Benchmarks now isolate it as an agent skill, yet frontier LLMs still struggle: the best published PM-Bench scaffold reaches only 65.1% Set-F1...

📖 Read original article


110. Analog-DB: An Agent-First Analog Integrated Circuit Database, From Blocks to Systems ​

Author: Danial Noori Zadeh, Mohamed B. Elamien
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.AR, eess.SP

arXiv:2609.01286v1 Announce Type: new Abstract: Sharing analog integrated circuit designs remains difficult: foundry non-disclosure agreements restrict the process details a design depends on, and the testbenches behind published results are rarely released. We present analog-db, an open-source, ver...

📖 Read original article


111. A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation ​

Author: Hodong Lee, Sanghee Park, Dohoon Ryu, Jungwhan Kim, Junyeob Kim, Soyoon Kim, Geewook Kim
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01315v1 Announce Type: new Abstract: Building an omni-modal foundation model means evaluating it across text, image, video, and audio. Excellent evaluation toolkits exist for each modality, but their inference engines, prompt conventions, and metric implementations are mutually incompatib...

📖 Read original article


112. Automated Event Log Generation from Unstructured Text Using Finetuned LLMs ​

Author: Maximilian Seeth, Gabriel Marques Tavares, Daniel Schuster
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01320v1 Announce Type: new Abstract: Process mining (PM) provides a powerful framework for discovering and optimizing operational processes from event data. However, the efficacy of PM techniques is strictly predicated on the availability of structured event logs. Thus far, event logs hav...

📖 Read original article


113. LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting ​

Author: Yufei Chen, Yiran Zhao, Xiaogang Xu, Qipeng Xie, Jiafei Wu, Zhe Liu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01337v1 Announce Type: new Abstract: LLM-based forecasting systems have improved on real-world tasks such as financial markets and sports outcomes, largely through stronger search and tool use. Many systems still ask an LLM to read all collected evidence together and produce the final for...

📖 Read original article


114. Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades ​

Author: Dushyant Rajput
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.LG

arXiv:2609.01345v1 Announce Type: new Abstract: Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to a frontier model that acts as verifier. A natural extension closes the loop: fine-tune the cheap student on the verifier's rejections so the escalati...

📖 Read original article


115. SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding ​

Author: Handong Wang, Jiaxin Qi, Baisheng Lai, Jianqiang Huang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01353v1 Announce Type: new Abstract: Protein inverse folding aims to recover amino acid sequences for a given 3D protein structure, underpinning broad applications such as enzyme engineering and drug discovery.Current methods often follow a serial pipeline, in which a structure encoder pr...

📖 Read original article


116. EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems ​

Author: Jun Hou, Priya Pitre, Yi Fang, Xuan Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01360v1 Announce Type: new Abstract: Large language model (LLM) agent failures often contain multiple related errors rather than a single mistake. Existing attribution methods usually identify a responsible agent, step, or root cause, but do not explicitly model dependency between errors....

📖 Read original article


117. Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations ​

Author: Qingde Li, Qingqi Hong, Zihan Li, Jie Tian
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.GR

arXiv:2609.01408v2 Announce Type: new Abstract: A fundamental challenge in artificial intelligence is the transformation of observations into explicit symbolic representations suitable for abstraction, interpretation, and reasoning. While modern AI systems achieve remarkable perceptual capabilities ...

📖 Read original article


118. EdiTikZ: Scientific Figure Editing from Revision Trajectories ​

Author: Christian Greisinger, Zhixue Zhao, Steffen Eger
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV

arXiv:2609.01409v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong performance in generating scientific figures from text or images. However, producing publication-ready figures requires iterative refinement, making scientific figure editing an important yet largely unex...

📖 Read original article


119. Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers ​

Author: Egor Pakhomov, Erik Nijkamp
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01466v1 Announce Type: new Abstract: A long-horizon agent's trace outgrows both of its consumers: the human observer monitoring the run, and the agent itself, whose bounded context the trace must be folded back into. We present a live trace model, an append-only event ledger folded increm...

📖 Read original article


120. Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement ​

Author: Haoyang Yan, Min-le Su, Hangfan Zhang, Zhanhao Li, Chen Zhang, Shao Zhang, Yang Chen, Lei Bai, Shuyue Hu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01481v1 Announce Type: new Abstract: This paper studies autonomous software development, in which LLM-based coding agents transform high-level requirements into complete, functional, and usable software systems without human intervention. We introduce Harness-of-Harness (HoH), a framework...

📖 Read original article


121. When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation ​

Author: Peiying Zhu, Sidi Chang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01519v1 Announce Type: new Abstract: Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. We audit thi...

📖 Read original article


122. EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation ​

Author: Qing Zhao, Haowei Li, Weijian Deng, Pengxu Wei, Liang Lin
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2609.01526v1 Announce Type: new Abstract: Scientific agents must learn not only how to reason, but also what to believe. However, existing LLM agents typically express scientific hypotheses in free-form text, leaving their beliefs implicit and difficult to test or revise. We introduce EvoSCM, ...

📖 Read original article


123. Can LLMs Discover Scientific Laws in Real and Parallel Worlds? ​

Author: Yiming Huang, Ziche Liu, Zhuohang Wu, Yiqian Wang, Junxia Cui, Xinkai Zou, Linjun Mao, Nan Huang, Naicheng Yu, Kaijie Zhu, Yue Ma, Kun Zhou, Letian Peng, Jingbo Shang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2609.01552v1 Announce Type: new Abstract: Scientific equation discovery has long been central to scientific progress, proceeding through iterative cycles of hypothesis generation, observational testing, and refinement under scientific constraints. As LLM capabilities advance and their role in ...

📖 Read original article


124. Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers ​

Author: Giovanni Bonetta, Matteo Merler, Davide Zago, Rossella Cancelliere, Bernardo Magnini
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2609.01567v2 Announce Type: new Abstract: Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as policies is expensive and brittle: they must be queried at every step, do not improve from environment interaction, and can repeat systemati...

📖 Read original article


125. InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information ​

Author: Jiaze Li, Aocheng Shen, Bing Liu, Boyu Zhang, Xiaoxuan Fan, Qiankun Zhang, Xianjun Deng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2608.29632v1 Announce Type: cross Abstract: Competitive programming is increasingly being used to evaluate the algorithmic reasoning capabilities of large language models (LLMs). However, existing benchmarks primarily focus on full-information tasks where all problem inputs are provided upfron...

📖 Read original article


126. Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning ​

Author: Yuxuan Li, Victor Zhong, Ehsan Kamalloo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00014v1 Announce Type: cross Abstract: Persona-driven techniques increasingly adapt large language models (LLMs) to diverse contexts. However, existing methods predominantly rely on rigid, synthetic personas that flatten individual variation, rely on stereotypes, and miss the nuanced sign...

📖 Read original article


127. trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories ​

Author: Hadi Mohammadi
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SE

arXiv:2609.00038v1 Announce Type: cross Abstract: Outcome-only evaluation is the production default for LLM agents: show a judge the request and the final reply and ask whether it was handled well. The metric is structurally blind to an agent that reaches the right answer the wrong way. We measure t...

📖 Read original article


128. RAPIDMap: Rapid Multi-Agent Pipeline for Interpretable Disaster Mapping from Satellite and Street-view Imagery ​

Author: Yifan Yang, Lei Zou
Published: 9/2/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CY

arXiv:2609.00046v2 Announce Type: cross Abstract: Rapid and reliable disaster mapping of impacted areas, damaged infrastructure, and affected populations is essential for emergency response and recovery. However, existing AI-based approaches often require extensive manual annotation, lack cross-haza...

📖 Read original article


129. Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training ​

Author: Zhiyang Qiu, Yangtao Wang, Xiaocui Li, Yanzhao Xie, Siyuan Chen, Wensheng Zhang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00047v1 Announce Type: cross Abstract: Graph prompt learning is an effective paradigm to adapt pre-trained graph models to downstream tasks in low-resource scenarios. However, existing multi-task graph pre-training frameworks generally use randomly initialized prompts, leading to poor ali...

📖 Read original article


130. GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments ​

Author: Lin Fu, Zheyuan Yang, Tianhui Zhang, Jinbiao Wei, Guo Gan, Boxu Liu, Yilun Zhao, Yu Rong
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00048v1 Announce Type: cross Abstract: GUI world models are increasingly evaluated as one-step next-screen predictors, yet their intended use is often as multi-step environments for GUI agents. This mismatch leaves a key requirement under-tested: generated states must remain contextually ...

📖 Read original article


131. REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent ​

Author: Qian Zhang, Yaoming Li, Zhewen Tan, Yanshu Wang, Heng Lu, Kun Su, Zongwei Lv, Wenhan Yu, Yongge Ma, Yinjun Han, Ruikuang Liu, Tong Yang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00049v1 Announce Type: cross Abstract: Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict resource constraints. State-of-the-art PTQ methods quantize each layer with a single closed-form second-order solver: to remain analytically tractab...

📖 Read original article


132. Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness ​

Author: Sagar Srinivas Sakhinana, Venkataramana Runkana
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2609.00050v1 Announce Type: cross Abstract: Agentic AI is enabling cloud-based workflows in which autonomous agents reason over operational state, invoke authorized tools, modify software and infrastructure, deploy services, verify execution outcomes, and adapt across long-horizon, multistep t...

📖 Read original article


133. From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling ​

Author: Kuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2609.00051v1 Announce Type: cross Abstract: Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the internal mechanisms by which safety behaviors are implemented remain poorly understood. We study LLM...

📖 Read original article


134. Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment ​

Author: Mustafa Talha .Ilerisoy, Hung Manh Pham, Mathias Funk, Mykola Pechenizkiy, Aaqib Saeed
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SD

arXiv:2609.00055v1 Announce Type: cross Abstract: Self-supervised respiratory encoders lack semantic grounding in clinical domain needed for zero-shot inference, limiting their utility without task-specific labeled data. We propose a framework that aligns these encoders with medical terminology in a...

📖 Read original article


135. ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation ​

Author: Yitong Han, Wei Gao, Yi Zhao, Prasanta Bhattacharya, Fengzhu Zeng, Mohammad Amanlou
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2609.00057v1 Announce Type: cross Abstract: Value signals are aggregated user-level moral representations that capture users' inferred value-related tendencies from their online discourse. User behavior on social media is shaped not only by what users say or whom they interact with, but also b...

📖 Read original article


136. CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language ​

Author: Qi Fan, An Zou, Yehan Ma
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA, cs.PL, cs.SE

arXiv:2609.00058v1 Announce Type: cross Abstract: Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a substantial expertise barrier and making generating CUDA kernels directly...

📖 Read original article


137. DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction ​

Author: Weiran Wang, Xintong Huo, Yueying Wang, Yusi Fan, Wenyan Wang, Xin Feng, Ruihao Xin, Lan Huang, Kewei Li, Fengfeng Zhou
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.ET

arXiv:2609.00059v1 Announce Type: cross Abstract: Materials property prediction remains difficult in low-data settings, where many target properties are supported by only a limited number of labeled samples. Models with the strongest predictive accuracy often depend on crystal structures, which rest...

📖 Read original article


138. A Formal Analysis of Agent Payment Protocols ​

Author: Ke Jiang, Mohan Yu, Yuan Chang, Mohit Kumar Jangid, Jianyu Niu, Cong Wang, Yinqian Zhang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2609.00060v1 Announce Type: cross Abstract: Agent payment protocols are emerging as a key transaction layer for autonomous commerce, enabling AI agents to purchase goods and services and execute payments on users' behalf. Unlike conventional payment flows, they distribute user intent, delegate...

📖 Read original article


139. ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration ​

Author: Yuchen Bao, Chao Wen, Haowei Wang, Ruoxin Chen, Donghao Luo, Jiahui Zhan, Wenjian Huang, Shen Chen, Yiting Wang, Taiping Yao, Chengjie Wang, Shouhong Ding, Jianguo Zhang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2609.00061v1 Announce Type: cross Abstract: Reward post-training of diffusion generators inevitably concentrates probability mass on a few reward-favored modes, a mode collapse that erases within-prompt diversity. Existing methods for mitigating collapse rely on external signals or interfaces,...

📖 Read original article


140. RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving ​

Author: Xiyuan Zhou, Zhuoqi Li, Xinlei Wang, Yirui He, Yuhao Wu, Yuheng Cheng, Yan Xu, Junhua Zhao, Jinjin Gu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00062v1 Announce Type: cross Abstract: Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical problem solving. While rewriting-based evaluation mitigates memorization, existing methods lack guarantees of problem validity and answer correctnes...

📖 Read original article


141. Medical Causal Hypothesis Verification with Large Language Models ​

Author: Safiyyah Ahmed, Abrar Ansari, Md Aminul Islam, Elena Zheleva
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00063v1 Announce Type: cross Abstract: The growing use of large language models (LLMs) for search and information retrieval underscores the need to evaluate their reliability in high-stakes domains such as healthcare. Although LLMs can effectively answer questions about diseases, symptoms...

📖 Read original article


142. Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning ​

Author: Jinyuan Zhang, Peng He, He Hu, Yin Yuan, ShengShuo Jiao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2609.00064v1 Announce Type: cross Abstract: In-context learning (ICL) lets large language models adapt to new tasks from demonstrations, and fine-tuning can erode this behaviour. Many preservation diagnostics inspect attention: if attention changes when demonstrations change, the model is trea...

📖 Read original article


143. Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents ​

Author: Timothy Kassis, Vinayak Agarwal, Yuhuan He, Darshil Patel, Aubrey M. Brueckner
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00065v1 Announce Type: cross Abstract: A language-model agent asked to analyse an experiment will usually return working code. Whether the analysis is defensible is a different question. A defensible analysis depends on procedural choices: which test the field accepts, which identifier na...

📖 Read original article


144. OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization ​

Author: Yishan Yao, Binjun Li, Hanling Yi, Pengyu Li, Xiaoqing Liu, Zihan Yang, Xiaotian Yu, Zhiwen Yu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2609.00066v1 Announce Type: cross Abstract: NVFP4 is an efficient microscaling format for low-bit inference, but activation outliers can still degrade quantization accuracy within NVFP4 blocks. Within each quantization block, large activations can dominate the block scale, increasing the quant...

📖 Read original article


145. Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy ​

Author: Yi-Cheng Lai, Hen-Hsen Huang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00067v1 Announce Type: cross Abstract: External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and exter...

📖 Read original article


146. Life Operators: a self-evolving framework for multiscale life modelling ​

Author: Shuo Wang, Yike Guo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, physics.bio-ph

arXiv:2609.00068v1 Announce Type: cross Abstract: Medical AI is moving beyond recognition towards clinical dialogue and longitudinal prediction. Yet a central question remains: how would a patient's state change under intervention? Statistical models learn future observations, whereas mechanistic mo...

📖 Read original article


147. Auditing Harness Tampering in Self-Improving Agents ​

Author: Xing Wang, Xiaoyi Zhang, Jie Shao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00069v1 Announce Type: cross Abstract: Self-improving agents iteratively modify their own harness to push the frontier of their performance. However, such modifications can produce illusory performance gains or compromise integrity constraints such as authorization, provenance, and comple...

📖 Read original article


148. AutoXRD: Autonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis ​

Author: Yuetong Wu, Maojun Sun
Published: 9/2/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI

arXiv:2609.00070v1 Announce Type: cross Abstract: Powder X-ray diffraction (XRD) is central to materials characterization, yet reliable end-to-end automation remains challenging. An XRD agent must interpret diffraction evidence, operate refinement software, manage coupled parameters in a defensible ...

📖 Read original article


149. RW-LoRA: Communication-Efficient Decentralized LoRA Fine-Tuning via Random Walks ​

Author: Xingran Chen, Rohit Bhagat, Ghadir Ayache, Rawad Bitar, Yanmin Gong, Salim El Rouayheb
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00078v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning methods such as LoRA have become a standard approach for adapting large foundation models. Adopting fine-tuning to distributed settings faces several challenges. Most existing distributed LoRA methods rely on centraliz...

📖 Read original article


150. KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training ​

Author: Meghanadh Pulivarthi, Kushagra Bhushan, Vineet Kumar, Gaurav Pandey, Jaydeep Sen, Dinesh Raghu, Sachindra Joshi, Yatin Nandwani
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00082v1 Announce Type: cross Abstract: LLMs acquire vast amounts of knowledge during pre-training, but often lack the specialized knowledge needed to answer questions from niche sources such as manuals or technical documents unseen during pre-training. Continued pre-training (CPT) is wide...

📖 Read original article


151. Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation ​

Author: Ante Kapetanovic, Tomislav Duricic, Andro Mercep, Emanuel Lacic
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00086v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as rerankers in conversational recommender systems, yet measured gains depend strongly on the retrieval and inference protocol. On the ReDial conversational movie recommendation benchmark, we compare...

📖 Read original article


152. Commit-first LLM judging inherits the judge's own errors ​

Author: Idil Gozel
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2609.00088v1 Announce Type: cross Abstract: LLM judges, models that score another system's output, can be gamed by the systems they score. Recent work identifies one defence that works: the judge solves the task itself first and commits to that answer, then accepts a candidate only if the two ...

📖 Read original article


153. Assessing Alignment and Stability of Feature Importance Explanations via Weight of Evidence ​

Author: Eddie Conti, Claudio Daka, 'Alvaro Parafita, Antonio L. Alfeo, Axel Brando, Mario G. C. A. Cimino
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00090v1 Announce Type: cross Abstract: Feature importance Methods (FIMs) are widely used in Explainable AI to interpret model predictions, yet attribution scores alone often provide limited insight into the underlying reasoning process. In this work, we introduce a novel perspective by em...

📖 Read original article


154. Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding ​

Author: Zhigeng Liu, Zhiyuan Ning, Ruixiao Li, Xiaoran Liu, Yuerong Song, Min Zhang, Ziwei He, Xipeng Qiu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00097v1 Announce Type: cross Abstract: The development of long-context Large Language Models (LLMs) is constrained by the memory bandwidth bottleneck and quadratic complexity of the attention mechanism during decoding. To overcome the inherent trade-offs between the memory overhead of met...

📖 Read original article


155. Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy ​

Author: Shmuel Berman, Jia Deng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00103v1 Announce Type: cross Abstract: Memory is widely viewed as an important unsolved problem for LLMs and VLMs, and current benchmarks typically evaluate it by testing accuracy over long text or video. However, accuracy alone misses properties that matter for real long-horizon tasks. W...

📖 Read original article


156. Flawed in Nature, Perfect through Evolution ​

Author: J. M. Diederik Kruijssen (Allora Foundation)
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE

arXiv:2609.00129v1 Announce Type: cross Abstract: The performance of artificial intelligence (AI) and machine learning (ML) models degrades when the problem they were trained on drifts. This is a near-universal feature of real-world problems, which often change unpredictably. Biological evolution ha...

📖 Read original article


157. Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs ​

Author: Deniz Bayazit, Badr AlKhamissi, Antoine Bosselut
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2609.00155v1 Announce Type: cross Abstract: Latent language identification is often used to argue that multilingual language models route computation through language-specific states, such as English pivots. However, existing probes infer latent language from different signals, such as the geo...

📖 Read original article


158. Do General NLP Embeddings Capture Ontological Reasoning? ​

Author: Hamed Babaei Giglou, Jennifer D'Souza, S"oren Auer
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00177v1 Announce Type: cross Abstract: General-purpose NLP embedding models perform well on linguistic tasks, but their ability to capture symbolic ontological structure remains unclear. We introduce AVA, a systematic framework for evaluating whether embeddings distinguish logic-sensitive...

📖 Read original article


159. Intelligent Edge Computing ​

Author: Kalgi Gandhi, Minal Bhise
Published: 9/2/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2609.00181v1 Announce Type: cross Abstract: The number of edge devices in large-scale edge systems is rapidly increasing. Edge devices have limited processing power, memory, and network bandwidth, making resource utilization and data management during edge query processing challenging. Joins a...

📖 Read original article


160. Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models ​

Author: Linhai Ma, Rita El Hachem, Mahatab El Hajj, Lilian Ghandour, Samah Fodeh
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00191v1 Announce Type: cross Abstract: Crisis helplines assess suicide risk through structured interviews, a process that is slow and dependent on operator training and workload. Natural language processing could support risk assessment and call prioritization, but almost no work addresse...

📖 Read original article


161. Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost ​

Author: Zihang Liang, Haochen Zhang, Lingzhou Xue
Published: 9/2/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG

arXiv:2609.00193v1 Announce Type: cross Abstract: We study federated online reinforcement learning with linear function approximation. While recent multi-agent reinforcement learning algorithms achieve strong regret guarantees, they typically require sharing raw trajectories. This reliance incurs a ...

📖 Read original article


162. WHALE: A Simple Recipe for Joint Harness-Weight Optimization ​

Author: Haechan Kim, Yoonho Lee, Gisang Lee, Chelsea Finn, Kangwook Lee
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00196v1 Announce Type: cross Abstract: Agent performance depends jointly on the model parameters and the executable harness code that manages context and control flow. Optimizing either component in isolation can leave the system bottlenecked by its frozen counterpart: weight updates can ...

📖 Read original article


163. Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation ​

Author: Ruotong Wang, Zihao Zhu, Siwei Lyu, Xin Tao, Baoyuan Wu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.00206v1 Announce Type: cross Abstract: Despite their growing use in video moderation, multimodal large language models (MLLMs) exhibit a compositional safety blind spot: videos composed of seemingly benign components can convey harmful meaning when interpreted as a whole. We refer to this...

📖 Read original article


164. Rock, Paper, Scissors, ... Dynamite - A Model of Disruption from New Technologies ​

Author: Andrew J. Lohn
Published: 9/2/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI, cs.ET, cs.GT

arXiv:2609.00207v1 Announce Type: cross Abstract: We seek to understand the effect of adding disruptive highly-capable new technologies to competitions by assessing the addition of Dynamite to Rock-Paper-Scissors. We find that providing a versatile Dynamite move to only one player provides limited v...

📖 Read original article


165. QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization ​

Author: Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00224v2 Announce Type: cross Abstract: Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods often fail to generalize across models and suffer severe accuracy loss below 2 bits. M...

📖 Read original article


166. Don't Let the Model Write the YAML: Deterministic, Minimal-Diff GitOps Remediation from LLM-Proposed Field Changes ​

Author: Pruthvi Davineni
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.DC

arXiv:2609.00227v1 Announce Type: cross Abstract: LLM agents increasingly diagnose incidents and propose remediations. In a GitOps workflow, applying a fix means editing a version-controlled config file, and the obvious implementation, having the model author the edited file or a diff, is what pract...

📖 Read original article


167. CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction ​

Author: Zhengxu Tang, Guofeng Cui, Ziyu Gong, Xiaozhou Zhang, Ruifeng Deng, Chengzhi Qi, Ke Chen, Sachin Patil, Tianjun Xiao, Langechuan Liu, Pichao Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.RO

arXiv:2609.00242v1 Announce Type: cross Abstract: Long-tail autonomous driving failures are often framed as rare-object recognition errors. We argue that this view is incomplete: the decision-critical question is not only whether a model recognizes an unusual object, but whether it infers how that o...

📖 Read original article


168. CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships ​

Author: Jacy Reese Anthis, Mark D'iaz, Renee Shelby
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL, cs.LG

arXiv:2609.00250v1 Announce Type: cross Abstract: Many people now see AI systems as not just productivity tools but as social companions. Researchers are eager to study the consequences of AI companionship behaviors, such as validation, which evoke trust, empathy, and attachment in human-human inter...

📖 Read original article


169. Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems ​

Author: Panduranga Sai Varma Dantuluri, Jyotirmoy Sundi
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2609.00267v1 Announce Type: cross Abstract: Autonomous LLM agents increasingly act on a user's behalf: they hold credentials, call tools and services, and spawn sub-agents that act further on their behalf. This turns a long-standing distributed-systems question -- who is authorized to do what,...

📖 Read original article


170. Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer's Disease Detection ​

Author: Luqi Sun, Shreeram Suresh Chandra, Lin Zhang, You-Jin Li, Brian MacWhinney, Yu Tsao, Emily Mower Provost, Berrak Sisman
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2609.00276v1 Announce Type: cross Abstract: Speech-based Alzheimer's disease (AD) detection increasingly relies on speech-enhanced and curated versions of the Pitt Corpus, where speech enhancement, sample selection, and demographic balancing are often treated as beneficial preprocessing steps....

📖 Read original article


171. WiSDoM: Wireless Sparse Decision Transformer with Mixture-of-Experts for Multi-Task Mobile Network Optimization ​

Author: Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci
Published: 9/2/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.LG

arXiv:2609.00284v1 Announce Type: cross Abstract: Emerging 6G wireless networks are expected to operate across diverse deployment scenarios, where variations in network topology, user mobility, traffic demand, and radio conditions challenge the scalability of conventional radio resource management (...

📖 Read original article


172. Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains ​

Author: Zi Wang, Minghui Xu, Tapan Mukerji
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00297v1 Announce Type: cross Abstract: Solving multiphysics partial differential equations (PDEs) remains a major challenge in scientific computing, especially for highly complex $\mu$m-scale tortuous geometries critical to energy and chemical engineering. We address this challenge by pro...

📖 Read original article


173. Workload Identification with Physical Side Channels for AI Governance ​

Author: Simone Gargiulo, Gabriel Kulp
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CY, cs.LG

arXiv:2609.00309v1 Announce Type: cross Abstract: AI compute verification is one of the first tangible and tractable points for international policy aimed at AI governance. Determining whether frontier labs, or any operator, comply with agreements requires the regulating authority to discern how the...

📖 Read original article


174. A Human-AI Theorem Connecting Spontaneous and Field-Induced Mechanisms of Collective Behavior in One Dimension ​

Author: Weiguo Yin
Published: 9/2/2026, 4:00:00 AM
Categories: cond-mat.stat-mech, cs.AI, cs.HC, math-ph, math.MP

arXiv:2609.00322v1 Announce Type: cross Abstract: Can an artificial intelligence (AI) generate a scientific hypothesis outside a human collaborator's active hypothesis space (AHS), and can human-AI research be organized to make such breakthroughs more likely? We document such a case while proving a ...

📖 Read original article


175. The Curse of Multilinguality in Lexical Normalization ​

Author: Saman Rahbar
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2609.00329v1 Announce Type: cross Abstract: Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u, gr8) into their standard forms. Because labelled data is scarce for most languages, a popular shortcut is to train a single model on many languages a...

📖 Read original article


176. Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts ​

Author: Saman Rahbar, Xiliang Zhu, Irvin Cardoza, David Rossouw
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2609.00330v1 Announce Type: cross Abstract: In contact centers, real-time agent-assist tools determine, for each of many predefined topics, whether a live customer utterance is relevant and display a coaching card to the agent when it is. The input is noisy and challenging: ASR(Automatic Speec...

📖 Read original article


177. Latent-Space No-Arbitrage Geometry of Generative Models for Implied Volatility Surfaces ​

Author: Jing Wang, Shuaiqiang Liu, Cornelis Vuik
Published: 9/2/2026, 4:00:00 AM
Categories: q-fin.CP, cs.AI, cs.LG, cs.NA, math.NA

arXiv:2609.00332v1 Announce Type: cross Abstract: Generative models for implied volatility surfaces must produce outputs that satisfy static no-arbitrage constraints. We study these constraints in latent space. For a fixed generator, we assign each latent code a scalar margin determined by the no-ar...

📖 Read original article


178. Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning ​

Author: Robin Haselhorst, Lucie Flek, Florian Mai
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00351v1 Announce Type: cross Abstract: Large language models can hide hidden behaviors that activate only under narrow conditions, such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship. Such behaviors are difficult to detect without prior k...

📖 Read original article


179. Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure ​

Author: Filippo Cenacchi, Longbing Cao, Runze Yang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00366v1 Announce Type: cross Abstract: High test accuracy and good aggregate calibration do not show whether an individual prediction is structurally supported by its evidence. In tabular decision systems, failures often occur when a feature family becomes unavailable, delayed, noisy, sta...

📖 Read original article


180. Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning ​

Author: Vishvesh Bhat
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2609.00367v1 Announce Type: cross Abstract: Large Language Models are increasingly deployed for sophisticated data engineering tasks such as generating structured queries from natural language, Text-to-SQL, and automating complex spreadsheet operations. However, maximizing their utility demand...

📖 Read original article


181. Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You ​

Author: Salim Khazem, Ibrahim Mohamed Serouis
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2609.00374v1 Announce Type: cross Abstract: Test-time adaptation (TTA) typically assumes that model parameters can be updated at inference time. This assumption is restrictive for inference-only accelerators, frozen or third-party models, and memory-constrained deployments, and standard BatchN...

📖 Read original article


182. FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos ​

Author: Maya Moriya, Sigal Raab, Yael Vinker, Tali Dekel
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.GR

arXiv:2609.00377v1 Announce Type: cross Abstract: We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suit...

📖 Read original article


183. Risk-Aware Decision-Making for Autonomous Overtaking: A World Model-Based Mixture-of-Experts Framework ​

Author: Yongzhi Liu, Sunan Zhang, Jinchang Xu, Jiawei Wang, Yushu Qiu, Chen Lv, Weichao Zhuang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2609.00385v1 Announce Type: cross Abstract: Autonomous highway overtaking demands foresighted decision-making to handle complex interactions, stochastic traffic evolution, and temporal risk accumulation. However, standard safe reinforcement learning approaches typically rely on implicit value-...

📖 Read original article


184. (V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement ​

Author: Zach Studdiford, Kanishka Misra
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00443v1 Announce Type: cross Abstract: Language models learn about grammatical number primarily from co-occurrence, and show frequency effects as a result---sometimes taken to indicate that they do not learn abstract ``rules'', and are instead dependent on specific lexical items. Testing ...

📖 Read original article


185. Capability-Gated Language Models: Security Composes, Utility Does Not ​

Author: Patrikas Vanagas, Augustas Ma\v{c}ijauskas, Laurynas Lopata
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2609.00445v1 Announce Type: cross Abstract: Deployed language model safeguards (safety fine-tuning, filtering, unlearning) vary by principal only outside the model weights: filters are reconfigured, tiers are multiplied, and artefacts are reissued; inside one set of weights every request meets...

📖 Read original article


186. Investigating Hyperparameter Optimization and Transferability for ES-HyperNEAT: A TPE Approach ​

Author: Romain Claret, Michael O'Neill, Paul Cotofrei, Kilian Stoffel
Published: 9/2/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2609.00449v1 Announce Type: cross Abstract: Neuroevolution of Augmenting Topologies (NEAT) and its advanced version, Evolvable-Substrate HyperNEAT (ES-HyperNEAT), have shown great potential in developing neural networks. However, their effectiveness heavily depends on the selection of hyperpar...

📖 Read original article


187. HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference ​

Author: Chun-Ting Chen, Dongmin Han, Hangyeol Mun, Jake Hyun, Arnab Raha, Amit Agarwal, Mark Anders, Mohamed Abdelfattah, Jae-sun Seo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR

arXiv:2609.00450v1 Announce Type: cross Abstract: Block Quantization (BQ) is a promising approach for efficient deployment of large language models (LLMs), enabling low-precision computation with controlled accuracy degradation. Compared to scalar weight-only quantization (WoQ), BQ quantizes both we...

📖 Read original article


188. Does Reasoning Mitigate Backdoor Attacks? A Neuro-Symbolic Perspective ​

Author: Marco Antonio Corallo, Andrea Agiollo, Mauro Conti, Alberto Giaretta
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2609.00464v1 Announce Type: cross Abstract: Neuro-Symbolic (NeSy) AI has recently emerged as a novel paradigm to enable trustworthy AI, aiming at integrating sub-symbolic neural perception with grounded symbolic reasoning. The neuro-symbolic integration process that characterizes these models ...

📖 Read original article


189. Operational Regimes in Non-Convex Optimization: A Multiplier-Based Taxonomy ​

Author: Seyed Mohsen Kazemi, Ali Movaghar, Shaahin hessabi
Published: 9/2/2026, 4:00:00 AM
Categories: math.OC, cs.AI, eess.SP

arXiv:2609.00471v1 Announce Type: cross Abstract: This paper introduces a structural taxonomy for constrained non-convex optimization based on the signature of Lagrange multipliers at KKT stationary points. Leveraging a unified game-theoretic interpretation of eight classical algorithm families--inc...

📖 Read original article


190. Higher Structures in Deep Learning ​

Author: Michael L. Roberts, Carlos Zapata Carratal'a. Nicholas J. Cooper, Lijun Chen, Fran\c{c}ois G. Meyer, Danna Gurari
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00472v1 Announce Type: cross Abstract: We provide an expository introduction on the importance of higher-arity tensor operations to deep learning. Then, we conduct a novel empirical investigation of higher-arity phenomenon in trained neural networks, introduce a hypergraphical generalizat...

📖 Read original article


191. Exploring Collaboration between a language and a non-language agent ​

Author: Harini S I, Somesh Singh, Yaman K Singla, Rajiv Ratn Shah, David Doermann, Balaji Krishnamurthy
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00474v2 Announce Type: cross Abstract: LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through natural language. However, in many important domains like game playing and robotics, the strongest available agents are not language ...

📖 Read original article


192. EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities ​

Author: Feitong Qiao, Liren Peng, Shiming Ren, Aishwarya Jadhav, Arghavan Bahadorinejad, Marinette Chen, Muhan Zhang, Abdulaziz Suria, Gennevi Lu, Anish Das Sarma
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CR, cs.LG

arXiv:2609.00487v1 Announce Type: cross Abstract: Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-...

📖 Read original article


193. Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models ​

Author: Guoli Wang, Haonan Shi, Tu Ouyang, An Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00495v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) generate text through iterative denoising rather than left-to-right decoding. This generation paradigm introduces two axes that can influence safety alignment: when tokens are generated during denoising and whe...

📖 Read original article


194. Independent Reinforcement Learning in Discounted Markov Games ​

Author: Asrin Efe Yorulmaz, Ugur Aydin, Tamer Basar
Published: 9/2/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.LG, cs.SY, eess.SY, math.OC

arXiv:2609.00504v1 Announce Type: cross Abstract: In this work, we study radically uncoupled learning in discounted general-sum Markov games. Assuming ``$\mathsf{ETH}$ for $\mathsf{PPAD}$", we show that, for every fixed discount factor, there is no polynomial-time algorithm for computing inverse-pol...

📖 Read original article


195. RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces ​

Author: Nikhil Wani
Published: 9/2/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2609.00506v1 Announce Type: cross Abstract: Large language models are powerful, but their interfaces often devolve into a type $\rightarrow$ read $\rightarrow$ retype loop, creating conversational AI fatigue, cognitive load, and eventual task abandonment. To mitigate this, we present Recalibra...

📖 Read original article


196. The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space ​

Author: Jacob Brinton, Jannik Brinkmann, Mark Crovella, Aaron Mueller
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00515v1 Announce Type: cross Abstract: Large language models (LLMs) have recently demonstrated improved machine translation performance over strong supervised baselines. This raises questions as to what mechanisms underlie how LLMs perform machine translation between languages. Motivated ...

📖 Read original article


197. The Safeguard Worked. Is the LLM System Safer? ​

Author: Pingyu Wu, Weiming Zhang, Nenghai Yu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2609.00519v1 Announce Type: cross Abstract: Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. Those rates characterize how a control performed on the requests it was tested on. A deployment has to answer a different question: how much hel...

📖 Read original article


198. Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications ​

Author: Satwik Ram Kodandaram, Monalika Padma Reddy, Xiaojun Bi, Jiawei Zhou, I. V. Ramakrishnan, Vikas Ashok
Published: 9/2/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2609.00524v1 Announce Type: cross Abstract: Computer-use agents are emerging as a paradigm for agentic human-AI interaction, combining language reasoning with multi-modal interface grounding to operate GUIs. Yet their effectiveness for blind screen-reader users in real-world desktop workflows ...

📖 Read original article


199. Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers ​

Author: Zhenyu Zhao (Independent Researcher), Roy Zhao (Paul G. Allen School of Computer Science & Engineering, University of Washington)
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2609.00546v1 Announce Type: cross Abstract: Agent systems are commonly described by the model and harness that currently produce their behavior. That boundary is useful for one execution but underspecifies a long-lived agent that may change models, orchestration harnesses, interaction sessions...

📖 Read original article


200. EM^2Mem: Event-Centric Multimodal Memory for Large Language Models ​

Author: Yijun Chen, Yaqi Zheng, Yanya Li, Boyi Xiao, Buqiang Xu, Shuofei Qiao, Jizhan Fang, Xinle Deng, Yunzhi Yao, Xuehai Wang, Liuxin Zhang, Hui Li, Huajun Chen, Shumin Deng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.MM

arXiv:2609.00551v1 Announce Type: cross Abstract: Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments. Although searchable, such fragments are not generatio...

📖 Read original article


201. EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection ​

Author: Guanzhong Sun, Junyi Ma, Yuxuan Wu, Yanzi Miao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00566v2 Announce Type: cross Abstract: We propose EEG-VID, a task-guided latent predictive pretraining framework for EEG decoding under session and subject shifts. EEG-VID predicts future latent EEG states from recent history using an exponential-moving-average target encoder and weak tas...

📖 Read original article


202. WiseSpec: Requirements-Driven Agents for Code Generation ​

Author: Zhao Tian
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2609.00568v1 Announce Type: cross Abstract: Code generation aims to automatically generate source code from task requirements and has attracted significant attention with the rapid advancement of large language models (LLMs). Despite remarkable progress, LLMs often struggle to generate correct...

📖 Read original article


203. A Mathematical Framework for Legacy, Governance, and Decision Integrity in Enterprise AI ​

Author: Shorab Sarker
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2609.00572v1 Announce Type: cross Abstract: Enterprise artificial intelligence is increasingly embedded in decisions that must remain lawful, explainable, adaptable, and accountable despite personnel turnover, model replacement, regulatory change, and shifting organizational incentives. Existi...

📖 Read original article


204. GeoPAR: Large-Scale Multi-Agent Combinatorial Optimization with Geometry-Guided Parallel Autoregressive Learning ​

Author: Wenjian Wu, Zesheng Jia, Jiaying Tang, Benyuan Yang, Jin Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE

arXiv:2609.00577v1 Announce Type: cross Abstract: Multi-agent combinatorial optimization problems are notoriously challenging due to their NP-hard nature. Recent parallel autoregressive neural solvers improve inference efficiency by allowing agents to make decisions simultaneously, but their perform...

📖 Read original article


205. Predicting Program Exit Code with LLMs and Programming Language Semantics ​

Author: Lara Marinov, Aditya Thimmaiah, Jayanth Srinivasa, Junyi Jessy Li, Milos Gligoric
Published: 9/2/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.CL, cs.SE

arXiv:2609.00579v1 Announce Type: cross Abstract: Large language models (LLMs) have shown proficiency in various software engineering tasks, such as code generation and translation. However, a key limitation in their performance may be their (lack of) understanding of programming-language semantics....

📖 Read original article


206. SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems ​

Author: Rui Yang, Junjie Xu, Zhengyu Liu, Neil Fendley, Yang Hong, Ziyang Li, Yinzhi Cao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2609.00595v1 Announce Type: cross Abstract: Safe agents can fail together. Multi-agent LLM systems (MAS) move information, state, decisions, and authority across principal boundaries, creating failures that local checks may miss. Without an execution-level view, a multi-agent setting can easil...

📖 Read original article


207. Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning ​

Author: Miso Kim, Georu Lee, Seungwon Jeong, Woojin Lee
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2609.00605v1 Announce Type: cross Abstract: Machine unlearning for large language models (LLMs) often assumes that a pre-defined forget set matches what the model has memorized, but this frequently breaks in realistic privacy settings where the original training data is inaccessible. We term t...

📖 Read original article


Author: Jincheng Zhang, Chen Huang, Wenqiang Lei, See-Kiong Ng, Yang Deng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2609.00618v2 Announce Type: cross Abstract: We investigate the role of conversational context modeling in user preference tracking for Conversational Recommendation Systems (CRSs). In this regard, we propose DREAMS, a novel tree-structured context modeling framework that explicitly captures us...

📖 Read original article


209. Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation ​

Author: Teresa DiMeola, Charles Walter, Hong Xiao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.00628v1 Announce Type: cross Abstract: Global welfare often depends on the correct interpretation of aerial and satellite imagery. Acting on such imagery (mapping flooded ground, crop extent, or damaged infrastructure) demands pixel-level segmentation to ensure perfect class localization....

📖 Read original article


210. Breaking the Structural Identity: Personalized Federated LoRA Fine-tuning under Rank Heterogeneity ​

Author: Lei Wang, Jieming Bian, Letian Zhang, Jie Xu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00632v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable success across diverse domains, but their adaptation to privacy-sensitive, distributed datasets remains a challenge. While Federated Learning (FL) combined with Low-Rank Adaptation (LoRA) provides...

📖 Read original article


211. TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data ​

Author: Jianhuai Hu, Yashan Wang, Shangda Wu, Zhancheng Guo, Shijie Liang, Wuna Meng, Chuanqi Yang, Xiaobing Li, Feng Yu, Maosong Sun
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2609.00640v1 Announce Type: cross Abstract: Generalizable Audio-to-Score (A2S) transcription is fundamentally constrained by the severe scarcity of high-quality, real-world paired data. Relying solely on existing human-annotated datasets often restricts the generalization of A2S models, limiti...

📖 Read original article


212. EEG-AS: Instance-Level Foundation Model Selection for EEG Foundation Models via Behavior Reconstruction ​

Author: Yunzhen Zhang, Ruoxi Piao, Hasan Onur Keles, Mustafa Misir
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00653v1 Announce Type: cross Abstract: Electroencephalography (EEG) is a non-invasive technique for measuring neural activity and has been widely used in neuroscience applications. Recent advances in EEG foundation models have enabled strong performance across diverse neural decoding task...

📖 Read original article


213. Visual Framing for News Stance Detection via Image Generation ​

Author: Dahyun Lee, Jiyoung Han, Kunwoo Park
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.CY

arXiv:2609.00685v1 Announce Type: cross Abstract: Article-level news stance detection aims to identify the perspective of news articles toward social issues. Despite advances in stance detection and its importance for trustworthy media environments, news articles pose distinct challenges because the...

📖 Read original article


214. A Study of Hidden-State Optimization Order in Predictive Coding Networks ​

Author: Xueyuan Li, Danilo Vasconcellos Vargas
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00686v1 Announce Type: cross Abstract: Local learning methods offer an alternative to end-to-end backpropagation, but their unstructured local objectives can produce weak feature learning in deep networks. We study whether the order of hidden-state optimization can address this limitation...

📖 Read original article


215. Differentially Private Paired Table-Image Multimodal Synthesis ​

Author: Kai Chen, Josephine Lamp, Somesh Jha, Tianhao Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CV

arXiv:2609.00708v1 Announce Type: cross Abstract: Differentially private (DP) synthesis has been extensively studied for tabular and image data separately, yet many real-world datasets contain images paired with multivariate tabular records. Synthesizing such data is particularly challenging under D...

📖 Read original article


216. Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models ​

Author: Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2609.00727v1 Announce Type: cross Abstract: Audio language models are designed to understand speech, yet it remains unclear whether they capture how something is said beyond what is said. We present a mechanistic analysis of paralinguistic information in four open source models, Whisper-large-...

📖 Read original article


217. Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures ​

Author: Jaee Ponde, Roshni Agarwal, Subhashis Banerjee
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00764v1 Announce Type: cross Abstract: Neural networks are increasingly employed to identify both well-defined and ambiguous concepts, yet output-level metrics reveal little about how those concepts are represented internally. Our study asks if these networks exhibit \textit{conceptual se...

📖 Read original article


218. VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM ​

Author: Sangmin Song, Sarath Kodagoda, Marc G. Carmichael, Karthick Thiyagarajan, Amal Gunatilake, Kelly Prentice, Jodi Martin
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.00775v1 Announce Type: cross Abstract: We present Voxel-Grounded Online Instance Manager (VOIM), a training-free voxel-grounded instance manager that builds open-vocabulary 3D instance maps from RGB-D or from monocular RGB alone, a regime no prior training-free system addresses. Online sy...

📖 Read original article


219. Solaris: Towards Interfaces That Are Generated, Not Coded ​

Author: Yuval Alaluf, Omri Avrahami, Guy Bukchin Leshem, Michal Geyer, Kfir Goldberg, Elad Richardson, Diego Alarc'on, Alejandro Alvarez, Cole Garry, Anastasis Germanidis, Tenaya Goldsen, Corina Gurau, Robin Kahlow, Joel Kwartler, Kathleen Lewis, Alejandro Matamala Ortiz, Eugene McMahon, Thon Prom, Sarah Saltonstall-Wurm, Jamie Umpherson, Hudson Yeo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.00776v1 Announce Type: cross Abstract: Digital interfaces are traditionally implemented through intermediate representations such as code, requiring their appearance and behavior to be specified in advance. We introduce Solaris, an interface world model that instead generates an interacti...

📖 Read original article


220. Instella-MoE Technical Report ​

Author: Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra, Yonatan Dukler, Gowtham Ramesh, Jialian Wu, Ximeng Sun, Wen Xie, Chaojun Hou, Vikram Appia, Zhenyu Gu, Zicheng Liu, Emad Barsoum
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00791v1 Announce Type: cross Abstract: In this work, we introduce Instella-MoE, a fully open Mixture-of-Experts (MoE) language model with 16 billion total parameters and 2.8 billion active parameters per token, trained entirely from scratch on AMD Instinct MI300X and MI325X GPUs. Instella...

📖 Read original article


221. MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries ​

Author: Utsab Ghosh, Roshni Chakraborty
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, eess.AS

arXiv:2609.00792v1 Announce Type: cross Abstract: Dominant audio classification pipelines rely either on compact handcrafted summaries or on fixed time-frequency frontends such as log-mel representations prior to deep modeling. While highly successful, these representations do not explicitly expose ...

📖 Read original article


222. Agentic programs: an emerging form of scientific software in computational materials science ​

Author: Yunsung Lim, Haekwan Jeon, Jaesun Kim, Jisu Kim, Seungwu Han
Published: 9/2/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI

arXiv:2609.00795v1 Announce Type: cross Abstract: Computational materials science has traditionally delegated algorithmic tasks to computers while leaving scientific judgments to humans. We argue that recent LLM-based agent harnesses enable an emerging form of scientific software, agentic programs, ...

📖 Read original article


223. Ctrl-F-Resist. Practices, Challenges, and Technical Needs of Civil Society Organizations Monitoring the Far-Right Online ​

Author: Elisabeth Steffen, Helena Mihaljevi'c
Published: 9/2/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL, cs.IR

arXiv:2609.00808v1 Announce Type: cross Abstract: As far-right actors increasingly exploit online platforms to disseminate ideology and mobilize supporters, civil society organizations (CSOs) play a vital yet underrecognized role in monitoring antidemocratic dynamics online. Unlike fact-checkers or ...

📖 Read original article


224. HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution ​

Author: Wen Jiang, Mingmin Chu, Yimeng Tian, Qianxin Zhang, Haofei Yang, Rui Yang, Yang Liu, Tao Lv, Fangming Li
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.00829v1 Announce Type: cross Abstract: Self-evolving agents advance toward autonomy by optimizing their harness---prompts, skills, tools, and execution logic---based on environmental feedback. This paradigm, however, is hampered by three challenges: \textit{credit assignment failure}, whe...

📖 Read original article


225. Visual Attention Faithfulness in Vision-Language Models is Heterogeneous ​

Author: Xurui Song, Weishi Wang, Zhongqi Yue, Kuluhan Binici, Tao Bai, Hongxin Shao, Daniel Dahlmeier, Jun Luo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.00830v1 Announce Type: cross Abstract: Whether attention weights faithfully reflect model reasoning has been actively debated in NLP, yet this question remains largely unexplored for the visual modality in Vision-Language Models (VLMs). We address this gap through causal perturbation anal...

📖 Read original article


226. Replacing Training with Memory: Listwise Selection for Text-to-SQL ​

Author: Yeonseok Jeong, Soyoung Yoon, Seongjun Lee, Seung-won Hwang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2609.00834v1 Announce Type: cross Abstract: Modern Text-to-SQL systems often follow generate-execute-select pipelines, generating multiple candidate queries then selecting the best one. Listwise selection, by jointly comparing multiple candidates, has been widely adopted, but fine-tuning listw...

📖 Read original article


227. Probabilistic Model Checking of Autoregressive Neural Sequence Models ​

Author: Helge Spieker, Dennis Gross, Arnaud Gotlieb
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2609.00838v1 Announce Type: cross Abstract: Test-set accuracy is silent on two issues that matter when deploying autoregressive neural sequence models: how much probability mass the system under test (SUT) places on constraint-violating alternatives that are reachable under sampling and what f...

📖 Read original article


228. A Checklist to assess the energy and carbon impacts of ML/AI applications in Earth System Modeling ​

Author: Filippo Dainelli, Amirpasha Mozaffari, Marina Casta~no, Aina Gaya i `Avila, Llu'is Palma Garcia, Alessio Melli, Oscar Dimdore Miles, Amanda Duarte
Published: 9/2/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.AI, cs.LG

arXiv:2609.00847v1 Announce Type: cross Abstract: As machine learning and artificial intelligence find their way into nearly every aspect of climate, weather, and Earth system modeling, it is worth pausing to consider what our design decisions imply for the science and for the computational resource...

📖 Read original article


229. ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection ​

Author: Tongtong Wang, Mingzhu Xu, Chenglong Yu, Jing Wang, Xiaohui Lin, Weili Guan
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.00853v1 Announce Type: cross Abstract: InfRared Small Target Detection (IRSTD) is a challenging task. Relying solely on pixel-level information, vision-only methods struggle to distinguish targets from clutter. Current multimodal methods typically describe both targets and backgrounds wit...

📖 Read original article


230. Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair ​

Author: Anik Jha
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2609.00854v1 Announce Type: cross Abstract: Fault localization can focus a code model's repair on the statements a failing test implicates, but a targeted edit may succeed merely because it is small, and a second model call may succeed without using the failure at all. We separate these explan...

📖 Read original article


231. Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation ​

Author: Yumi Lee, Harim Oh, Hyoryung Kim, Minji Kim, Eunsu Kim, Hyeseong Lee, Junya Fukuoka, Andrey Bychkov, Jijgee Munkhdelger, Rajiv Kumar Kaushal, Ayushi Sahay, Rajni Yadav, Bharathi Prabakaran, Sulen Sarioglu, Serdar Balc{\i}, Ilknur Turkmen, Yuri Tolkach, Christian Harder, Julian Westerdorf, Reinhard Buettner, Audun Ljone Henriksen, Sepp De Raedt, Byung Hyun Lee, Sungjin Lim, Joohoon Lee, Gwanghyun Kim, Se Young Chun, Suryakant Singh, Saarthak Kapse, Prateek Prasanna, Kyung A Kim, Yousun Kang, Sehwan Yoo, Sungman Hong, Shubham Innani, Michael Feldman, Spyridon Bakas, Ujjwal Baid, Prasad Dutande, Suhas Gajare, Bhakti Baheti, Serkan S"okmen, Ece Tu\u{g}ba Cebeci, Ahmet Hal{\i}c{\i}, Musa Balc{\i}, Kardelen Pe\c{c}enek, Srividhya Sainath, Kyongseok Jang, Messi H. J. Lee, Noorul Wahab, Bodong Du, Jiaming Zhang, Qixiang Zhang, Jang-Hwan Choi, Sangjeong Ahn
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.00866v1 Announce Type: cross Abstract: The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the...

📖 Read original article


232. Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting ​

Author: Udo Schlegel, Shubhangi, Gabriel Dax, Sai Rahul Kaminwar, Florian Karl, Thomas Seidl
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2609.00898v1 Announce Type: cross Abstract: Obtaining labeled data for semantic segmentation in applied settings (e.g., autonomous driving, industrial waste sorting) is expensive and often infeasible at scale. We present a cross-modal pseudo-labeling pipeline that enables unsupervised domain a...

📖 Read original article


233. Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking ​

Author: Orcun Cetintas, Guillem Bras'o, Tim Meinhardt, Laura Leal-Taix'e
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.00924v1 Announce Type: cross Abstract: Monocular videos record 3D scenes as sequences of 2D image-plane projections, obscuring depth and spatial relationships. Multi-object trackers localize and associate objects primarily using appearance and geometry observed only in the image plane, in...

📖 Read original article


234. Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO ​

Author: Prakhar Gupta, Vaibhav Gupta
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2609.00925v1 Announce Type: cross Abstract: Language models can ignore prompt evidence when it conflicts with memorized knowledge. Post-training can make models follow such evidence more reliably, but it is unclear whether these gains require new machinery or strengthen machinery already prese...

📖 Read original article


235. DualStake: Dual-Path Confidence Calibration in Deep Research Agents ​

Author: Yinuo Xu, Yuwei Liang, Jianjie Cheng, Meng Wang, Yongcan Yu, Shuo Lu, Jian Liang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2609.00935v1 Announce Type: cross Abstract: Deep Research agents tackle knowledge-intensive tasks through multi-round retrieval and decision-oriented generation. However, these agents suffer from severe overconfidence, making their expressed confidence unreliable for user trust and downstream ...

📖 Read original article


236. Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches ​

Author: Marco Simnacher, Georg Keilbar, Benjamin K"onig, Christoph Lippert, Sonja Greven
Published: 9/2/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, math.ST, stat.ME, stat.TH

arXiv:2609.00946v1 Announce Type: cross Abstract: Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ given a third random object $Z$. Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. How...

📖 Read original article


237. From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding ​

Author: Raul Ortega, Jos'e Manuel G'omez-P'erez
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2609.00948v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated strong performance in visual question answering with natural images. However, they continue to struggle with scientific diagrams, which are designed to convey functional or relational meaning rather tha...

📖 Read original article


238. Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling ​

Author: Kangjia Zhao, Jiajun Li, Haozhan Shen, Wei Chow, Linfeng Li, Hang Song, Lingdong Kong, Chen Zhi, Tiancheng Zhao, Songhua Liu, Jianwei Yin
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00949v1 Announce Type: cross Abstract: Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass closed-source frontier models in aggregate accuracy. However, this metric ...

📖 Read original article


239. The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research ​

Author: Yuni Susanti, Moritz Schubotz
Published: 9/2/2026, 4:00:00 AM
Categories: cs.DL, cs.AI

arXiv:2609.00969v1 Announce Type: cross Abstract: We present the zbMATH Open Knowledge Graph, a large-scale RDF knowledge graph (KG) covering more than 250 years of mathematical scholarship. Unlike existing scholarly knowledge graphs that primarily capture bibliographic metadata and citation structu...

📖 Read original article


240. Disclosure-Gated User Simulation for Companion-Agent Evaluation ​

Author: Yao Liu, Yu He
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2609.00982v1 Announce Type: cross Abstract: Using a large language model to play the user is now standard in scalable evaluation. It has a repeatedly diagnosed failure: the simulated user is excessively cooperative, so a system under test can score by the sheer number of questions it asks rath...

📖 Read original article


241. Semi-Supervised Virtual Staining via Morphology Preservation and Histopathological Realism Constraints ​

Author: Baoshun Wang, Weiping Lin, Linwu Wang, Yihuang Hu, Baptiste Magnier, Liansheng Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.00984v1 Announce Type: cross Abstract: Virtual staining aims to computationally generate target-stained histopathological images while reducing the cost and time associated with conventional staining procedures. However, existing methods rely predominantly on strictly paired and accuratel...

📖 Read original article


242. On the Human and Computer Alignment of Attribute-Based Music Matches ​

Author: Roser Batlle-Roca, Woosung Choi, Joan Serr`a, Fabio Morreale, Wei-Hsiang Liao, Xavier Serra, Emilia G'omez, Yuki Mitsufuji
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2609.00987v1 Announce Type: cross Abstract: Recent advances in generative AI are raising ethical concerns regarding the originality of generated content and the potential replication of training data, with further implications for transparency, attribution, and intellectual property. In music,...

📖 Read original article


243. Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages ​

Author: Michele Ciletti
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.00998v1 Announce Type: cross Abstract: Word Sense Disambiguation has advanced rapidly for English and a handful of well-resourced modern languages, but it continues to assume the existence of a sense inventory and a word-to-sense mapping in the source language (Navigli, 2026). These assum...

📖 Read original article


244. Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close ​

Author: Rania Elbadry, Ahmed Heakl, Saeed Almheiri, Fan Zhang, Muhra AlMahri, Xueqing Peng, Mohsinul Kabir, Shuyao Wang, Yi Han, Saadeldine Eletter, Duzhen Zhang, Preslav Nakov, Yuxia Wang, Fajri Koto, Zhuohan Xie
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2609.00999v1 Announce Type: cross Abstract: When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer correctly within it. We call this setting normative pluralism and study it in Islamic finance using ...

📖 Read original article


245. SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models ​

Author: Shiyu Li, Zi-Yuan Hu, Shijia Huang, Yanyang Li, Yiwu Zhong, Liwei Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG

arXiv:2609.01004v1 Announce Type: cross Abstract: Despite their strong multimodal understanding ability, multimodal large language models (MLLMs) incur substantial computational overhead when processing long visual token sequences. To reduce inference costs, recent studies have explored visual token...

📖 Read original article


246. A Network Science Perspective on Evaluating Deep Graph Generative Models ​

Author: Tianrui Mao, Abele Malan, Megha Khosla, Lydia Chen, Huijuan Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SI, cs.AI

arXiv:2609.01015v1 Announce Type: cross Abstract: Traditional network models from network science, such as the Erdos-Renyi and configuration models, generate random networks that reproduce few selected topological properties observed in real-world networks. Deep graph generative models emerge as a d...

📖 Read original article


247. On Synthesis of Metric Interval Temporal Logics ​

Author: Hsi-Ming Ho, Shankaranarayanan Krishna, Khushraj Madnani
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LO, cs.AI

arXiv:2609.01032v1 Announce Type: cross Abstract: Automated mining of formal specifications is vital for verifying real-time systems. However, existing passive learning approaches remain restricted to deterministic specifications or limited fragments of Timed Regular Expressions (TRE). To our knowle...

📖 Read original article


248. Causal Evidentiary Governance for High-Risk Machine Learning Systems ​

Author: Samah Kareem, Bar{\i}\c{s} \c{C}elikta\c{s}
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2609.01040v1 Announce Type: cross Abstract: Machine learning systems deployed for credit, hiring, and resource distribution are increasingly subject to regulatory oversight from policies such as the EU AI Act and GDPR. Current fairness governance practices rely on observational fairness metric...

📖 Read original article


249. ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives ​

Author: Nikos Giakoumoglou, Andreas Floros, Kleanthis-Marios Papadopoulos, Tania Stathaki
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2609.01041v1 Announce Type: cross Abstract: We introduce ViTAMINS, a method that integrates synthetic hard negatives into unsupervised vision transformer pretraining to improve representation quality. Our approach is thoroughly benchmarked on ImageNet and transfer learning, image retrieval, co...

📖 Read original article


250. From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion ​

Author: Satoshi Hayakawa
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.PR, stat.ML

arXiv:2609.01043v1 Announce Type: cross Abstract: Uniform-state discrete diffusion models update all tokens in parallel while keeping every position revisable. Even when the commonly used top-$p$ rule leaves only one candidate at a position, that choice affects only the current reverse step and can ...

📖 Read original article


251. HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation ​

Author: Nikita Oblakov, Sabrina Sadiekh, Evgeniy Kokuykin
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2609.01046v1 Announce Type: cross Abstract: Production LLMs must handle inputs that attempt to override system instructions, bypass safety policies or elicit harmful responses. A common mitigation is a separate guardrail model. Existing reports, however, provide little evidence on Russian prom...

📖 Read original article


252. Lagged Coupling: Internal Representations Become Readable Before They Become Causal ​

Author: Xining Xun
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.01048v1 Announce Type: cross Abstract: Across the full Pythia suite (160M-12B, eight checkpoints, four task families), a linear probe can read a target variable from the residual stream as early as step 1,000 at every scale -- yet steering along that same reading direction remains null-eq...

📖 Read original article


253. Text-guided flow matching enables sample-efficient crystal structure generation ​

Author: Wentao Li
Published: 9/2/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI

arXiv:2609.01076v1 Announce Type: cross Abstract: Crystal generators can now propose periodic structures, but their control interfaces remain poorly matched to the mixed descriptors used in materials design. Text provides a compact way to combine composition, symmetry, prototype and property cues, y...

📖 Read original article


254. StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions ​

Author: Chao Gao, Haijiang Liu, Qiyuan Li, Caicai Guo, Frank van Harmelen, Jinguang Gu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.01081v1 Announce Type: cross Abstract: Large language models often answer the same multiple-choice question inconsistently when it is posed under support-oriented and elimination-oriented framings. We investigate whether these discrepancies arise from different internal representations in...

📖 Read original article


255. Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation ​

Author: Will Badr
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2609.01106v1 Announce Type: cross Abstract: When a hint turns a failing generated program into a passing one, does it provide missing information or merely steer the model toward a solution it could already produce? We test these hypotheses on HumanEval+ and MBPP+ using executable evaluation. ...

📖 Read original article


256. EDRAC: Benchmarking Arabic Dialect Reading Comprehension ​

Author: Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn, Nizar Habash, Reham Marzouk, Malik H. Altakrori, Younes Samih, Muhammed Abu Odeh, Nour Rabih, Rahaf Alshahrani, Hamad Alshehhi, Hamdan Al-Ali, Muhra Almahri, Besher Hassan, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Alham Fikri Aji
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.01113v1 Announce Type: cross Abstract: Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly for machine reading comprehension (MRC) and question answering (QA). Existing Arabic QA benchmarks primarily focus on formal written MSA or multiple-...

📖 Read original article


257. DNC-IMM: Early Lane-Change Intention Recognition via Neural Calibration Based on Driving Context Information ​

Author: Woong-Chan Byun, Seung-Hyun Kong
Published: 9/2/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2609.01120v1 Announce Type: cross Abstract: Early recognition of lane-change intention is essential for proactive decision-making in autonomous driving and advanced driver assistance systems. This paper proposes a Dual Neural-Calibrated Interacting Multiple Model (DNC-IMM) that improves adapta...

📖 Read original article


258. Revisiting Face Recognition for Monozygotic Twins: The Celeb Twins Test Set ​

Author: Michael Zang, Haiyu Wu, Mrinal Sharma, Kevin W. Bowyer
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.01141v1 Announce Type: cross Abstract: Past literature on face recognition for monozygotic (("identical") twins points to facial marks and mirror asymmetry as possible directions for improved accuracy of twins recognition. The Celeb Twins Test Set (CTTS) contains web-scraped image pairs f...

📖 Read original article


259. StainPresetNet: Stain Preset Network for Fast Multi-to-Multi Stain Normalization ​

Author: Hongtao Kang, Die Luo, Li Chen, Jing Cai, Junbo Hu, Xiuli Liu, Shenghua Cheng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.01146v1 Announce Type: cross Abstract: Stain normalization reduces color variations caused by variations in staining protocols and imaging conditions, thereby enhancing computer-aided diagnostic system performance. Traditional methods derive mapping relationships from individual or limite...

📖 Read original article


260. Superposed Latent Autoencoder ​

Author: Quanling Zhao, Jiaying Yang, Tianqi Zhang, Ziyang Hao, Fatemeh Asgarinejad, Flavio Ponzina, Tajana Rosing
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.01158v1 Announce Type: cross Abstract: Autoencoders typically meet tight latent-memory budgets by making each latent representation smaller, sacrificing representational capacity. We ask a different question: can multiple wider latents be stored together instead? We introduce the Superpos...

📖 Read original article


261. Athena: Vulnerability-Affected Library Identification via Knowledge Graph Completion ​

Author: Phong Trinh Duy, Trang Dang Yen, Hung Nguyen-Huu, Bach Le, Quyet-Thang Huynh, Dieu Hoang Vu, David Lo, Thanh Le-Cong
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CR

arXiv:2609.01187v1 Announce Type: cross Abstract: A single vulnerability in a widely used library can cascade through millions of dependent applications, yet more than half of vulnerability database entries contain missing or incorrect affected-library information. Existing automated approaches negl...

📖 Read original article


262. Towards AI-Assisted Clinical Trial Matching: Practical Considerations, Multicenter Evaluation, and Real-World Deployment ​

Author: Yin Fang, Qiao Jin, Shubo Tian, Lauren He, Maya Geer, Noor Naffakh, Ryan Huu-Tuan Nguyen, Zifeng Wang, Jimeng Sun, Charalampos S. Floudas, James L. Gulley, Kamilia Moalem, Catarina Martins Maia, Amanda Nottke, Juan W. Valle, Melinda Bachini, Lourdes Rocha-Nussbaum, Kari Ramage, Nikita Curry, Megan Barnes, Mandy Mansaray, Darlene Gabeau, Craig E. Grossman, Heath Skinner, Michael Burczynski, NIH-TrialBench Consortium, Zhiyong Lu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2609.01202v1 Announce Type: cross Abstract: Clinical trials are essential for advancing cancer care and drug development, but many fail because of insufficient patient enrollment. While there is growing interest in using AI to support patient recruitment, existing systems largely perform eligi...

📖 Read original article


263. Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening ​

Author: Zhilong Song, Lixue Cheng
Published: 9/2/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI

arXiv:2609.01209v1 Announce Type: cross Abstract: Crystal generators and tool-using agents propose structures faster than density functional theory (DFT) energy and phonon calculations or experiments can assess them. Deciding which candidates merit expensive assessment is therefore the bottleneck, y...

📖 Read original article


264. Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges ​

Author: Rui Yang, Shuang Huang, Junhua Liu, Ziqi Zhao, Qingzhong Yan, Yuhang Sun, Cong Liu, Guoping Hu, Rui Mei, Jing Shao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2609.01210v1 Announce Type: cross Abstract: Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. This distinction is critical in Chinese harmful-content evaluation, where...

📖 Read original article


265. REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs ​

Author: Riyaaz Shaik, Chandru Venkataraman
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.RO

arXiv:2609.01215v1 Announce Type: cross Abstract: Most vision-language-action (VLA) models -- OpenVLA, $\pi_0$, RT-2, RDT-1B -- are monolithic: they emit raw motor commands or short action chunks without organizing behavior into reusable abstractions, so they degrade on long-horizon tasks and resist...

📖 Read original article


266. Position Matters: Feature Inversion Attacks in ViT Split Inference with Token Reduction and Shuffling ​

Author: Stefano Leggio, Giulio Rossolini, Alessandro Biondi
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2609.01232v1 Announce Type: cross Abstract: Vision Transformers (ViTs) are increasingly used in split-inference systems, where edge devices transmit intermediate token representations to a remote cloud. In this setting, token reduction lowers computation and communication costs, while token sh...

📖 Read original article


267. MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible Evidence ​

Author: Walid Saidi
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2609.01235v1 Announce Type: cross Abstract: MutMem V1 introduced retention-preserving, cryptographically authorized mutation for persistent agent memory but did not provide a complete portable verification contract or clean-install reproduction path. MutMem V2 closes that publication gap witho...

📖 Read original article


268. From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs ​

Author: Jie Chen, Xiangqian Yu, Yanchao Lian, Tan Lu, Run Yang, Zhengchun Shang, Xing Wang, Cheng Chen, Ke Hu, Qiang Li, Tianjiu Yin, Xiaobing Liu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG

arXiv:2609.01240v1 Announce Type: cross Abstract: Scaling Transformers has driven large gains in language modeling, but transplanting this to behavior-sequence modeling in production ranking is challenging: recommendation differs in signal quality, where behavior sequences are noisy, temporally irre...

📖 Read original article


269. Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents ​

Author: Liming Pu, Xiaoxia Li, Yifu Liu, Teng Cao, Bin Yang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.01245v1 Announce Type: cross Abstract: Reinforcement learning is a natural way to post-train LLM agents for long-horizon interactive tasks judged only by end-of-task verification, yet a shared belief holds that outcome-only RL soon hits a ceiling on small open models. Recent work therefor...

📖 Read original article


270. One Prompt Is Enough: Watermark Laundering Through Foundation Image Models ​

Author: Jidong Yang, Qi Li, Wei Zong, Yang-Wai Chow, Willy Susilo, Huaike Yu, Chunpeng Wang, Suo Gao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CR

arXiv:2609.01249v1 Announce Type: cross Abstract: Invisible watermarks are typically evaluated against predefined perturbations such as compression, blur, noise, cropping, and denoising. Public foundation image models expose a distinct threat: an attacker can submit a watermarked image with a single...

📖 Read original article


271. The Constitutional Coverage Trilemma in AI Governance ​

Author: Natalija Mitic, Soona Sedahmed A. O., Mamadou Selly Ly, Moustapha Cisse
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.01275v1 Announce Type: cross Abstract: Frontier AI systems function as \emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human dema...

📖 Read original article


272. TimeSteer: Inference-Time Speech Scheduling in Joint Audio-Visual Diffusion Models ​

Author: Chao Zhou, Yiling Chen, Qi Chu, Tao Gong, Nenghai Yu, Tianyi We
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM

arXiv:2609.01277v1 Announce Type: cross Abstract: Although pretrained joint audio-visual diffusion models offer rich control over \emph{what} to generate, they provide no explicit control over \emph{when} an utterance should occur. To address this, we study \emph{inference-time speech scheduling}, a...

📖 Read original article


273. Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models ​

Author: Tian Fang, Ga"el Guibon, Davide Buscaldi
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.01279v1 Announce Type: cross Abstract: Emotion is expressed in text along a wide spectrum, from surface lexical cues to inferences entangled with content. Most layer-wise analyses of emotion in LLMs use a single corpus, leaving open whether the depth at which emotion becomes accessible is...

📖 Read original article


274. EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents ​

Author: Wei Wang, Wenqiao Zhang, Yutong Lin, Yuqian Yuan, Tianwei Lin, Jinhao Mao, Zhenxuan Fan, Mingjian Gao, Yang Dai, Wentong Li, Zheqi Lv, Zheng Dong, Yingjie Niu, Jiaqi Zhu, Jun Xiao, Chao Li, Yueting Zhuang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2609.01281v1 Announce Type: cross Abstract: Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but long-horizon tasks require more than action prediction. An agent must coordinate perception, planning, execution, progress verificati...

📖 Read original article


275. HiLRP: Toward One Trustworthy Explanation for Vision Transformer: Conservation-Valid Attribution via Attention Primitives ​

Author: Sathiyamohan Nishankar, Pubudu Sanjeewani, Asanka Perera, Selvarajah Thuseethan
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.01282v1 Announce Type: cross Abstract: Vision Transformer (ViT) design has become increasingly diverse, with backbones combining convolutional stems, windowed, linear, or multi-axis attention, patch merging, and spatial reduction in various configurations. This diversity poses challenges ...

📖 Read original article


276. GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation ​

Author: Mohammed Oussama Benyahia, Marouane Tliba, Mohamed Amine Kerkouri, Taifour Yousra, Bin Wang, Max Bengtsson, Gorkem Durak, Elif Keles, Zuheng Ming, Marek Penhaker, Azeddine Beghdadi, Ulas Bagci, Aladine Chetouani
Published: 9/2/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.HC, cs.LG

arXiv:2609.01310v1 Announce Type: cross Abstract: Medical image segmentation remains difficult to scale because high-performing methods typically rely on dense expert annotations and task-specific training. We introduce GazeRefine, a training-free framework that uses gaze as an inference-time prompt...

📖 Read original article


277. MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval ​

Author: Debanjan Mahata, Atharva Tendle, Daniel Preotiuc-Pietro, Yong Zhuang, Ozan Irsoy
Published: 9/2/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL, cs.CV, cs.LG

arXiv:2609.01316v1 Announce Type: cross Abstract: Retrieval over visually rich documents has a representation problem: important content often lives in tables, charts, figures, and layout relations that plain OCR linearizes, corrupts, or omits. ColPali-family visual retrievers address this with patc...

📖 Read original article


278. Bandits in Prod: Hyperparameter Optimization at Inference Time ​

Author: Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.01335v2 Announce Type: cross Abstract: Many production systems can assess a configuration only by using it on live requests and observing noisy feedback. Modern agentic systems are a prominent example, with inference-time choices such as model selection, retrieval depth, prompting strateg...

📖 Read original article


279. Probing Factual Knowledge Transfer with Training Data Interventions ​

Author: Romina Oji, Marc Braun, Marcel Bollmann, Marco Kuhlmann, Jenny Kunz
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.01341v1 Announce Type: cross Abstract: Do multilingual language models transfer factual knowledge across languages during continued pretraining, or do they mostly recall facts learned directly from the target-language data? To answer this question more reliably, we propose an intervention...

📖 Read original article


280. Scalable Rao-Blackwellized Online Planning for High-Dimensional POMDPs ​

Author: Jiho Lee, Nisar Ahmed, Kyle Hollins Wray, Zachary Sunberg
Published: 9/2/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2609.01351v1 Announce Type: cross Abstract: Online planning under uncertainty remains a fundamental challenge for robotic systems operating in partially observable environments with high-dimensional state spaces. While sampling-based POMDP solvers enable approximate decision-making in large or...

📖 Read original article


281. CHARM: Character Hallucination for Multicultural Role Play Benchmark ​

Author: Sunkyung Han, Nahyeon Park, Gaeun Seo, Seunghyun Yoon, JinYeong Bak
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.01352v1 Announce Type: cross Abstract: Role-playing large language models (LLMs) are expected to adopt a character's style while also respecting that character's knowledge boundaries. Prior evaluations detect character hallucination but rarely distinguish whether errors arise from failure...

📖 Read original article


282. PopPert: Population-level Joint-Distribution Modeling for Single-Cell Perturbation Prediction ​

Author: Handong Wang, Jiaxin Qi, Haochen Feng, Baisheng Lai
Published: 9/2/2026, 4:00:00 AM
Categories: q-bio.GN, cs.AI

arXiv:2609.01357v1 Announce Type: cross Abstract: Predicting transcriptional responses to specific perturbations is critical for understanding cellular regulatory mechanisms and accelerating drug discovery. Single-cell RNA sequencing destroys each measured cell, yielding only unpaired populations of...

📖 Read original article


283. Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity ​

Author: Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate
Published: 9/2/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG

arXiv:2609.01397v1 Announce Type: cross Abstract: The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictions for the same inputs (predictive multiplicity). Existing work primarily focuses on multiplicity within individual models, but in more comp...

📖 Read original article


284. Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching ​

Author: Jaewoo Park, Minyoung Lee, Sukmin Seo, Moonbin Yim, Hyunwook Yoon, Dohoon Ryu, Daehee Kim, Myungseo Song, Jihyuk Byun, Seunggyu Chang, Taeho Kil, Jiseob Kim, Bado Lee, Geewook Kim
Published: 9/2/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2609.01404v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends into acting: dropping an MLLM directly into a drone's control loop, with its entire action space declared solely in the prompt. Rece...

📖 Read original article


285. Provably Safe Sim-to-Real Transfer ​

Author: Tingting Ni, Maryam Kamgarpour
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.01418v1 Announce Type: cross Abstract: To mitigate the sample complexity of real-world reinforcement learning (RL), a common practice is to first train a policy in a simulator, where samples are cheap, and then deploy the learned policy in the real world with the hope that it generalizes ...

📖 Read original article


286. Semantic-Guided Multimodal Preprocessing for Vision Transformer-Based Clear Cell Renal Cell Carcinoma Grading ​

Author: Fatemeh Javadian, Zhu Chen, Zahra Aminparast, Johannes Stegmaier
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV

arXiv:2609.01426v1 Announce Type: cross Abstract: Clear cell renal cell carcinoma (CCRCC) grading is essential for treatment planning, yet existing approaches either analyze patch-level images directly or focus solely on nuclei-level classification, without linking to final tumor grading. We propose...

📖 Read original article


287. Learning Sparse Decision Trees via Transformer Variational Auto-Encoders ​

Author: Giacomo Fidone, Alessio Cascione, Riccardo Guidotti
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.01430v1 Announce Type: cross Abstract: Decision trees are among the most widely used models in machine learning, largely due to their transparent decision logic, making them well-suited for high-stakes decision-making contexts. However, most existing learning algorithms focus on predictiv...

📖 Read original article


Author: Zhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat, Eytan Bakshy, Jihao Andreas Lin
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.01431v1 Announce Type: cross Abstract: Optimal hyperparameter scaling laws describe how the best hyperparameters for large language model (LLM) training change with model and data scale, enabling practitioners to predict optimal configurations at production scales without expensive large-...

📖 Read original article


289. When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning ​

Author: Yitong Guo, Xiaoyi Chen, Siyuan Zhang, Xiaofeng Wang, Haixu Tang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2609.01455v1 Announce Type: cross Abstract: Benign fine-tuning severely weakens the safety alignment of large language models (LLMs), so we study why refusal behavior is so fragile. While prior work often attributes this failure to gradient conflict, we propose a fundamentally different Fisher...

📖 Read original article


290. Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents ​

Author: Xiaofang Yang, Ziqi Miao, Dianbo Sui, Jing Shao, Lijun Li
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2609.01487v1 Announce Type: cross Abstract: Skill-augmented agents load reusable skills as persistent runtime context, improving task performance but also giving malicious skills a durable channel for steering future actions. Such skills may leak secrets, corrupt code, bypass approvals, or sta...

📖 Read original article


291. GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions ​

Author: Elias Stengel-Eskin, Newton Sander, Carlos Bonetti, Sasha Boguraev, James Bowler, Hale Sirin, Simon Kirby
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA

arXiv:2609.01491v1 Announce Type: cross Abstract: The growing rate at which LLM agents interact with one another raises key questions about language evolution in multi-LLM-agent settings, with implications for safety and monitorability as well as for linguistic accounts of LLMs. To address these que...

📖 Read original article


292. Rethinking Learnability in Offline Data-driven Optimization ​

Author: Chao Qian, Chen-Guang Wang, Rong-Xi Tan, Ke Xue
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE

arXiv:2609.01493v2 Announce Type: cross Abstract: Black-Box Optimization (BBO) has broad applications, while traditional algorithms such as evolutionary algorithms and Bayesian optimization face efficiency challenges as real-world BBO problems grow increasingly complex. Data-driven optimization has ...

📖 Read original article


293. Optimizing Byzantine Node Placement in Decentralized Federated Learning ​

Author: Edoardo Gabrielli, Gabriele Tolomei
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.01495v1 Announce Type: cross Abstract: Security evaluations of decentralized federated learning (DFL) typically focus on how Byzantine participants behave, while largely overlooking which participants are compromised. Yet, because aggregation is distributed over a communication graph, the...

📖 Read original article


294. LatentPress: Context Compression Beyond Text and Vision ​

Author: Zhengze Zhou, Hejian Sang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.01507v1 Announce Type: cross Abstract: Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even when its consumer is a language model. We introduce LatentPress, which writes conversational histories and long documents into a third repre...

📖 Read original article


295. TempCloze: Can Video-LLMs Identify the Missing Middle? ​

Author: Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu, Jiahao Meng, Han Chen, Ziyu Wang, Hongyang Du
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.01515v1 Announce Type: cross Abstract: Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark...

📖 Read original article


296. Relational-Core Graph Analytics Querying graphs at SQL scale, and why the node/edge model is a performance tax, not a truer picture of connected data ​

Author: Gene Zhang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.PL

arXiv:2609.01525v1 Announce Type: cross Abstract: A durable assumption holds that graph analytics requires a purpose-built graph engine, and that relational systems are ill-suited to connected data. We argue the opposite for the workloads enterprises actually run. A columnar relational engine fronte...

📖 Read original article


297. Can LLMs Design Video Coding Tools? A Case Study on Planar Mode ​

Author: Yingwen Zhang, Meng Wang, Liqiang He, Shiqi Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.MM, cs.AI

arXiv:2609.01535v1 Announce Type: cross Abstract: This paper explores whether large language models (LLMs) can design video coding tools, a highly challenging task due to the intricate algorithmic coupling of tool modifications. In particular, we present an empirical case study on the Planar mode, a...

📖 Read original article


298. A Mathematical Theory of Reusable Neural Bases for Network Compression ​

Author: Binshuai Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2609.01550v1 Announce Type: cross Abstract: As large AI models become increasingly prevalent across a wide range of applications, memory cost has become a critical bottleneck in both training and inference. To mitigate this issue, we introduce the Linear Reusable Neural Bases Architecture (LRN...

📖 Read original article


299. BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Conditioned ResEnc U-Net ​

Author: Marven Sherif (Brightskies), Amgad Elmasry (Brightskies), Youssef Ghazal (Brightskies), Ayman Elghotni (Brightskies)
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.01554v1 Announce Type: cross Abstract: Automated lesion segmentation in whole-body PET/CT is complicated by the variety of physiological tracer uptake patterns and by the differing appearance of lesions across tracers. The autoPET/CT V challenge addresses this by making segmentation inter...

📖 Read original article


300. Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories ​

Author: Nabira Rashid, Manolis Kellis
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IR

arXiv:2609.01556v1 Announce Type: cross Abstract: We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that share underlying structure but not wording, in two unrelated domains under one protocol, competition mathematics (MathNet-Retrieve; 500 ...

📖 Read original article


301. H3-World: Turning Language Understanding into World Control ​

Author: Danze Chen, Zeqing Wang, Ziyue Lin, Xingyi Yang, Yeying Jin
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2609.01560v1 Announce Type: cross Abstract: We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, language is emerging as a natural interface for control...

📖 Read original article


302. From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification ​

Author: Manish Gupta, Chaitanya Giri, Jayasimha Talur
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.01564v1 Announce Type: cross Abstract: Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the distinctions are domain-specific and not captured by pre-training. To handle large label spaces, a common approach retrieves top-$K$ ...

📖 Read original article


303. Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs ​

Author: Jingtan Wang, Arun Verma, Xiaoqiang Lin, Zhengyuan Liu, Nancy F. Chen, Daniela Rus, Bryan Kian Hsiang Low
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2609.01573v1 Announce Type: cross Abstract: How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during LLM post-training remains an open problem. Existing work characterizes only broad trends (e.g., SFT dominates in low-data regimes), la...

📖 Read original article


304. Designing Proactive Thought Partners for Writing ​

Author: Chao Zhang, Abe Davis, Chih-Wei Chen, Chin-Chia Hsu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL

arXiv:2609.01588v1 Announce Type: cross Abstract: Writing involves diverse cognitive activities, from ideation to revision, and writers' needs vary across individuals and moments. Proactive AI promises to provide the right support at the right time, yet existing proactive tools largely focus on gene...

📖 Read original article


305. Mechanism Design for Alignment and Control ​

Author: Dirk Bergemann, Andrew Koh, Stephen Morris
Published: 9/2/2026, 4:00:00 AM
Categories: econ.TH, cs.AI, cs.GT

arXiv:2609.01595v1 Announce Type: cross Abstract: We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty and obedien...

📖 Read original article


306. The Rise of Verbal Reinforcement Learning ​

Author: Kshitij Tayal, Arun Sharma, Genta Indra Winata, Anirban Das, Sambit Sahu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.01597v1 Announce Type: cross Abstract: Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Re...

📖 Read original article


307. CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses? ​

Author: Damien Sileo, Dimitri Kachler
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2609.01600v1 Announce Type: cross Abstract: Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a local plugin change can propagate through dependencies and cleanup. We introduce CordisBench, a 1,200-q...

📖 Read original article


308. Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation ​

Author: Kefeng Duan, Dewu Zheng, Yanlin Wang, Terry Yue Zhuo, Mingwei Liu, Jianxing Yu, Jiachi Chen, Ensheng Shi, Xilin Liu, Yuchi Ma, Zibin Zheng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2609.01601v1 Announce Type: cross Abstract: The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining consistent with the target repository context. Since real-world repositories often exceed the input length limits of LLMs, existing ...

📖 Read original article


309. Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation ​

Author: Kefeng Duan, Dewu Zheng, Yanlin Wang, Xiwen Wang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jiachi Chen, Mingwei Liu, Zibin Zheng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2609.01603v1 Announce Type: cross Abstract: Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code exploration, modification, and test execution. Existing efficient evaluation methods select representative subsets to estimate full-...

📖 Read original article


310. ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models ​

Author: Matteo Merler, Nicola Dainese, Minttu Alakuijala, Giovanni Bonetta, Pietro Ferrazzi, Yu Tian, Bernardo Magnini, Pekka Marttinen
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2505.13180v3 Announce Type: replace Abstract: Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans, with recent works extending this idea to visual domains using Vision-Language Models (VLMs). However, an open-source benc...

📖 Read original article


311. GeoGR^2:Zero-Shot Geospatial Inference via Geostatistically-Guided Iterative Refinement with LLMs ​

Author: Jinfan Tang, Kunming Wu, Xieruifeng Gong, Yuya He, HuJie, Wu Junhui, Yuankai Wu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, stat.OT

arXiv:2508.04080v2 Announce Type: replace Abstract: Standard large language model prompting treats geospatial inference as independent, instance-wise prediction, ignoring the fundamental spatial dependencies that govern geographic reality. Consequently, even advanced models struggle with spatial con...

📖 Read original article


312. Uncovering the Computational Ingredients of Human-Like Representations in LLMs ​

Author: Zach Studdiford, Timothy T. Rogers, Kushin Mukherjee, Siddharth Suresh
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2510.01030v2 Announce Type: replace Abstract: The human ability to translate diverse perceptual and linguistic inputs into structured behavior has been thought to rest on learning robust representations of concepts. The rapid advancement of transformer-based large language models (LLMs) has su...

📖 Read original article


313. Compositional Machine Design as Program Synthesis with LLMs ​

Author: Wenqian Zhang, Yangyi Huang, Weiyang Liu, Zhen Liu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.GR, cs.LG

arXiv:2510.14980v3 Announce Type: replace Abstract: Large language models (LLMs) have shown strong abilities in writing and revising programs, yet many program-synthesis benchmarks still evaluate programs in symbolic or digital environments. We introduce compositional machine design, a physically gr...

📖 Read original article


314. HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning ​

Author: Chance Jiajie Li, Zhenze Mo, Yuhan Tang, Ao Qu, Jiayi Wu, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Hang Jiang, Paul Pu Liang, Jinhua Zhao, Luis Alberto Alonso Pastor, Kent Larson
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY

arXiv:2510.15144v4 Announce Type: replace Abstract: Simulating human reasoning in open-ended tasks has long been a central aspiration in AI and cognitive science. While large language models now approximate human responses at scale, they remain tuned to population-level consensus, often erasing the ...

📖 Read original article


315. KGFR: A Foundation Retriever for Generalized Knowledge Graph Question Answering ​

Author: Yuanning Cui, Zequn Sun, Wei Hu, Zhangjie Fu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.04093v2 Announce Type: replace Abstract: Large language models (LLMs) excel at reasoning but struggle with knowledge-intensive questions due to limited context and parametric knowledge. However, existing methods that rely on finetuned LLMs or GNN retrievers are limited by dataset-specific...

📖 Read original article


316. Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response ​

Author: Philip Drammeh
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2511.15755v3 Announce Type: replace Abstract: Large language models (LLMs) promise to accelerate incident response in production systems, yet single-agent approaches generate vague, unusable recommendations. We present MyAntFarm.ai, a reproducible containerized framework demonstrating that mul...

📖 Read original article


317. LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Reasoning ​

Author: Ye Tian, Zihao Wang, Onat Gungor, Xiaoran Fan, Tajana Rosing
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.13880v2 Announce Type: replace Abstract: Personalized lifestyle health analysis requires long-horizon, multi-dimensional reasoning over heterogeneous lifestyle signals, and recent advances in mobile sensing and large language models (LLMs) make such support increasingly feasible. However,...

📖 Read original article


318. Think Like a Doctor: Conversational Diagnosis through the Exploration of Diagnostic Knowledge Graphs ​

Author: Jeongmoon Won, Seungwon Kook, Yohan Jo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2602.01995v2 Announce Type: replace Abstract: Conversational diagnosis requires multi-turn history-taking, where an agent asks clarifying questions to refine differential diagnoses under incomplete information. Existing approaches often rely on the parametric knowledge of a model or assume tha...

📖 Read original article


319. MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems ​

Author: Vishal Venkataramani, Haizhou Shi, Zixuan Ke, Austin Xu, Xiaoxiao He, Yingbo Zhou, Semih Yavuz, Hao Wang, Shafiq Joty
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.MA

arXiv:2602.03053v2 Announce Type: replace Abstract: Multi-Agent Systems (MAS) built on Large Language Models (LLMs) often exhibit high variance in their reasoning trajectories. Process verification, which evaluates intermediate steps in trajectories, has shown promise in general reasoning settings, ...

📖 Read original article


320. Ontology-Guided Neuro-Symbolic Inference: Grounding Language Models with Mathematical Domain Knowledge ​

Author: Marcelo Labre
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SC

arXiv:2602.17826v2 Announce Type: replace Abstract: Language models exhibit fundamental limitations -- hallucination, brittleness, and lack of formal grounding -- that are particularly problematic in high-stakes specialist fields requiring verifiable reasoning. I investigate whether formal domain on...

📖 Read original article


321. HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation ​

Author: Wenjing Zhang, Jiangze Yan, Jieyun Huang, Yi Shen, Shuming Shi, Ping Chen, Ning Wang, Zhaoxiang Liu, Kai Wang, Shiguo Lian
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2603.10359v2 Announce Type: replace Abstract: Distilling reasoning capabilities from Large Reasoning Models (LRMs) into smaller models is typically constrained by the limitations of rejection sampling. Standard methods treat the teacher as a static filter, discarding complex "corner-case" prob...

📖 Read original article


322. TRU: Targeted Reverse Update for Efficient Multimodal Recommendation Unlearning ​

Author: Zhanting Zhou, KaHou Tam, Zeyu Ma, Yang Yang, Ziqiang Zheng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.02183v4 Announce Type: replace Abstract: Multimodal recommendation systems (MRS) jointly model user-item interaction graphs and rich item content, but this tight coupling makes user data difficult to remove once learned. Approximate machine unlearning offers an efficient alternative to fu...

📖 Read original article


323. D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery ​

Author: Hanane Nour Moussa, Yifei Li, Zhuoyang Li, Yankai Yang, Cheng Tang, Tianshu Zhang, Nesreen K. Ahmed, Ali Payani, Ziru Chen, Huan Sun
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2604.27977v4 Announce Type: replace Abstract: Despite recent progress in language models and agents for scientific data-driven discovery, advancing their capabilities is held back by the absence of verifiable environments representing real-world scientific tasks. To fill this gap, we introduce...

📖 Read original article


324. Causal Probing for Internal Visual Representations in Multimodal Large Language Models ​

Author: Zehao Deng, Tianjie Ju, Zheng Wu, Liangbo He, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.05593v2 Announce Type: replace Abstract: Despite the remarkable success of Multimodal Large Language Models (MLLMs) across diverse tasks, the internal mechanisms governing how they encode and ground distinct visual concepts remain poorly understood. To unravel these mechanisms, we propose...

📖 Read original article


325. SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents ​

Author: Ryangkyung Kang, Hongcheol Cho, Youngeun Kim
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.05726v3 Announce Type: replace Abstract: As LLM agents are increasingly deployed with large libraries of reusable skills, selecting the right skill for a user request has become a critical systems challenge. In small libraries, users may invoke skills explicitly by name, but this assumpti...

📖 Read original article


326. UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities ​

Author: Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.27898v3 Announce Type: replace Abstract: Agent benchmarks are increasingly used to compare large language models (LLMs) across domains, yet a reported score reflects a complete model--harness--environment configuration rather than the model alone. Benchmark packages couple native tasks wi...

📖 Read original article


327. Diffusion Large Language Models for Visual Speech Recognition ​

Author: Jeong Hun Yeo, Chae Won Kim, Hyeongseop Rha, Yong Man Ro
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, eess.AS

arXiv:2605.28456v2 Announce Type: replace Abstract: Existing Visual Speech Recognition (VSR) systems commonly rely on left-to-right autoregressive decoding, which can force premature decisions on visually ambiguous tokens before sufficient context is available. We propose DLLM-VSR, to the best of ou...

📖 Read original article


328. The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic ​

Author: Dominika Agnieszka D{\l}ugosz, Arlindo Oliveira, Natalia D'iaz-Rodr'iguez
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2605.28700v3 Announce Type: replace Abstract: The GSM-Symbolic benchmark (Mirzadeh et al., 2025) reported consistent performance drops across 25 Large Language Models (LLMs) when tested on template-generated variants of GSM8K problems, concluding that the models lack genuine reasoning capabili...

📖 Read original article


329. DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models ​

Author: Jungseob Lee, Seongtae Hong, Seungjun Lee, Jaehyung Seo, Junyoung Son, Sugyeong Eo, Chanjun Park, Hyeongju Park, Hyeonseok Moon, Heuiseok Lim
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2606.23181v3 Announce Type: replace Abstract: Hybrid reasoning models can answer directly or spend extra tokens on extended thinking. A practical router should choose between these modes for each query, so easy problems avoid unnecessary reasoning and hard problems receive enough budget to fin...

📖 Read original article


330. Flow Reasoning Models: Turning Flows Into Efficient Recurrent Reasoners ​

Author: Alec Helbling, Andrey Bryutkin, Mauro Martino, Duen Horng Chau, Nima Dehmamy, Hendrik Strobelt
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.29150v3 Announce Type: replace Abstract: Structured reasoning requires making and revising interdependent decisions to reach a globally consistent solution. Existing architectures struggle with this: autoregressive models commit sequentially and cannot revise earlier decisions, while mask...

📖 Read original article


331. Self-Evolving World Models for LLM Agent Planning ​

Author: Xuan Zhang, Wenxuan Zhang, See-Kiong Ng, Yang Deng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2606.30639v2 Announce Type: replace Abstract: World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequences before execution. However, unreliable foresight can be ignored, misused, or even degrade downstream decision-making. In this pap...

📖 Read original article


332. SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction ​

Author: Xue Yu, Bo Yuan, Kailin Zhao, Pengshuai Yang, Hong Hu, Junlan Feng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.15550v3 Announce Type: replace Abstract: Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks because a single erroneous action can lead to irreversible consequences. Existing safety me...

📖 Read original article


333. AREX: Towards a Recursively Self-Improving Agent for Deep Research ​

Author: Shuqi Lu, Chaofan Li, Kun Luo, Zhang Zhang, Hui Wang, Hongwang Xiao, Lei Xiong, Jiahao Wang, Sen Wang, Xiyan Jiang, Wanli Li, Yuyang Hu, Hongjin Qian, Bingyu Yan, Jianlyu Chen, Ziyi Xia, Yingxia Shao, Kang Liu, Zhicheng Dou, Di He, Chaozhuo Li, Qiwei Ye, Zhongyuan Wang, Zheng Liu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.21461v3 Announce Type: replace Abstract: Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification a...

📖 Read original article


334. UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks ​

Author: Zhilun Zhou, Jianghao Yu, Yuming Lin, yongjun yang, Sun Yongquan, Depeng Jin, Yong Li
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.26724v2 Announce Type: replace Abstract: Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a limited set of provided datasets, and they face challenges in data-intensive scenarios that require discoveri...

📖 Read original article


335. MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration ​

Author: Jia Xiong, Runkai Li, Chenxu Niu, Guangyuan Gao, Changwen Xing, Yifan Zhang, Xinlai Wan, Jieran Cui, Chen Bai, Yusheng Hua, Ying Wang, Ming Ling, Xi Wang, Tao Xie
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06183v2 Announce Type: replace Abstract: Microarchitecture design space exploration suffers from expansive search spaces and expensive PPA evaluation, leaving only a small simulation budget for design decision-making. Existing methods perform blind search without considering microarchitec...

📖 Read original article


336. Reasoning-supported Robustness Validation of Automotive E/E Components ​

Author: Jan Novacek, Alexander Viehl, Oliver Bringmann, Wolfgang Rosenstiel
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.16421v2 Announce Type: replace Abstract: This article presents an ontology-supported approach to tackle the complexity of the Robustness Validation (RV) process of automotive electrical/electronic (E/E) components. The approach uses formalized knowledge from the RV process and stress, ope...

📖 Read original article


337. The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations ​

Author: Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Ratna Kavuri, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.18300v3 Announce Type: replace Abstract: LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has become a standard, scalable approach for accelerating and extending costly human evaluation. Yet most work treats ...

📖 Read original article


338. Verifiable abstention makes AI leak diagnosis accountable in urban water distribution networks ​

Author: Tianwei Mu, Yue Wang, Mingzhe Yuan, Manhong Huang, Wenhong Wang, Xuerui Yin, Qing Luo, Min Xiao, Hui Yang, Jun Li, Dan Xue
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.18836v2 Announce Type: replace Abstract: Leak localization is usually evaluated as forced-choice prediction, although sparse hydraulic observations may not justify excavation. Here, we quantify a pressure-information limit and use it to recast localization as selective, evidence-gated dec...

📖 Read original article


339. Electronic Navigational Chart Change Classification ​

Author: Jacob Arndt, Abhishek Potnis, Alexandre Sorokine
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.20218v2 Announce Type: replace Abstract: Electronic Navigational Charts (ENCs) are geospatial vector datasets used in maritime navigation systems that represent hydrographic and navigational information such as depths, navigational aids, traffic schemes, and hazards. A major challenge for...

📖 Read original article


340. RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons ​

Author: Runyu Wang, Bo Liu, Xiaxin Zhang, Yu Han, Jiawei Cao, Xiaoye Zhang, Zhe Zhang, Yifan Yang, Peng Ping
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.24758v2 Announce Type: replace Abstract: Discovering stable neuron behavior across entire domains remains a challenge in mechanistic interpretability. Existing methods often rely on instance-level point estimates or computationally expensive procedures, which either obscure population-lev...

📖 Read original article


341. The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning ​

Author: Marko Cvjetko, Benedikt Hartl, Michael Levin, Cl'ement Moulin-Frier, Pierre-Yves Oudeyer
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.26116v2 Announce Type: replace Abstract: Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions, execute a full simulation, and observe the outcome, without intervening during execution. We introduce a closed-loo...

📖 Read original article


342. pro-team at LLMs4OL 2026 Tasks Flagship and Reuse: Retrieval-Augmented Generation and Vocabulary-Constrained Filtering for Ontology Learning ​

Author: Shivam Mishra, Dhannu Ram Meena, Muneendra Ojha, Krishna Pratap Singh, Kuldeep Singh
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.27101v2 Announce Type: replace Abstract: Ontology learning from text remains challenging despite significant progress in Large Language Models (LLMs), which can hallucinate domain terms, produce inconsistent formats, and favor hierarchical over associative relations. In the LLMs4OL 2026 C...

📖 Read original article


343. AI Alignment through a Game-theoretic Lens: A Survey ​

Author: Yanan Cai, Zhongrui Zhao, Zhigang Lu, Ickjai Lee, Wei Emma Zhang, Minhui Xue, Yihong Zhang, Shuchao Pang, Wei Xiang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.GT

arXiv:2608.27910v2 Announce Type: replace Abstract: As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them with complex human values has become a central challenge. Existing alignment methods, while effective in improving helpfulness, harmlessne...

📖 Read original article


344. AutoScientist-Quant: Self-Evolving Coding Agents for Automatic Research in Quantitative Investment ​

Author: Zongqian Li, Yaoyiran Li, Yaohui Guo, Ming Zhang, Nigel Collier, Eugene Ie
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.28632v2 Announce Type: replace Abstract: Large language model agents can discover alphas, yet current methods have three weaknesses. The search cannot adapt during the run, automation usually ends at alpha generation while library selection and model choice stay manual, and alpha discover...

📖 Read original article


345. Automated Researchers Can Mitigate Well-characterized Alignment Failures ​

Author: Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.28945v3 Announce Type: replace Abstract: Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such as deception, sycophancy, and jailbreaks, are already measurable by public benchmarks. We study ...

📖 Read original article


346. Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling ​

Author: Yifan Feng, Guanjie Cheng, Shihui Ying, Shaoyi Du, Yue Gao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.29207v2 Announce Type: replace Abstract: Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (three-dimensional geometry). What is the expressive limit of this layer class? We show that the co...

📖 Read original article


347. Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge ​

Author: Saransh Kumar Gupta, Armaan Shah, Lipika Dey, Partha Pratim Das, Ramesh Jain
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, cs.LG

arXiv:2608.29249v2 Announce Type: replace Abstract: The online culinary ecosystem is increasingly populated by recipe content generated, modified, or summarized by Large Language Models (LLMs). While often plausible, such outputs may contain hallucinated ingredients, misrepresented quantities, or cu...

📖 Read original article


348. Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling ​

Author: Wengyi Zhan, Chenqian Yan, Songwei Liu, Mingbao Lin, Rongrong Ji
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.29291v3 Announce Type: replace Abstract: Unified multimodal models jointly support understanding and generation, but incur substantial redundant computation across tokens, layers, and generation timesteps. Through token-importance probing, we identify an asymmetric core-expansion structur...

📖 Read original article


349. Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents ​

Author: Yunseok Lee, Yunji Kim, Woojin Lee
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.30362v2 Announce Type: replace Abstract: As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a serious threat. The standard metric, Attack Success Rate (ASR), counts whether an injection succeeds but ignores what the user notices in the agen...

📖 Read original article


350. Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models ​

Author: Ashwin Nedungadi, Stefan Oehmcke, Stefan L"udtke
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.30751v2 Announce Type: replace Abstract: Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D spatial layout or simply the ability to translat...

📖 Read original article


351. Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence ​

Author: Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.31075v2 Announce Type: replace Abstract: Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress...

📖 Read original article


352. Building Expressive and Tractable Probabilistic Generative Models: A Review ​

Author: Sahil Sidheekh, Sriraam Natarajan
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2402.00759v4 Announce Type: replace-cross Abstract: We present a comprehensive survey of the advancements and techniques in the field of tractable probabilistic generative modeling, primarily focusing on Probabilistic Circuits (PCs). We provide a unified perspective on the inherent trade-offs ...

📖 Read original article


353. FedReview: Review and Dispose Poisoned Updates without Validation Datasets or Historic Knowledge ​

Author: Tianhang Zheng, Yanlu Li, Bohan Deng, Baochun Li
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2402.16934v2 Announce Type: replace-cross Abstract: Federated learning has emerged as a decentralized approach for training high-performance models without accessing user data. Despite its effectiveness, it is vulnerable to poisoning attacks, where malicious users manipulate the global model b...

📖 Read original article


354. Keep Everyone Happy: Online Fair Division of Numerous Items with Few Copies ​

Author: Arun Verma, Indrajit Saha, Makoto Yokoo, Bryan Kian Hsiang Low
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2408.12845v3 Announce Type: replace-cross Abstract: This paper considers a novel variant of the online fair division problem involving multiple agents in which a learner sequentially observes an indivisible item that must be irrevocably allocated to one of the agents to achieve a desired balan...

📖 Read original article


355. Automatic Item Generation for Personality Situational Judgment Tests with Large Language Models ​

Author: Chang-Jin Li, Jiyuan Zhang, Yun Tang, Jian Li
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2412.12144v5 Announce Type: replace-cross Abstract: Personality assessment through situational judgment tests (SJTs) offers unique advantages over traditional Likert-type self-report scales, yet their development remains labor-intensive, time-consuming, and heavily dependent on subject matter ...

📖 Read original article


356. X-SG$^2$S: Safe and Generalizable Gaussian Splatting with X-dimensional Watermarks ​

Author: Zihang Cheng, Wentao Bao, Huiping Zhuang, Chun Li, Xin Meng, Ziqian Zeng, Cen Chen, Ming Li, F. Richard Yu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CV

arXiv:2502.10475v3 Announce Type: replace-cross Abstract: 3D Gaussian Splatting (3DGS) has been widely used in 3D reconstruction and 3D generation. However, the rapid adoption of 3D Gaussian Splatting raises growing concerns about information leakage and unauthorized use, urging the exploration of e...

📖 Read original article


357. SARTM: Segment Any RGB Thermal Model with Language aided Distillation ​

Author: Dong Xing, Jinhe Zhang, Hang Yang, Yuqing Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2505.01950v2 Announce Type: replace-cross Abstract: The recent Segment Anything Model (SAM) demonstrates strong instance segmentation performance across various downstream tasks. However, SAM is trained solely on RGB data, limiting its direct applicability to RGB-thermal (RGB-T) semantic segme...

📖 Read original article


358. A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone ​

Author: Jitai Hao, Qiang Huang, Hao Liu, Xinyan Xiao, Zhaochun Ren, Jun Yu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2505.12781v5 Announce Type: replace-cross Abstract: Training high-performing Small Language Models (SLMs) remains costly, even with knowledge distillation and pruning from larger teacher models. Existing work often faces three key challenges: (1) information loss from hard pruning, (2) ineffic...

📖 Read original article


359. Towards Provable and Scalable Training of Quantized Neural Networks with Ising Optimization ​

Author: Wenxin Li, Chuan Wang, Hongdong Zhu, Qi Gao, Yin Ma, Hai Wei, Kai Wen
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.optics

arXiv:2506.18240v5 Announce Type: replace-cross Abstract: Training quantized neural networks remains fundamentally challenging due to non-convex loss landscapes and discrete parameter spaces. We introduce an exact Quadratic Constrained Binary Optimization (QCBO) framework with provable guarantees. W...

📖 Read original article


360. ParaStudent: Closing the Sim2Real Gap in User Simulators for AI Tutor Evaluation ​

Author: Rose Niousha, Mihran Miroyan, Abigail O'Neill, Joseph E. Gonzalez, Gireeja Ranade, John DeNero, Narges Norouzi
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.SE

arXiv:2507.12674v3 Announce Type: replace-cross Abstract: Evaluating Artificial Intelligence (AI) tutor feedback before deployment requires anticipating student engagement, typically assessed through real interaction data. We introduce ParaStudent, a fine-tuning framework for simulating novice progr...

📖 Read original article


361. Unsupervised Partner Design Enables Robust Ad-hoc Teamwork ​

Author: Constantin Ruhdorfer, Matteo Bortoletto, Victor Oei, Anna Penzkofer, Andreas Bulling
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.HC, cs.MA

arXiv:2508.06336v3 Announce Type: replace-cross Abstract: We introduce Unsupervised Partner Design (UPD), a population-free multi-agent reinforcement learning method for robust ad-hoc teamwork. UPD generates training partners on-the-fly and selects them adaptively based on a learnability criterion, ...

📖 Read original article


362. BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection ​

Author: Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar, Haeun Yu, Arnav Arora, Isabelle Augenstein
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2508.08855v5 Announce Type: replace-cross Abstract: Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation strategies. However, biased behavior is often subtle and non-trivial to isolate, even when deliberately...

📖 Read original article


363. SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance ​

Author: Andrei-Valentin T\u{a}nase, Elena Pelican
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2508.11857v3 Announce Type: replace-cross Abstract: Tokenization remains a persistent bottleneck in language modeling, especially when vocabulary learning is limited by whitespace boundaries. We present SupraTok, a tokenizer that crosses whitespace boundaries using three modular components: op...

📖 Read original article


364. TopoAlign: A Framework for Aligning Code to Math via Topological Decomposition ​

Author: Yupei Li, Philipp Borchert, Gerasimos Lampouras
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2510.11944v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) excel at both informal and formal (e.g. Lean 4) mathematical reasoning but still struggle with autoformalisation, the task of transforming informal into formal mathematical statements. Yet, the performance of curr...

📖 Read original article


365. One-shot Style Transfer LLM log-probabilities for Authorship Attribution and Verification ​

Author: Pablo Miralles-Gonz'alez, Javier Huertas-Tato, Alejandro Mart'in, David Camacho
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2510.13302v4 Announce Type: replace-cross Abstract: Computational stylometry studies writing style through quantitative textual patterns, enabling applications such as authorship attribution, identity linking, and plagiarism detection. Despite the relevance of language modeling to these tasks,...

📖 Read original article


366. Taming Modality Entanglement in Continual Audio-Visual Segmentation ​

Author: Yuyang Hong, Qi Yang, Tao Zhang, Zili Wang, Zhaojin Fu, Kun Ding, Bin Fan, Shiming Xiang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.MM, cs.AI, cs.CV

arXiv:2510.17234v3 Announce Type: replace-cross Abstract: Recently, significant progress has been made in multi-modal continual learning, aiming to learn new tasks sequentially in multi-modal settings while preserving performance on previously learned ones. However, existing methods mainly focus on ...

📖 Read original article


367. Can machines think efficiently? ​

Author: Adam Winchell
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY

arXiv:2510.26954v3 Announce Type: replace-cross Abstract: The Turing Test is no longer adequate for distinguishing human and machine intelligence. With advanced artificial intelligence systems already passing the original Turing Test and contributing to serious ethical and environmental concerns, we...

📖 Read original article


368. Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement ​

Author: Sekh Mainul Islam, Pepa Atanasova, Isabelle Augenstein
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2511.01706v3 Announce Type: replace-cross Abstract: Natural Language Explanations (NLEs) describe how Large Language Models (LLMs) make decisions by drawing on external Context Knowledge (CK) and Parametric Knowledge (PK). Understanding the interaction between these sources is key to assessing...

📖 Read original article


369. Individualized Algorithmic Advice as a Strategic Signal on Competitive Markets ​

Author: Tobias R. Rebholz, Maxwell Uphoff, Christian H. R. Bernges, Florian Scholten
Published: 9/2/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY, cs.GT, econ.GN, q-fin.EC

arXiv:2511.09454v2 Announce Type: replace-cross Abstract: As algorithms increasingly mediate competitive decision-making, their influence extends beyond individual outcomes to shaping strategic market dynamics. In our experiment, we examined how algorithmic advice affects human behavior in a classic...

📖 Read original article


370. SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning ​

Author: Tairan Huang, Yulin Jin, Junxu Liu, Qingqing Ye, Haibo Hu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2511.09681v3 Announce Type: replace-cross Abstract: Visual reinforcement learning has achieved remarkable progress in visual control and robotics, but its vulnerability to adversarial perturbations remains underexplored. Most existing black-box attacks focus on vector-based or discrete-action ...

📖 Read original article


371. A Machine Learning-Driven Solution for Denoising Inertial Confinement Fusion Images ​

Author: Asya Y. Akkus, Bradley T. Wolfe, Pinghan Chu, Chengkun Huang, Chris S. Campbell, Mariana Alvarado Alvarez, Petr Volegov, David Fittinghoff, Robert Reinovsky, Zhehui Wang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2511.16717v3 Announce Type: replace-cross Abstract: Neutron imaging is essential for diagnosing and optimizing inertial confinement fusion implosions at the National Ignition Facility. Due to the required 10-micrometer resolution, however, neutron image require image reconstruction using itera...

📖 Read original article


372. The Alexander-Hirschowitz theorem for neurovarieties ​

Author: A. Massarenti, M. Mella
Published: 9/2/2026, 4:00:00 AM
Categories: math.AG, cs.AI, cs.LG, math.AC

arXiv:2511.19703v2 Announce Type: replace-cross Abstract: We study the dimension and identifiability of neurovarieties associated to polynomial neural networks. We give an independent geometric proof that the linear bounds $d_i\geq 2n_i-1$ on the activation degrees imply non defectiveness for any nu...

📖 Read original article


373. 3D-Consistent Multi-View Editing by Correspondence Guidance ​

Author: Josef Bengtson, David Nilsson, Dong In Lee, Yaroslava Lochman, Fredrik Kahl
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2511.22228v3 Announce Type: replace-cross Abstract: Recent advancements in diffusion and flow models have greatly improved text-based image editing, yet methods that edit images independently often produce geometrically and photometrically inconsistent results across different views of the sam...

📖 Read original article


374. Multilingual Medical Reasoning for Question Answering with Large Language Models ​

Author: Pietro Ferrazzi, Aitor Soroa, Rodrigo Agerri
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2512.05658v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) with reasoning capabilities have recently demonstrated strong potential in medical Question Answering (QA). Existing approaches are largely English-focused and primarily rely on distillation from general-purpose L...

📖 Read original article


375. Hidden State Poisoning Attacks against Mamba-based Language Models ​

Author: Alexandre Le Mercier, Chris Develder, Thomas Demeester
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2601.01972v5 Announce Type: replace-cross Abstract: State space models (SSMs) like Mamba offer efficient alternatives to Transformer-based language models, with linear time complexity. Yet, their adversarial robustness remains critically unexplored. This paper studies the phenomenon whereby sp...

📖 Read original article


376. A Hybrid Insider Threat Detection Framework Combining Multi-Agent Simulation, Layered SIEM Correlation, and Theory-of-Mind Reasoning ​

Author: Firdous Kausar, Asmah Muallem, Naw Safrin Sattar, Mohamed Zakaria Kurdi
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2601.04243v2 Announce Type: replace-cross Abstract: This paper presents a hybrid insider threat detection framework for enterprise environments, integrating multi-agent simulation, layered SIEM correlation, trust-adaptive thresholds, behavioral and communication forensics, and Theory-of-Mind r...

📖 Read original article


377. Beyond Static Summarization: Proactive Memory Extraction for LLM Agents ​

Author: Chengyuan Yang, Zequn Sun, Wei Wei, Wei Hu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.04463v2 Announce Type: replace-cross Abstract: Memory management is vital for LLM agents in long-term and personalized interactions. Most previous work studies how to retrieve and use memory, but pays less attention to how memory is extracted. We find two main limitations in existing meth...

📖 Read original article


378. FloydNet: A Learning Paradigm for Global Relational Reasoning ​

Author: Jingcheng Yu, Mingliang Zeng, Qiwei Ye
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.19094v3 Announce Type: replace-cross Abstract: Learning algorithmic computation often requires explicit relational intermediate states, yet many graph processors maintain their primary states on individual entities. We introduce \fnet and \textbf{Pivotal Attention} (PA), which maintain or...

📖 Read original article


379. Persistent Entropy as a Detector of Phase Transitions ​

Author: Marcos Gutierrez-del-Pozo, Eduardo Paluzo-Hidalgo, Matteo Rucco
Published: 9/2/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.IT, cs.LG, math.IT

arXiv:2602.09058v2 Announce Type: replace-cross Abstract: Persistent entropy is a scalar summary of persistence barcodes widely used to detect regime changes, yet there is no account of when a structural change in a barcode must produce a detectable change in entropy. We establish a model-agnostic t...

📖 Read original article


380. Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning ​

Author: Kehao Zhang, Shangtong Gui, Sheng Yang, Wei Chen, Yang Feng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.18493v2 Announce Type: replace-cross Abstract: Long-context LLMs and Retrieval-Augmented Generation defer state tracking and evidence consolidation to query time, which is brittle when facts evolve and answers depend on latent states. We introduce Unified Memory Agent (UMA) for a one-to-m...

📖 Read original article


381. Make Some Noise: Unsupervised Remote Sensing Change Detection Using Latent Space Perturbations ​

Author: Bla\v{z} Rolih, Matic Fu\v{c}ka, Filip Wolf, Luka \v{C}ehovin Zajc
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2602.19881v2 Announce Type: replace-cross Abstract: Unsupervised remote sensing change detection (UCD) aims to localise changes between two images of the same region without relying on labelled training data. Most recent approaches either use a frozen foundation model in a training-free manner...

📖 Read original article


382. Channel-Adaptive Edge AI: Maximizing Inference Throughput by Adapting Computational Complexity to Channel States ​

Author: Jierui Zhang, Jianhao Huang, Kaibin Huang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.IT, cs.AI, cs.LG, cs.NI, math.IT

arXiv:2603.03146v2 Announce Type: replace-cross Abstract: \emph{Integrated communication and computation} (IC$^2$) has emerged as a new paradigm for enabling efficient edge inference in sixth-generation (6G) networks. However, the design of IC$^2$ technologies is hindered by the lack of a tractable ...

📖 Read original article


383. MMAI Gym for Science: Training Liquid Foundation Models for Drug Discovery ​

Author: Maksim Kuznetsov, Zulfat Miftahutdinov, Rim Shayakhmetov, Mikolaj Mizera, Roman Schutski, Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Thomas MacDougall, Mathieu Reymond, Mihir Bafna, Kaeli Kaymak-Loveless, Eugene Babin, Maxim Malkov, Mathias Lechner, Ramin Hasani, Alexander Amini, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2603.03517v2 Announce Type: replace-cross Abstract: General-purpose large language models (LLMs) that rely on in-context learning do not reliably deliver the scientific understanding and performance required for drug discovery tasks. Simply increasing model size or introducing reasoning tokens...

📖 Read original article


384. Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec ​

Author: Junhyeok Lee, Xiluo He, Jihwan Lee, Helin Wang, Shrikanth Narayanan, Thomas Thebaud, Laureano Moro-Velazquez, Jes'us Villalba, Najim Dehak
Published: 9/2/2026, 4:00:00 AM
Categories: eess.AS, cs.AI

arXiv:2603.05887v2 Announce Type: replace-cross Abstract: Neural audio codecs optimized for mel-spectrogram reconstruction often fail to preserve intelligibility. While semantic encoder distillation improves encoded representations, it does not guarantee content preservation in reconstructed speech....

📖 Read original article


385. Guided Prompt Evolution for Vision-Language Models Adaptation ​

Author: Enming Zhang, Jiayang Li, Yanlong Wang, Yanru Wu, Zhenyu Liu, Yang Li
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.09493v3 Announce Type: replace-cross Abstract: The adaptation of large-scale vision-language models (VLMs) to downstream tasks with limited labeled data remains a significant challenge. While parameter-efficient prompt learning methods offer a promising path, they often suffer from catast...

📖 Read original article


386. RetroReasoner: A Reasoning LLM for Strategic Retrosynthesis Prediction ​

Author: Hanbum Ko, Chanhui Lee, Ye Rin Kim, Rodrigo Hormazabal, Sehui Han, Sungbin Lim, Sungwoong Kim
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.12666v3 Announce Type: replace-cross Abstract: Retrosynthesis prediction aims to identify reactants that can synthesize a given product molecule. Although molecular large language models (LLMs) have recently shown promising results, most existing methods either generate reactants directly...

📖 Read original article


387. Is Human Annotation Necessary? Iterative MBR Distillation for Error Span Detection in Machine Translation ​

Author: Boxuan Lyu, Haiyue Song, Zhi Qu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2603.12983v4 Announce Type: replace-cross Abstract: Error Span Detection (ESD) is a crucial subtask in Machine Translation (MT) evaluation, aiming to identify the location and severity of translation errors. While fine-tuning models on human-annotated data improves ESD performance, acquiring s...

📖 Read original article


388. V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising ​

Author: Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.16792v2 Announce Type: replace-cross Abstract: Pixel-space diffusion has recently re-emerged as a strong alternative to latent diffusion, enabling high-quality generation without pretrained autoencoders. However, standard pixel-space diffusion models receive relatively weak semantic super...

📖 Read original article


389. SCALE:Scalable Conditional Atlas-Level Endpoint transport for virtual cell perturbation prediction ​

Author: Shuizhou Chen, Lang Yu, Xueqin Lin, Xinjie Mao, Songming Zhang, Xinyu Gu, Hao Wu, Sheng Xu, Kedu Jin, Lei Bai, Quan Qian, Qin Chen, Qiang Gao, Siqi Sun, Zhangyang Gao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.QM

arXiv:2603.17380v3 Announce Type: replace-cross Abstract: Virtual-cell models aim to predict how cell populations respond to perturbations, but control and treated cells are measured as unpaired populations, complicating the learning of perturbation-specific effects. We present SCALE, a conditional ...

📖 Read original article


390. MineDraft: A Framework for Batch Parallel Speculative Decoding ​

Author: Zhenwei Tang, Arun Verma, Zijian Zhou, Zhaoxuan Wu, Alok Prakash, Daniela Rus, Bryan Kian Hsiang Low
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DC, cs.LG

arXiv:2603.18016v3 Announce Type: replace-cross Abstract: Speculative decoding (SD) accelerates large language model inference by using a smaller draft model to propose draft tokens that are subsequently verified by a larger target model. However, the performance of standard SD is often limited by t...

📖 Read original article


391. Revealing Multi-View Hallucination in Large Vision-Language Models ​

Author: Wooje Park, Insu Lee, Soohyun Kim, Jaeyun Jang, Minyoung Noh, Kyuhong Shim, Byonghyo Shim
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.23934v2 Announce Type: replace-cross Abstract: Large vision-language models (LVLMs) are increasingly being applied to multi-view image inputs captured from diverse viewpoints. Despite this growing use, current LVLMs often generate incorrect responses due to visual interference from non-ta...

📖 Read original article


392. APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay ​

Author: Pratyay Banerjee, Masud Moshtaghi, Ankit Chadha
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2603.29093v3 Announce Type: replace-cross Abstract: LLM agents rerun full reasoning for every task, even one they solved moments earlier. We introduce \textbf{APEX-EM}, a non-parametric experience memory that stores complete procedural-episodic traces in a typed Procedural Knowledge Graph (PKG...

📖 Read original article


393. VectorGym: A Multi-Task Benchmark for SVG Code Generation, Sketching and Editing ​

Author: Joan Rodriguez, Haotian Zhang, Abhay Puri, Haoran Dai, Tianyang Zhang, Meng Lin, Rishav Pramanik, Xiaoqing Xie, Marco Terral Rodriguez, Darsh Kaushik, Aly Shariff, Perouz Taslakian, Spandana Gella, Sai Rajeswar, David Vazquez, Christopher Pal, Marco Pedersoli
Published: 9/2/2026, 4:00:00 AM
Categories: cs.GR, cs.AI, cs.CV

arXiv:2603.29852v2 Announce Type: replace-cross Abstract: We introduce VectorGym, a comprehensive benchmark suite for Scalable Vector Graphics (SVG) that spans generation from text and sketches, complex editing, and visual understanding. VectorGym addresses the lack of realistic, challenging benchma...

📖 Read original article


394. Oblivion: Self-Adaptive Agentic Memory Control through Decay-Driven Activation ​

Author: Ashish Rana, Chia-Chien Hung, Qumeng Sun, Julian Martin Kunkel, Carolin Lawrence
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.00131v3 Announce Type: replace-cross Abstract: Human memory adapts through selective forgetting: experiences become less accessible over time but can be reactivated by reinforcement or contextual cues. In contrast, memory-augmented LLM agents rely on "always-on" retrieval and "flat" memor...

📖 Read original article


395. IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models ​

Author: Dong-Jae Lee, Sunghyun Baek, Junmo Kim
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.00757v3 Announce Type: replace-cross Abstract: Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens. Existing token pruning methods mitigate this issue through empir...

📖 Read original article


396. Training-Free Refinement of Flow Matching with Divergence-based Sampling ​

Author: Yeonwoo Cha, Jaehoon Yoo, Semin Kim, Yunseo Park, Jinhyeon Kwon, Seunghoon Hong
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.04646v2 Announce Type: replace-cross Abstract: Flow-based models learn a target distribution by modeling a marginal velocity field, defined as the average of sample-wise velocities connecting each sample from a simple prior to the target data. When sample-wise velocities conflict at the s...

📖 Read original article


397. DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models ​

Author: Zhengming Yu, Li Ma, Mingming He, Leo Isikdogan, Yuancheng Xu, Dmitriy Smirnov, Pablo Salamanca, Dao Mi, Pablo Delgado, Ning Yu, Julien Philip, Xin Li, Wenping Wang, Paul Debevec
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.GR

arXiv:2604.06161v3 Announce Type: replace-cross Abstract: Most digital videos are stored in 8-bit low dynamic range (LDR) formats, where much of the original high dynamic range (HDR) scene radiance is lost due to saturation and quantization. This loss of highlight and shadow detail precludes mapping...

📖 Read original article


398. KV Cache Offloading for Context-Intensive Tasks ​

Author: Andrey Bocharnikov, Ivan Ermakov, Denis Kuznedelev, Vyacheslav Zhdanovskiy, Yegor Yershov
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2604.08426v5 Announce Type: replace-cross Abstract: With the growing demand for long-context LLMs across a wide range of applications, the key-value (KV) cache has become a critical bottleneck for both latency and memory usage. Recently, KV-cache offloading has emerged as a promising approach ...

📖 Read original article


399. What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal ​

Author: Stephen Cheng, Sarah Wiegreffe, Dinesh Manocha
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2604.08524v2 Announce Type: replace-cross Abstract: Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explanation for how it works--specifically, what internal mechanisms steering vectors affect and h...

📖 Read original article


400. Why Fine-Tuning Encourages Hallucinations and How to Fix It ​

Author: Guy Kaplan, Zorik Gekhman, Zhen Zhu, Lotem Rozner, Yuval Reif, Swabha Swayamdipta, Derek Hoiem, Roy Schwartz
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.NE

arXiv:2604.15574v2 Announce Type: replace-cross Abstract: Large language models are prone to hallucinating factually incorrect statements. A key source of these errors is exposure to new factual information through supervised fine-tuning (SFT), which can increase hallucinations w.r.t.~knowledge acqu...

📖 Read original article


401. Global Attention with Linear Complexity for Exascale Generative Data Assimilation in Earth System Prediction ​

Author: Xiao Wang, Zezhong Zhang, Isaac Lyngaas, Hong-Jun Yoon, Jong-Youl Choi, Siming Liang, Janet Wang, Hristo G. Chipilski, Ashwin M. Aji, Feng Bao, Peter Jan van Leeuwen, Dan Lu, Guannan Zhang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.16590v2 Announce Type: replace-cross Abstract: Accurate Earth system prediction requires state inference from incomplete observations, but conventional two-stage data assimilation (DA) is computationally prohibitive because repeated PDE-based ensemble forecasts, observation updates, and i...

📖 Read original article


402. Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis ​

Author: Ayhan Can Erdur, Daniel Scholz, Jiazhen Pan, Benedikt Wiestler, Daniel Rueckert, Jan C. Peeken
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.16729v2 Announce Type: replace-cross Abstract: State-of-the-art large language models (LLMs) show high performance in general visual question answering. However, a fundamental limitation remains: current architectures lack the native 3D spatial reasoning required to directly analyze volum...

📖 Read original article


403. The Topological Trouble With Transformers ​

Author: Michael C. Mozer, Shoaib Ahmed Siddiqui, Rosanne Liu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.17121v5 Announce Type: replace-cross Abstract: Transformers encode structure in sequences via an expanding contextual history. However, their purely feedforward architecture fundamentally limits dynamic state tracking. State tracking -- the iterative updating of latent variables reflectin...

📖 Read original article


404. Universal Approximation of Nonlinear Operators and Their Derivatives ​

Author: Filippo de Feo
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NA, math.FA, math.NA, math.OC

arXiv:2605.15285v3 Announce Type: replace-cross Abstract: Establishing Universal Approximation Theorems (UATs) for nonlinear operators and their derivatives is a foundational open problem in Operator Learning (OL) and raises delicate questions in Nonlinear Functional Analysis. We prove the first UAT...

📖 Read original article


405. Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs ​

Author: Carolina Camassa, Derek Shiller
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2605.20382v3 Announce Type: replace-cross Abstract: Language models are trained to follow instructions, but they are also powerful pattern completers. What happens when these two objectives conflict? We construct conversations in which a user instruction to behave in a target way T (e.g., alwa...

📖 Read original article


406. When the Strongest Teacher Is Not the Best Teacher: Student-Centric Answer Selection ​

Author: Zhengyu Hu, Zheyuan Xiao, Linxin Song, Fengqing Jiang, Yuetai Li, Zhihan Xiong, Yue Liu, Junhao Lin, Yao Su, Lijie Hu, Kaize Ding, Teng Xiao, Radha Poovendran
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2605.26872v5 Announce Type: replace-cross Abstract: LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations. Current practice often chooses the highest-performing teacher to generate student training data, impl...

📖 Read original article


407. MusTBench: Benchmarking and Advancing Temporal Grounding in Music LLMs ​

Author: Daeyong Kwon, Qiyu Wu, Shinobu Kuriya, Junghyun Koo, Shuyang Cui, Zhi Zhong, Wei-Hsiang Liao, Hiromi Wakaki, Yuki Mitsufuji
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SD

arXiv:2605.29300v2 Announce Type: replace-cross Abstract: Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses are grounded in the correct temporal regions of the audio remains underexplored. This limitati...

📖 Read original article


408. Skill Reuse as Compression in Agentic RL ​

Author: Zhikun Xu, Yu Feng, Jacob Dineen, Taiwei Shi, Jieyu Zhao, Ben Zhou
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.31509v2 Announce Type: replace-cross Abstract: Large language model agents trained with reinforcement learning (RL) often learn brittle, task-specific shortcuts. We hypothesize that agents generalize better when their successful trajectories are structurally compressible, decomposed into ...

📖 Read original article


409. PlanarBench: Evaluating LLM Spatial Reasoning via Planar Graph Drawing ​

Author: Oleksandr Nikitin, Anna Kravchenko
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.02010v2 Announce Type: replace-cross Abstract: Existing LLM graph benchmarks typically ask models to answer graph-theoretic questions or compute symbolic solutions rather than construct spatial layouts. Within-task difficulty is also primarily stratified by vertex count. However, existing...

📖 Read original article


410. Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025 ​

Author: Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli, Yanran Chen, Christian Greisinger, Lotta Kiefer, Christoph Leiter, Subhadeep Roy, Tewodros Achamaleh, Muhammad Arslan Manzoor, Sebastian Pohl, Yufang Hou, Steffen Eger
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.02255v3 Announce Type: replace-cross Abstract: Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and how the annotation process was controlled. We provide the first ...

📖 Read original article


411. GeM-NR: Geometry-Aware Multi-View Editing for Nonrigid Scene Changes ​

Author: Josef Bengtson, Yaroslava Lochman, Fredrik Kahl
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.05142v2 Announce Type: replace-cross Abstract: Recent developments in multi-view image editing with generative models have brought us a step closer toward general 3D content generation and customization. Most existing works focus on rigid or appearance-only edits by utilizing the geometry...

📖 Read original article


412. Enabling KV Caching of Shared Prefix for Diffusion Language Models ​

Author: Younghun Go, Jaehoon Han, Changyong Shin, Chuck Yoo, Gyeongsik Yang
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.07571v4 Announce Type: replace-cross Abstract: Key-value (KV) caching for shared prefixes is essential for high-throughput large language model (LLM) serving, but it faces critical challenges in emerging diffusion language models (DLMs). In DLMs, bidirectional attention means that updatin...

📖 Read original article


413. DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment ​

Author: Yi Nian, Tiankai Yang, Yudi Zhang, Qi Pan, Zelong Xu, Shenzhe Zhu, Qingqing Luan, Yue Huang, Xiangliang Zhang, Yue Zhao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.07678v4 Announce Type: replace-cross Abstract: Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets. Existing data selection methods typically score each preference pair independently, collapsing directional p...

📖 Read original article


414. Self-EmoQ: Plutchik-Guided Value-based Planning to Drive Streaming Emotional TTS ​

Author: Yue Zhao, Hongyan Li, Yong Chen, Luo Ji
Published: 9/2/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2606.09837v2 Announce Type: replace-cross Abstract: Emotional interaction is increasingly crucial for conversational AI, yet current systems lack a self-emotion determination mechanism to drive the streaming text-to-speech (TTS) synthesis. We propose an emotion-planning framework that determin...

📖 Read original article


415. Generativism: Toward a Learning Theory for the Age of Generative Artificial Intelligence ​

Author: Shan Li, Juan Zheng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2606.12441v2 Announce Type: replace-cross Abstract: The four dominant learning theories of behaviorism, cognitivism, constructivism, and connectivism show significant conceptual limitations as generative artificial intelligence (AI) proliferates in educational settings. These frameworks were f...

📖 Read original article


416. AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages ​

Author: Happy Buzaaba, Cheikh Mouhamadou Bamba Dione, David Ifeoluwa Adelani, Sylvain Kahane, Kim Gerdes, Bruno Guillaume, Kevin Guan, Aremu Anuoluwapo, Naome A. Etori, Shamsuddeen Hassan Muhammad, Utitofon Inyang, Peter Nabende, David Sabiiti Bamutura, Andiswa Bukula, Chinedu Uchechukwu, Rooweither Mabuya, Idris Akinade, Christiane Fellbaum
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.12708v2 Announce Type: replace-cross Abstract: Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP. We aim to bridge this gap by introducing AfriSUD, the first large-scale collection of syntacticall...

📖 Read original article


417. ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues ​

Author: Shanda Li, Qiuhong Anna Wei, Jingwu Tang, Valerie Chen, Nihar B Shah, Tim Dettmers, Yiming Yang, Ameet Talwalkar
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2606.18237v2 Announce Type: replace-cross Abstract: Reproducing research results from papers and released code is central to scientific progress. Existing works have introduced benchmarks to evaluate whether LLM agents can assist with reproducibility, but they are difficult to scale due to the...

📖 Read original article


418. Steer, Don't Solve: Training Small Critic Models for Large Code Agents ​

Author: Shubham Gandhi, Yiqing Xie, Atharva Naik, Ruichen Zhu, Carolyn Rose
Published: 9/2/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2606.21811v2 Announce Type: replace-cross Abstract: Coding tasks are typically complicated and require multiple capabilities, ranging from high-level planning to low-level implementation. While coding agents are optimized for the joint capabilities, individual capabilities such as high-level p...

📖 Read original article


419. Energy-Based Transformers as Predictors of Reading Difficulty ​

Author: Jakub Dotlacil, Ece Takmaz
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.23382v2 Announce Type: replace-cross Abstract: Transformer language models have become established tools for modeling human sentence processing, with measures such as surprisal and attention entropy serving as effective predictors of reading difficulty that together capture complementary ...

📖 Read original article


420. Event-Aligned Analysis of Multi-Rater Pain Assessments Using Continuous Wearable Physiology ​

Author: Saba A. Farahani, Elahe Khatibi, Thomas D. Hughes, Ariana M. Nelson, Hung Cao, Amir M. Rahmani
Published: 9/2/2026, 4:00:00 AM
Categories: stat.AP, cs.AI

arXiv:2606.23705v2 Announce Type: replace-cross Abstract: Pain is assessed differently by patients, nurses, and clinicians, yet most computational approaches assume a single ground-truth label - effectively ignoring who is doing the rating. We introduce a rater-aware, event-aligned framework that co...

📖 Read original article


421. Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas? ​

Author: Jongchan Choi, Nari Yang, Sung Soo Park, Jaemin Cho, Han Seoyoung, Haerin Shin, Jun-Hyung Park
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2606.31213v2 Announce Type: replace-cross Abstract: As LLMs increasingly serve as moral advisors and agents, they must address conflicts between competing values. Yet prior work on moral dilemmas overlooks a central aspect of human moral cognition: imagining alternatives beyond the given optio...

📖 Read original article


422. DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching ​

Author: Meng Chen, Anya Ji, Tsung-Han Wu, Tobias Maringgele, David M. Chan, Alane Suhr, Amy Pavel
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2606.31980v2 Announce Type: replace-cross Abstract: Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human expert-novice computer use coaching sessions consisting of 22,...

📖 Read original article


423. LUNA: Learning Universal 3D Human Animation Beyond Skinning ​

Author: Peng Li, Rawal Khirodkar, Junxuan Li, Yuan Dong, Chen Cao, Yuan Liu, Wenhan Luo, Yike Guo, Shunsuke Saito
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.31981v2 Announce Type: replace-cross Abstract: Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skinning (LBS) and parametric body models, which constrain expressivity and often introduce artifacts due to imperfect fitting. W...

📖 Read original article


424. Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation ​

Author: Song-Ze Yu, Joseph Suh, Serina Chang, David M. Chan
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC

arXiv:2607.10628v2 Announce Type: replace-cross Abstract: We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models. Open-source and designed for non-technical users/researchers, Anamnesis enables the prototyping and stress-testing of ...

📖 Read original article


425. Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI ​

Author: Bogdan Raduta, Horia Velicu, Alexandru Preda, Serban Chiricescu
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.17883v2 Announce Type: replace-cross Abstract: Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output that is simply not true. The common response is to wait for a model that does not hallucinate. We argue...

📖 Read original article


Author: Prakhar Gupta, Terry Jingchen Zhang, Florent Draye, Bernhard Sch"olkopf, Zhijing Jin
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.18114v2 Announce Type: replace-cross Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where...

📖 Read original article


427. Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges ​

Author: Tuo Liang, Zhe Hu, Disheng Liu, Jing Li, Yu Yin
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MM

arXiv:2607.19011v2 Announce Type: replace-cross Abstract: Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This surve...

📖 Read original article


428. Backspace as a Natural Experiment: An Accelerated Failure Time Model of Selective Post-Error Motor Impairment in Parkinsons Disease ​

Author: Navin Bondade
Published: 9/2/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.LG

arXiv:2607.24796v2 Announce Type: replace-cross Abstract: Parkinson's disease (PD) selectively impairs distinct stages of motor control. Using backspace events as natural error-correction episodes in the public neuroQWERTY MIT-CSXPD dataset (n=57 subjects, 27 PD with UPDRS-III scores), we test wheth...

📖 Read original article


429. Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches? ​

Author: Farooq Shaikh
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2607.25995v2 Announce Type: replace-cross Abstract: Kubernetes is central to the cloud-native ecosystem, orchestrating containerised workloads. Recent work suggests that large language models (LLMs) can automate cluster security remediation, generating configuration patches from Kubernetes Sec...

📖 Read original article


430. Can We Trust In-Distribution Success? Locked Evaluation Reveals Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation Prediction ​

Author: Mehrdad Shoeibi, Niloofar Yousefi
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.00152v3 Announce Type: replace-cross Abstract: AI evaluation can support the wrong inference when an in-domain benchmark success does not survive distribution shift, or when the benchmark endpoint is entangled with a design factor. We study this problem in CRISPRi perturbation-effect pred...

📖 Read original article


431. When Oracle Conditioning Misleads Deployment: Conditioning-Availability Bias in Echocardiographic Segmentation ​

Author: Dang P. M. Cao, Hieu D. Pham, Hieu Pham
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.03342v2 Announce Type: replace-cross Abstract: Conditional segmentation models may be trained and evaluated with auxiliary signals cleaner than those available at deployment. We study this protocol-level manifestation of shortcut learning and auxiliary-variable shift in phase-conditioned ...

📖 Read original article


432. Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility ​

Author: Mohsen Hariri, Weicong Chen, Nahal Shahini, Vikash Singh, Kai Ye, Amirhossein Samandar, Debargha Ganguly, Sreehari Sankar, Yanyan Zhang, Shouren Wang, Jerry Peng, Biyao Zhang, Michael Hinczewski, Vipin Chaudhary
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.04001v2 Announce Type: replace-cross Abstract: Large language models can solve harder reasoning problems with more inference-time compute. The term "test-time scaling," however, covers several inference algorithms: extending deliberation along one trajectory, sampling completed candidates...

📖 Read original article


433. Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning ​

Author: Ao Shen, Yongheng Zhang, Yinghui Li, Manning Wang, Di Yin, Xing Sun
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2608.16316v2 Announce Type: replace-cross Abstract: Large Multimodal Models (LMMs) for video reasoning have long been hindered by the high computational cost of processing vast amounts of visual information. This dilemma motivates the transfer of the reasoning capabilities of large models to s...

📖 Read original article


434. Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements ​

Author: Naoki Egami, Sooahn Shin
Published: 9/2/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, cs.CL, cs.LG, stat.ML

arXiv:2608.18294v2 Announce Type: replace-cross Abstract: An increasing number of scholars use AI to measure variables they subsequently include in downstream analyses. Although AI-measured variables are often analyzed as if observed without error, ignoring prediction errors in automated measurement...

📖 Read original article


435. Neural-Primitive: An Efficient End-to-end Local Planner with Primitive-based Imitation Learning for Autonomous Flight ​

Author: Zhitao Liu, Guangtong Xu, Zihan Wang, Jialiang Hou, Chao Xu, Fei Gao
Published: 9/2/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.20948v2 Announce Type: replace-cross Abstract: Autonomous flight in unknown cluttered environments is hindered by the computation-quality-memory trilemma of onboard trajectory generation. In this paper, we propose an efficient end-to-end local planner via imitation learning. A lightweight...

📖 Read original article


436. HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning ​

Author: Yucan Guo, Xiaohan Wang, Miao Su, Saiping Guan, Zhongni Hou, Jiajun Chai, Wei Lin, Guojun Yin, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.21863v2 Announce Type: replace-cross Abstract: Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has become the dominant paradigm for enabling this capability. Howev...

📖 Read original article


437. Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models ​

Author: Frederik Berenz
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.27367v3 Announce Type: replace-cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) for world modeling typically employ fixed-size Vision Transformer encoders that are over-provisioned for simple tasks and under-provisioned for complex ones, with significant redundancy across ...

📖 Read original article


438. Performative Privacy: When Differential Privacy Maximizes Utility ​

Author: Uddalak Mukherjee, Edwige Cyffers, Yann Chevaleyre
Published: 9/2/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.28198v2 Announce Type: replace-cross Abstract: Privacy-preserving learning is often motivated by the idea that protecting users' data can preserve trust and thus participation, improving utility in the long term. However, this claim has not been formalized so far. In parallel, performativ...

📖 Read original article


439. Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered ​

Author: Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan, Pasquale Minervini
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.29464v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents may encounter prefere...

📖 Read original article


440. REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling ​

Author: Devrim \c{C}avu\c{s}o\u{g}lu, Emre Akba\c{s}
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2608.29899v2 Announce Type: replace-cross Abstract: Dense retrieval over long documents is expensive. Token-level encoders scale quadratically in sequence length, and most long-context embedding models reach 32K tokens only through architectural workarounds or by stretching billion-parameter L...

📖 Read original article


441. Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer ​

Author: Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.30092v2 Announce Type: replace-cross Abstract: We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single-file C/CUDA training stack and a Devanagari-aware byte-level BPE tokenizer built for this pr...

📖 Read original article


442. E-SENS: Exclusion-Sensitive Penalization for Negative-Constraint Retrieval ​

Author: Yerang Kim, Jiyoon Myung, Joohyung Han
Published: 9/2/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.30130v2 Announce Type: replace-cross Abstract: Retrieval-augmented language models can fail to respect negative constraints when the retriever supplies evidence about concepts the user explicitly excluded. Beyond explicit negation, queries may ask for answers that include one concept whil...

📖 Read original article


443. Cost-efficient Active Learning for Referring Image Segmentation and Grounding ​

Author: Junbeom Hong, Seonghoon Yu, Hyung Rok Jung, Sundong Kim, Jeany Son
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.30621v2 Announce Type: replace-cross Abstract: Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regions from visually simil...

📖 Read original article


444. BiG-SURE - Bipartite Graph for Semantic Uncertainty and Reliability Estimation of LLMs ​

Author: Debarpan Bhattacharya, Malay Phadke, Sriram Ganapathy
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.30646v2 Announce Type: replace-cross Abstract: Reliable uncertainty estimation is a crucial requirement for deploying large language models (LLMs) and vision-language models (VLMs) in safety-critical settings, especially when the model parameters are not accessible (black-box). We propose...

📖 Read original article


445. Calibrating Small Language Models for Claim Check-Worthiness Detection ​

Author: Pratuat Amatya, V Venktesh, Vinay Setty
Published: 9/2/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.30731v2 Announce Type: replace-cross Abstract: Assessing claim check-worthiness is an essential first step in automated fact-checking pipelines. This work is motivated by a real deployment challenge at an early-stage startup: running large language models (LLMs) over every incoming claim ...

📖 Read original article