Skip to content

arXiv cs.AI - 2026-08-11 ​

753 items collected.


1. Towards an Argumentative Foundation for Evaluative AI ​

Author: Xiang Yin, Tim Miller, Nico Potyka, Antonio Rago, Francesca Toni
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.07473v1 Announce Type: new Abstract: Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting competing hypotheses together with evidence for and against each. In this position paper, we advocate ...

📖 Read original article


2. Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains ​

Author: Hiroki Naito
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2608.07474v1 Announce Type: new Abstract: Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains when AI output velocity V exceeds human cognitive capacity C_max. The operative constraint, however, is not V alone but V x L, where L denotes per-it...

📖 Read original article


3. Determinization in Structure Theories: A Unified Framework via Closure, Comparability, and Joint Admissibility ​

Author: Hai Hai Fu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2608.07476v1 Announce Type: new Abstract: We develop a formal framework for constructing canonical interpretations from plural structure theories. A structure theory is a triple T = ({\Sigma}, A, I) consisting of a signature, axioms, and an inference policy, whose admissible interpretation fam...

📖 Read original article


4. Emotion in an active inference model of human driving ​

Author: Julian F. Schumann, Johan Engstr"om, Ran Wei, Jens Kober, Martijn Wisse, Arkady Zgonnikov
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.LG, cs.RO

arXiv:2608.07480v1 Announce Type: new Abstract: Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-directed action with uncertainty reduction. It has been successfully applied across biological and artificial systems, including recent work on huma...

📖 Read original article


5. Training Variable Long Sequences with Data-Centric Parallel ​

Author: Geng Zhang, Xuanlei Zhao, Kai Wang, Yang You
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.07524v1 Announce Type: new Abstract: Training deep learning models on variable long sequences poses significant computational challenges. Existing methods force a difficult trade-off between efficiency and ease-of-use. Simple approaches use static configurations that cause workload imbala...

📖 Read original article


6. The Knowing-Saying Gap: When Probes See Errors that Confidence Misses ​

Author: Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithm...

📖 Read original article


7. NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation ​

Author: Yuchen Zhou, Niels Bobet, Maribel Acosta
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.DB

arXiv:2608.07530v1 Announce Type: new Abstract: SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical expertise that most domain experts lack. Translating natural language requirements into SHACL (NL2SHACL) would lower...

📖 Read original article


8. Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems ​

Author: Mojtaba Eslami
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, econ.TH

arXiv:2608.07532v1 Announce Type: new Abstract: Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow full broadcast. Both can be inefficient because token cost, latency, redundancy, and e...

📖 Read original article


9. MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents ​

Author: Gengyang Xu, Dongwei Xiao, Yiteng Peng, Shuai Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.07533v1 Announce Type: new Abstract: An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodied agents primarily relies on two paradigms: (1) manually annotated Visual Question Answering (VQA) pairs and (2)...

📖 Read original article


10. When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains ​

Author: Chen Liang, Fasheng Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, econ.GN, q-fin.EC

arXiv:2608.07538v1 Announce Type: new Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a ...

📖 Read original article


11. TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations ​

Author: Fateme Mazdarani, Carlos Toxtli
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07540v1 Announce Type: new Abstract: AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is recognizing when an unfamiliar formulation denotes a known formal object. We study this challenge through theorem rec...

📖 Read original article


12. An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop ​

Author: Yiwen Zhang, Eloise Zeng, Jaeha Lee, Tony Yue Yu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA, cs.RO

arXiv:2608.07542v1 Announce Type: new Abstract: Autonomous research loops driven by large language models can run machine-learning experiments at scale but tend to drift toward local refinements of whichever metric they optimise rather than testing the hypotheses that motivate the experiments. We ad...

📖 Read original article


13. The Field Knows: Cross-Dimensional Geometry from Navigation to Black Holes ​

Author: Chenghao Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2608.07566v1 Announce Type: new Abstract: We introduce a continuous metric field framework trained by a single causal contrastive loss. The framework encodes a scene into coefficients of a fixed symmetric matrix basis, assembles them into a Lie algebra element, and exponentiates the result to ...

📖 Read original article


14. TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair ​

Author: Prajwal S. Venkateshmurthy
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07617v1 Announce Type: new Abstract: Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatched environments, broken imports, or package conflicts. Existing document-repair evaluations inject fa...

📖 Read original article


15. CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models ​

Author: Hsu-kuang Chiu, Stephen F. Smith
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.RO

arXiv:2608.07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative p...

📖 Read original article


16. Controlled Memory Interference in Continual LLM Agents ​

Author: Ao Ding, Hongzong LI, Shiqin Tang, Li Zhang, Liang Chen, Xuyang Chen, Zi Liang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.IR, cs.LG

arXiv:2608.07622v1 Announce Type: new Abstract: Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience. Yet memory evolution is not simply a process of storing more information: new experiences may reinforce, revise,...

📖 Read original article


17. From Single Chatbots to Governed Agent Ecosystems: An Agentic AI Pattern Catalogue and Orchestration Framework for Mission-Critical Hospital Information Management Systems ​

Author: Manideep Dhar, Ritwik Singh, Sharat Chandra Kumar Manikonda
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA

arXiv:2608.07627v1 Announce Type: new Abstract: Hospitals are racing to embed AI, while coping with the surge in adaptation of the technology in other industries, into the triage management, documentation, scheduling, and revenue-cycle workflows, yet most deployments remain as fragmented pilots that...

📖 Read original article


18. Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD Campaigns ​

Author: Yijie Wang, Zhen-Yu Yin, Zhenheng Tang, Xiaowen Chu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci

arXiv:2608.07637v1 Announce Type: new Abstract: Long-running molecular simulation campaigns require repeated continuation from saved states, provenance-aware progression, adaptive assessment, and occasional interpretation of workflow conditions that cannot be resolved safely by fixed rules. Here, we...

📖 Read original article


19. Contextual Value Alignment via Multilayer Combinatorial Fusion ​

Author: Yuanhong Wu, Djallel Bouneffouf, D. Frank Hsu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA

arXiv:2608.07642v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches such as RLHF, CAI, and their variants have achieved promising results, they often rely on a single-agent framewo...

📖 Read original article


20. Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution ​

Author: Changzhi Liu, Yilun Liu, Sikuan Yan, Volker Tresp, Yunpu Ma
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.07645v1 Announce Type: new Abstract: Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing solutions generally derive self-modification from a single failure trajectory at a time, overlooking...

📖 Read original article


21. An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography ​

Author: Jalil Jalili, Hossein Taghizad, Anuwat Jiravarnsirikul, Christopher Bowd, Akram Belghith, Raheleh Kafieh, Christopher A. Girkin, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning ...

📖 Read original article


22. IntelliAudit: Using Large Language Models to Evaluate Audit Controls ​

Author: Allison Wilson, Sina Moradi Sabet, Diar Shakimov, Panteha Shahrivar, Mohammad Reza Bagheri, Dean Konenkamp, Mohammad A. Tayebi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CR, cs.HC, cs.MA

arXiv:2608.07688v1 Announce Type: new Abstract: IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic security and compliance controls. This judgment is difficult to automate because relevant evidence is distributed across policies, records, spreadsheet...

📖 Read original article


23. Towards Researcher Agents for Knowledge-Graph Question Answering ​

Author: Tommaso Soru, Abdulsobur Oyewale
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.DB

arXiv:2608.07700v1 Announce Type: new Abstract: Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, grounding surface terms in the target ontology, and producing graph patterns that are both syntactica...

📖 Read original article


Author: Sana Tonekaboni, Lena Stempfle, Sasha Ronaghi, Corinna Coupette, I. Glenn Cohen, Emily Alsentzer, Marzyeh Ghassemi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.07705v1 Announce Type: new Abstract: Clinical foundation models trained on large-scale patient data are increasingly used for decision support, screening, and public health. As deployment expands, privacy risk increasingly arises from model-mediated leakage, yet its prevalence and severit...

📖 Read original article


25. QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing ​

Author: Yijing Zuo, Zhe Fu, Zihan Nie, Zhihui Zhu, Haohan Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.07743v1 Announce Type: new Abstract: Identifying a meaningful quantum speedup requires more than matching a classical problem to a familiar quantum primitive: the claim must preserve the task, respect access and output models, expose required promises, and remain within a defensible compl...

📖 Read original article


26. Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service Classes ​

Author: Simone Mainardi, Kaushal Bansal, Prabhat Singh
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.NI

arXiv:2608.07747v1 Announce Type: new Abstract: We study how to share a single conserved capacity budget across many locations and two service classes when demand is uneven, time-varying, and can exceed supply. The shape recurs: an origin's request-rate cap split across its edge locations, a license...

📖 Read original article


27. Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation ​

Author: Sahil Pardasani, Madhusudan Singh
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed OpenAI's o1 contributed to market panic on January 27, 2025, when Nvidia ...

📖 Read original article


28. AndroidReality: How Far Are Mobile Agents from the Real World? ​

Author: Xiaoou Liu, Longchao Da, Hanyang Chen, Yuan Ling, Hua Wei
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-world deployment due to environmental variations and imperfect interface conditions. In this work, we in...

📖 Read original article


29. The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era ​

Author: Majid Memari, George Rudolph
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07779v1 Announce Type: new Abstract: Artificial intelligence is changing the task composition of computing work faster than curricula and training typically adapt. This is a curriculum-framework paper, grounded in a structured narrative review of labor-market and software-engineering evid...

📖 Read original article


30. Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space ​

Author: Yiwei Chen, Bingqi Shang, Sijia Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.07786v1 Announce Type: new Abstract: Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage pipelines, leading to intricate lineage relationships that reflect model origin, ownership, and evolution. Understanding these relationships is important f...

📖 Read original article


31. CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR ​

Author: Veronica Chatrath, Bryan Zhu, George Pu, Jingxuan Fan, Apaar Shanker, Varun Ursekar, Anahita Sharma, Jason Qin, Keqi Han, Soham Dinesh Tiwari, Soham Dan, Vijay Kalmath, Yuan Li, Daniel Yue Zhang, Chenguang Wang, Zainab Doctor, Zhijun Yin, Nigam H. Shah, Yuan Xue
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and ...

📖 Read original article


32. CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift ​

Author: Yiyao Zhang, Diksha Goel, Hussain Ahmad, Shixun Huang, Jun Shen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We study both halves of that requirement with CausalNav, a controller built around a signed, action-conditioned tran...

📖 Read original article


33. When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines ​

Author: Yiyao Zhang, Diksha Goel, Hussain Ahmad, Shixun Huang, Jun Shen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07813v1 Announce Type: new Abstract: An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of that decision depends less on judge accuracy than on the decision rule the judge is embedded in. On frozen candid...

📖 Read original article


34. Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge ​

Author: Shibo Chu, Yuze Liu, Tiehua Zhang, Zhishu Shen, Lianghua He, Haofen Wang, Zhijun Ding
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous structures. However, existing benchmarks provide limited assessment of whether LLMs can faithfully perform multi-hop...

📖 Read original article


35. Back to the Future: A workbook time machine for spread sheet creation benchmarks ​

Author: Mansi Uniyal, Agamdeep Singh, Ananya Singha, Priyanshu Gupta, Mukul Singh, Gust Verbruggen, Vu Le, Sumit Gulwani
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived objects in spreadsheets (formulas, charts, pivot tables, and conditional formatting). Applied to public...

📖 Read original article


36. SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control ​

Author: Rulin Zhou, Qiujie Song, Yujie Ma, An Wang, Wanhao Liu, Guoheng Ma, Yidu Wang, Guankun Wang, Xingrong Diao, Jiankun Wang, Chaowei Zhu, Xianming Liu, Hongliang Ren
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07876v1 Announce Type: new Abstract: Autonomous laparoscopic camera control requires continuous understanding of the surgeon's operative intent in dynamic surgical scenes, where the target operative region is not a stable physical object but a latent and temporally evolving attention stat...

📖 Read original article


37. GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering ​

Author: Zihua Yang, Zhencheng Xie, Junyang Chen, Liang Xie, Yiqun Zhang, Mengke Li, Yang Lu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IT, cs.LG, math.IT

arXiv:2608.07881v1 Announce Type: new Abstract: Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurements and discrete categorical symbols. Traditionally, algorithms rely entirely on dataset-internal statistics to est...

📖 Read original article


38. Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills ​

Author: Agamdeep Singh, Srishti Gautam, Priyanshu Gupta, Nikita Mehrotra, Tanmay Bakshi, Sumit Gulwani
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same ...

📖 Read original article


39. TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis? ​

Author: Yuxuan Zhu, Peng Pu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07899v1 Announce Type: new Abstract: Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure originated. We introduce TelemetrySuffBench, a controlled benchmark that separates failure detection, fa...

📖 Read original article


40. GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning ​

Author: Chen Li, Sijie Cheng, Yuelin Zhang, Junxi Li, Maozhi Huang, Yang Liu, Wenbing Huang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2608.07905v1 Announce Type: new Abstract: Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environmental awareness. We propose GraphThink, a novel framework that integrates a task graph to provide struc...

📖 Read original article


41. When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits ​

Author: Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark in which an unknown fraction alpha of items was seen during training. W...

📖 Read original article


42. TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents ​

Author: Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07917v2 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis. Optical character recognition (OCR) can bridge ...

📖 Read original article


43. ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration ​

Author: Yang Liu, Shiwei Hou, Xiyuan Chen, Yu Wang, Sen Yuan, Qirui Gan, Shao You, Feifan Chen, Wencheng Li, Shuyang Hu, Yongzhou Liu, Emma Xia, Xiaojing Lu, Hao Wang, Fan Xu, Yanfeng Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07925v1 Announce Type: new Abstract: EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong, an execution-grounded LLM coding agent for PyAether and SKILL that combines API retrieval, documen...

📖 Read original article


44. REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment ​

Author: Zhengze Huang, Luyang Yu, Di Hong, Xinzhe Huang, Wanyu Lin, Zhixuan Chu, Zhan Qin, Tianhang Zheng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinct failure sources: reasoning hallucination, where flawed inference steps...

📖 Read original article


45. Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution ​

Author: Lewei Xu, Yihao Ding, Zihan Xu, Daniel Yitian Su, Daochang Liu, Siwen Luo, Yifan Peng, Wei Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07943v1 Announce Type: new Abstract: Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has produced competing, largely untested claims about how these systems sho...

📖 Read original article


46. Directed Neuro-Symbolic Stochastic Execution for Verification of Distributed Parallel AI Programs ​

Author: Gautham Koorma, Vikas Sharma, George Edwards, Mahdi Eslamimehr
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.DC

arXiv:2608.07947v1 Announce Type: new Abstract: Distributed parallel Artificial Intelligence (AI) programs expose reliability gaps that conventional testing cannot close: parallel executions are non-deterministic, and AI workloads bring high-dimensional inputs and non-linear operations that defeat f...

📖 Read original article


47. Guixu: Valuation-Driven Data Discovery for Autonomous AI Agents with On-Chain Attestation ​

Author: Yifan Wu, Yuchen Peng, Jiaqi Chai, Yufei Qian, Xilin Li, Ke Chen, Lidan Shou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.DB, cs.IR, cs.MA

arXiv:2608.07949v1 Announce Type: new Abstract: Autonomous agents increasingly rely on external data to complete downstream tasks such as model training and decision support. However, existing data discovery systems remain largely retrieval-oriented: they surface candidate datasets from heterogeneou...

📖 Read original article


48. KGCache: Amortized Subgraph Retrieval for KG Reasoning with LLMs ​

Author: Uros Stanic, Changcheng Yuan, Sabuj Laskar, Ariful Azad
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07954v1 Announce Type: new Abstract: Large language models can answer knowledge-intensive questions more reliably when they are grounded with knowledge graphs, but systems such as Think-on-Graph and Reasoning-on-Graph repeatedly query the same graph neighborhoods across different question...

📖 Read original article


49. Self-Evolving Neuro-Symbolic Skills for Tool-Augmented Spatial Reasoning ​

Author: Shi-Yu Tian, Zhuo-Xia Wang, Xuan-Yi Zhu, Zhi Zhou, Xinwei Yang, Kun-Yang Yu, Ming Yang, Yang Chen, Yu-Feng Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07955v1 Announce Type: new Abstract: Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial tasks that demand both precise spatial perception and fine-grained geometric computation beyond end-to-end generat...

📖 Read original article


50. SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning ​

Author: Keyang Zhong, Kuo Wang, Peng Liu, Quanlong Zheng, Junlin Xie, Zhijia Liang, Yanhao Zhang, Guanbin Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07959v1 Announce Type: new Abstract: Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments. While Chain-of-Tool-Thoug...

📖 Read original article


51. CyberAGENTS: Structured Autonomy for Agentic Gamified Learning in Cybersecurity ​

Author: Ivan Hornung, Deepthi Marasinghe Arachchige, Tharindu Kumarage, Garima Agrawal, Yuli Deng, Ying-Chih Chen, Huan Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07965v1 Announce Type: new Abstract: Gamification is especially effective in learning domains requiring active problem-solving and iterative skill-building, such as cybersecurity education. Generative AI agents offer a path to delivering such experiences adaptively at scale, but introduce...

📖 Read original article


52. VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge ​

Author: Wenqi Chen, Haofei Yang, Rui Yang, Fangming Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2608.07994v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex product documentation like telecommunications. However, existing RAG approaches largely overlook the holistic integ...

📖 Read original article


53. Thought-Level Beam Search for Reasoning ​

Author: Lijie Yang, Hongyin Luo, Jiawei Zhao, Tri Dao, Ravi Netravali
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08020v2 Announce Type: new Abstract: Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from \emph{how much} compute to spend, to \emph{where} to allocate it. We ...

📖 Read original article


Author: Mark Burgess
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2608.08022v1 Announce Type: new Abstract: Recent incidents involving Artificial Intelligence (AI) agents, which were reported escaping their containment `unintentionally' to gain unauthorized access, pose looming questions about who or what should be held legally responsible for resultant crim...

📖 Read original article


55. The Authority Expectancy Effect in Multi-User Conflict ​

Author: Eunna Lee
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08026v1 Announce Type: new Abstract: We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a model-elicited baseline -- the triage hierarchy and the SA hierarchy. Across four LLMs (Claude, Gemin...

📖 Read original article


56. Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE ​

Author: Ramakrishna P. Kompella, Aadit Mahajan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.08032v1 Announce Type: new Abstract: Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language. We trace this gap mechanistically in sarvam, an Indic-multilingual mixtu...

📖 Read original article


57. SkillSmith: Enhancing Locally Deployed Agents via Automatic Skill Construction and Evolution ​

Author: Xinle Jiang, Remy Xie, Ming Tang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08037v1 Announce Type: new Abstract: LLM-based agent frameworks now act as personal assistants for multi-step tasks. Existing agent frameworks such as OpenClaw commonly follow the Cloud Agent depolyment mode using closed-source cloud LLMs as backbone model, which may expose private user i...

📖 Read original article


58. Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities ​

Author: Xiaohe Li, Yiru Wang, Junhao Fan, Mingyuan Liu, Jie Huang, Kaixin Zhang, Jiahao Li, Chen Qian, Zide Fan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08045v1 Announce Type: new Abstract: Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Simulators therefore provide a scalable foundation for developing and evaluating such coordination. Exi...

📖 Read original article


59. JustLLMGRPO: Radiographic Control for Chest X-Ray Generation ​

Author: Pengxiang Cai, Xiaohan Li, Anglin Liu, Qingyuan Zeng, Zexun Li, Jintai Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08046v1 Announce Type: new Abstract: Text-conditioned chest X-ray generation aims to synthesize realistic radiographs that faithfully depict specified findings. Existing work has primarily improved quality by updating image generators, implicitly treating prompts as fixed after CXR-domain...

📖 Read original article


60. SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents ​

Author: Fengrong Wan, Chengcan Wu, Ningtao Lyu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat RAG diaries and Markdown logs optimize needle retrieval but under-serve currency, provenance, and o...

📖 Read original article


61. H2: A Dual Hybrid Semantic Data Lake Architecture for Medical Data Harmonization with Human-In-the-Loop verified, LLM Driven Metadata Annotation System ​

Author: Ioannis N. Tzortzis, Georgia Kapetadimitri, Agapi Davradou, Nefeli Kousta, Nikolaos Bakalos, Ioannis Rallis, Dimitrios Kalogeras, Nikolaos Doulamis, Anastasios Doulamis
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08056v1 Announce Type: new Abstract: Medical data, by its nature, exhibit a high degree of heterogeneity on multiple levels ranging from (a) different modalities like images, text and time series, (b) diverse tabular schemata introduced by institutions and (c) completely unstructured text...

📖 Read original article


62. CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models ​

Author: Siddarth Singh, Victoria Williams, Simon Rosen, Ebenezer Gelo, Helen Sarah Robertson, Ibrahim Suder, Benjamin Rosman, Geraud Nangue Tasse, Steven James
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict. Current evaluations of large language models (LLMs) remain limited: most test whether mo...

📖 Read original article


63. Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios? ​

Author: Gabriele La Malfa, Nitay Alon, Emanuele La Malfa, Reuth Mirsky, Stefan Sarkadi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.RO

arXiv:2608.08077v1 Announce Type: new Abstract: Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observability. As AI techniques are increasingly applied to safety-critical scenarios, it is crucial to understand whethe...

📖 Read original article


64. PATH: Next-Interval Prediction via Autoregressive Tree Hierarchy on Tabular Data ​

Author: Pengxiang Cai, Wanchen Lian, Chenyang Liu, Xiaohan Li, Qingyuan Zeng, Jinhong Wang, Jintai Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.08078v1 Announce Type: new Abstract: Interval prediction aims to achieve a target coverage level while producing intervals that are as short as possible. Many conformal regression pipelines first predict an uncertainty surrogate and then convert it into an interval through calibration or ...

📖 Read original article


65. Generative Models: Principles, Architectures, and Applications ​

Author: Jun Lu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08101v1 Announce Type: new Abstract: Generative AI has emerged as one of the most transformative forces in modern artificial intelligence, reshaping how we create, imagine, and interact with digital content. From photorealistic images to coherent text, from immersive videos to novel molec...

📖 Read original article


66. Think Deep, Speak Once: Relit, A Recursive Latent Implicit Transformer Framework ​

Author: Abhishek Panwar, Maheep Singh, Saksham Bansal
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08113v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has become the dominant paradigm for eliciting reasoning in Large Language Models (LLMs), yet it creates substantial computational overhead by forcing models to externalize intermediate reasoning steps as discrete token...

📖 Read original article


67. Neurosymbolic Discovery of Algebraic Graph Constructions ​

Author: David Seka, Stefan Szeider
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SC, math.CO

arXiv:2608.08118v1 Announce Type: new Abstract: There are several methods for searching for graphs with prescribed properties, such as SAT solvers and specialized generators. These methods return the result as raw data: an adjacency matrix or a string encoding. The raw data certifies that the graph ...

📖 Read original article


68. Constraining ontology mappings using metaphysical choices ​

Author: Giacomo De Colle, Helena Blackmore, Chris Partridge
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.DB, cs.LO

arXiv:2608.08122v1 Announce Type: new Abstract: In this paper we discuss the foundations behind a novel methodology for the validation of semantic mappings between different data sources based upon different foundation ontologies, where the methodology builds a framework based upon the metaphysical ...

📖 Read original article


69. Improving Constraint Models with LLM Agents ​

Author: Florentina Voboril, Stefan Szeider
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LO, cs.SE

arXiv:2608.08127v1 Announce Type: new Abstract: The runtime of Constraint Programming (CP) solvers is highly sensitive to modeling choices, such as symmetry breaking, implied constraints, global constraints, constraint reformulation, and variable representation. Improving these constraint models has...

📖 Read original article


70. TokenPrint: A Calibrated Token-Space Fingerprint for Language-Model Provenance ​

Author: Yuqi Wu, Shengming Zhao, Jie Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08139v1 Announce Type: new Abstract: Establishing the provenance of a language model---including its base checkpoint and possible overlap in training distributions---is a governance challenge that metadata alone cannot resolve. We introduce a training-free fingerprint based on the top-$k$...

📖 Read original article


71. Long SKILL Compliance as Logical Reasoning: Closure-Grounded Detection with Scaling-Guided On-Policy Distillation ​

Author: Shuaitao Zhao, Feng Ni, Lichao Ma, Jiaye Lin, Fei Han, Yang Wei, Lu Pan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08146v1 Announce Type: new Abstract: The increasing complexity of enterprise business scenarios has promoted the widespread adoption of long SKILL documents in agent systems, posing new challenges for compliance detection: large models incur substantial inference costs, while small models...

📖 Read original article


72. A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning ​

Author: Fouad Bahrpeyma
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.08158v1 Announce Type: new Abstract: Sparse, delayed, and weakly informative rewards remain central obstacles to efficient reinforcement learning. Reward shaping addresses these limitations by supplementing the task reward with an auxiliary signal that can accelerate learning while, in th...

📖 Read original article


73. When Is a Steerable Concept Representation Real? Measurement Confounds in a Cross-Family Audit of Neuroscience Parallels in LLMs ​

Author: Yuqi Wu, Shengming Zhao, Jie Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08159v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly reported to exhibit human-like neural and cognitive signatures, including concept cells, mental number lines, and cognitive maps. These claims often rely on linear probing and activation steering applied to...

📖 Read original article


74. Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) Brachytherapy ​

Author: Ronghua Xu, Kepha Barasa, Manoj Kumal, Xinyun Liu, Weihua Zhou, Xin Qian
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08163v1 Announce Type: new Abstract: The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized pedagogical frameworks in medical education. This paper presents a novel agentic AI-driven immersive sim...

📖 Read original article


75. Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation ​

Author: Yongkang Yang, Zhezheng Hao, Hong Zhang, Yi Liu, Xiankun Lin, Wence Ji, Fanjunduo Wei, Jiarui Yu, Qiang Lin, Xiaoyun Liang, Hande Dong
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.08176v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-distillation. Two recent research lines promote vanilla OPSD by choosing which tokens to learn from and b...

📖 Read original article


76. Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges ​

Author: Muhammad Ayub Sabir, Shaohong Zheng, Zhiyu Qu, Fatima Ashraf, Junbiao Pang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08184v1 Announce Type: new Abstract: Large multimodal agents (LMAs) are increasingly proposed for intelligent transportation systems (ITS), but existing studies often conflate multimodality, agency, empirical performance, and deployment readiness. This review provides an auditable evidenc...

📖 Read original article


77. Quantization Degradation in Large Language Models: A Signal-Noise Perspective ​

Author: Chenxi Zhou, Pengfei Cao, Jinyu Ye, Bohan Yu, Haida Yu, Jiang Li, Jun Zhao, Kang Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.08188v1 Announce Type: new Abstract: Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across bit-widths, quantiz...

📖 Read original article


78. Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation Budgets ​

Author: Ximeng Liu, Qianlong Wang, Yingming Mao, Annan Li, Yatao Li, Shizhen Zhao, Jianmin Wu, Dawei Yin, Dou Shen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08189v1 Announce Type: new Abstract: LLM-driven program discovery relies on rapid evaluator feedback, but many scientific and engineering tasks require high-fidelity simulations, hardware execution, or physical experiments, making each evaluation expensive. Cheap surrogate evaluators can ...

📖 Read original article


79. A Minimal $\kappa$--$\tau$ Logic for Risk-Sensitive Abduction ​

Author: Remo Pareschi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08192v1 Announce Type: new Abstract: Standard approaches to abductive reasoning can retain multiple candidate explanations, but they do not generally combine explicit compositional cross-hypothesis interaction with an internal, rival-sensitive commitment judgment. This paper argues that i...

📖 Read original article


80. Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models ​

Author: Wenwen He, Wenke Huang, Wei Yang Bryan Lim, Dacheng Tao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08199v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in group decision-making with other LLMs and humans. Yet it remains unclear whether their influence is driven by persuasion-oriented expression or compliance-oriented accommodation. We introduce De...

📖 Read original article


81. Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue ​

Author: Kaiming Liu, Fuwen Luo, Ziyue Wang, Jinrui Ju, Yuxuan Liu, Xuanyu Lei, Yunghwei Lai, Peng Li, Yang Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08210v1 Announce Type: new Abstract: Collaborative dialogue can end with apparent agreement while participants still differ on goals, assumptions, or execution plans, creating an \textbf{illusion of alignment (IoA)}. A real-user study across 18 meetings confirms that IoA arises routinely ...

📖 Read original article


82. Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment ​

Author: Peiyang Liu, Xi Wang, Ziqiang Cui, Di Liang, Wei Ye
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.08212v1 Announce Type: new Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Existing prompts, however, conflate harmful-text exposure with an invitation to continue assistant behavior. We hold...

📖 Read original article


83. Metanormative Theory for RL-Based Moral Agents ​

Author: Aleks Knoks, Marija Slavkovik
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08220v1 Announce Type: new Abstract: The overlapping disciplines of machine ethics and value alignment are concerned with designing artificial agents that are aligned with human values and that act in ethically acceptable ways. A recent trend in these disciplines is the use of reinforceme...

📖 Read original article


84. LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems ​

Author: Heng Zhou, Lian Zhang, Yutao Fan, Tiancheng He, Siki Chen, Hejia Geng, Philip Torr, Zhenfei Yin
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.08236v1 Announce Type: new Abstract: Multi-agent LLM systems often fail not for lack of candidate answers, but because they have no persistent mechanism for deciding which incompatible claim should currently be trusted. Majority vote, debate, and judge-based selection choose an output wit...

📖 Read original article


85. A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences ​

Author: Jobst Heitzig, Ram Potham
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, cs.MA, econ.TH

arXiv:2608.08240v1 Announce Type: new Abstract: This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially ax...

📖 Read original article


86. FemWear: A Specialized Wearable Foundation Model for Women's Health ​

Author: Yifan Wang, Chenzhong Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08244v1 Announce Type: new Abstract: General wearable foundation models are pretrained across broad sensor streams and populations, but are not designed around women's-health tasks. We introduce FemWear, a specialized wearable foundation model that parameter-efficiently repurposes a pretr...

📖 Read original article


87. SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents ​

Author: Varun Pratap Bhardwaj, Garima Singh, Arun Pratap Bhardwaj
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.IR

arXiv:2608.08253v1 Announce Type: new Abstract: AI agents are becoming shared infrastructure, yet durable memory is commonly assembled from separate retrieval, governance, and operational components. We present SuperLocalMemory 4.0, a governed, local-first memory operating system for AI agents. The ...

📖 Read original article


88. Your Prompt Is Not the Only Prompt: How Much Do LLMs Weight Structured-Output Schema Descriptions? ​

Author: Sin-Ying Lin
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08254v1 Announce Type: new Abstract: Structured output, where an LLM populates a predefined JSON schema, has become a default mechanism for data labeling and information extraction, but it also introduces a second instruction channel through schema descriptions. We tested whether classifi...

📖 Read original article


89. OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents ​

Author: Zhengyang Shan, Xu Qian, Jiayun Xin, Kun Li, Yue Zhang, Minghui Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08264v1 Announce Type: new Abstract: Large language model agents are becoming operational interfaces to files, memories, registries, and external tools. This deployment shift creates a new skill revocation problem: after a skill is removed from an explicit registry, an agent may still rec...

📖 Read original article


90. Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios ​

Author: Julius Wirbel, P. Nicholas Hansen, Line K. H. Clemmensen, Roberto Galeazzi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.RO

arXiv:2608.08281v1 Announce Type: new Abstract: Recently, Large Language Models (LLMs) have shown considerable capability for situational understanding, reasoning, and decision making in different domains, most notable in the automotive sector. Therefore, we explore current state-of-the-art LLMs as ...

📖 Read original article


91. Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders ​

Author: Chan Aristella Lu, Arya Fayyazi, Junhao Zhang, Saeid Shokoufa, Yue Xing, Zhen Xiang, Kyu Hyung Lee, Mehdi Kamal, Massoud Pedram
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08284v1 Announce Type: new Abstract: Fairness audits for LLM-based recommenders have largely focused on observable outputs, implicitly assuming that stable recommendations reflect stable internal processing. We challenge this assumption with FairGap, the first benchmark to jointly evaluat...

📖 Read original article


92. Mitigating Over-Personalization in LLMs via Structured Memory ​

Author: Hakeem Hannoon, Andrew Zhao, Mihir Narayan, Sharvin Goyal, Ivaxi Sheth
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.08300v1 Announce Type: new Abstract: Conversational assistants increasingly rely on persistent long-term memory to personalize responses across sessions. However, when stored user information is reintroduced into the model context, it can also influence responses in inappropriate or unrel...

📖 Read original article


93. Query-Only Backdoor Attacks on Self-Evolving Skills via Trajectory Poisoning ​

Author: Yuyang Luo, Haoran Wang, Kai Shu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.08303v1 Announce Type: new Abstract: Agentic skills improve large language model (LLM) agents by encoding reusable procedures for complex tasks. However, manually authored skills often adapt poorly to long-horizon tasks and changing environments. To address the limitation, self-evolving s...

📖 Read original article


94. StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning ​

Author: Yifan Li, Ruxin Sun, Tongzhou Zhao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08326v2 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, most existing methods evaluate an entire response using a binary reward based only on final-answer correctness, ther...

📖 Read original article


95. LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving ​

Author: Shuowei Jin, Xueshen Liu, Jiaxin Shan, Le Xu, Tieying Zhang, Liguang Xie, Z. Morley Mao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.DC

arXiv:2608.08382v1 Announce Type: new Abstract: As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control. Enabling fractional sharing of the inference engine requires a real-time, per-request attribution primitive that is ...

📖 Read original article


96. Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents ​

Author: Harshitha Kolukuluru, Reshma Ashok, Kirat Arora, Evan William Ciccarelli, Nischal Ashok Kumar, Lunyiu Nie, Franck Dernoncourt, Samyadeep Basu, Ryan A. Rossi, Nedim Lipka
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.IR, cs.MA

arXiv:2608.08389v1 Announce Type: new Abstract: Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the marginal value of additional evidence often declines. This leads to unnecessary token cost, higher latency,...

📖 Read original article


97. CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception ​

Author: Zejun Xu, Taiyi Chen, Jin Li, Yongtong Gu, Qi Cheng, Aixuan Lv, Shuai Zhu, Pengfei Zhu, Kaichen Yang, Boyu Sun, Yixian Yang, Mulong Xie, Xin Liu, Dagang Li, Xiaoteng Ma, Hongru Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven by benchmarks that evaluate end-to-end task success, these evaluations largely overlook two fundamen...

📖 Read original article


98. Estimating Uncertainty in Galaxy Morphology Classification ​

Author: Kai Cheng, Ruoqi Wang, Qiong Luo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, astro-ph.IM

arXiv:2608.08398v1 Announce Type: new Abstract: Astronomers classify galaxy morphology to investigate cosmic evolution. While deep foundation models are increasingly utilized in Galaxy Morphology Classification (GMC), little work has been done on evaluating the uncertainty of GMC results. Uncertaint...

📖 Read original article


99. Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective ​

Author: Xiaoyan Zhao, Yujie Cai, Yang Zhang, Grace Hui Yang, Tat-Seng Chua
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08445v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of large language models (LLMs)--a mechanism to ground their outputs in external knowledge. This view, however, is incomplete when considered within a...

📖 Read original article


100. TRACE-Memory: Public-Conditioned Retrieval and Utility-Aware Evidence Admission for Personalized Generation ​

Author: Jing Wang, Zhu Wang, Yifan Guo, Yulong Yang, Yunji Liang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08446v1 Announce Type: new Abstract: Personalized generation systems retrieve user history by request--memory relevance and inject it into the model context. Yet relevant history may concern the wrong preference aspect, duplicate public information, or provide insufficient support. We arg...

📖 Read original article


101. What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files ​

Author: Chi Zhang, Yimin Liu, Xinze Chen, Ping Ji
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CR

arXiv:2608.08453v1 Announce Type: new Abstract: Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting resources, enabling Large Language Model (LLM) agents to reuse procedures beyond a single conversation. Yet many public skills appear to originate fro...

📖 Read original article


102. Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses ​

Author: Tailin Zhou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the \emph{harness}---is typically treated as a fixed artifact after deployment. This work studies an alternat...

📖 Read original article


103. LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs ​

Author: Minhan Cho, Soyoung Park, Kihyeon Jeong, Byeongkyu Jeon, Daejin Choi, Jinyoung Han
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR

arXiv:2608.08467v1 Announce Type: new Abstract: The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequently used reference data, such as identifier lookup tables, directly in the server instructions: the sy...

📖 Read original article


104. Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation ​

Author: Kaichen Zhang, Wei Huang, Keming Wu, Bo Li, Xiaojuan Qi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08469v1 Announce Type: new Abstract: Existing streaming multimodal models process observations incrementally but still follow a turn-based prefill-then-decode pattern, making them non-duplex: new observations cannot naturally enter an active generation stream. Proactive alternatives use m...

📖 Read original article


105. Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production ​

Author: Cong Ming, Jingyi Chen, Bin Liu, Qi Chu, Tao Gong, Nenghai Yu, Yingfei Xiang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08471v1 Announce Type: new Abstract: Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new jailbreak techniques and previously un-addressed harmful categories emerge within days, leaving the defense perpetually a step behind. We present SES...

📖 Read original article


106. HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails ​

Author: Tak Ho Alex Li, Kaijie Liu, Lik-Hang Lee, Kin Chung Ho, Ping Shum, Michael K. Ng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2608.08485v1 Announce Type: new Abstract: Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inference costs. We challenge the prevailing paradigm by asking: can safety be achieved through pure ge...

📖 Read original article


107. TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models ​

Author: Yidong Wang, Yan Zhan, Ziteng Feng, Zhenyu Cui, Ziyi Zhou, Renzhao Liang, Jiaxuan Zhu, Zilei Yang, Yiran Zhao, Zhongkuan Mao, Bo Jia, Hanchu Ni, Chenggang Xie, Biao Liu, Yi Zhang, Yong Dai, Xiaozhu Ju, Wei Ye, Shikun Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08491v1 Announce Type: new Abstract: Reward models are a bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision feedback beyond handcrafted rewards or task-specific annotations. Existing open-source VLM reward judges like RoboRewar...

📖 Read original article


108. MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models ​

Author: Rahma Simin Ali, Jawad Hossain
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.08503v1 Announce Type: new Abstract: Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Thought (CoT) supervision provides benefits beyond ordinary supervised fine-tuning. We construct \textsc{MathShikkha...

📖 Read original article


109. Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs ​

Author: Mohanad Odema, Gabrielle De Micheli, Dayin Gou, Nilesh Malpeddi, Prathamesh Vaste, Jacob Song
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.PF

arXiv:2608.08506v1 Announce Type: new Abstract: Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-level accuracy. However, existing SOTA frameworks share two key limitation...

📖 Read original article


110. Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding ​

Author: Mahbub E Sobhani, Md. Faiyaz Abdullah Sayeedi, Fahmid Hasan Chowdhury, Md Adnan Arefeen, Farig Sadeque, Md. Faizul Bari, Swakkhar Shatabda
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08512v1 Announce Type: new Abstract: Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a question has different correct answers at different dates. In contrast to encyclopedic knowledge, where an old fact i...

📖 Read original article


111. Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing ​

Author: Minhan Cho, Jimin Kweon
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08514v1 Announce Type: new Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B models). The first, RPC, a...

📖 Read original article


112. Discovering Diverse Planning Policies for Multimodal Embodied Agents with Quality-Diversity Optimization ​

Author: Pengfei Xu, Yong Liu, Xiaoya Nan, Qiang Yang, Peilan Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08523v1 Announce Type: new Abstract: Multimodal embodied agents are increasingly required to solve long-horizon tasks by integrating visual observations, textual goals, and interaction history into closed-loop decision making. However, state-of-the-art large-model-based planners often rel...

📖 Read original article


113. Deep probabilistic logic programming for diagnostic reasoning from incomplete information: A case study in stroke detection ​

Author: Felix Weitk"amper, Monchito Avila, Elizabeth Nanjala, Siska, Grace Zawadi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08561v1 Announce Type: new Abstract: In medical applications, raw data is frequently associated with significant privacy concerns, lending particular importance to the encoding of summary statistics from the literature. On the other hand, deep learning has become an invaluable tool for as...

📖 Read original article


114. VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference ​

Author: Wenxu Jia, Dongjie Fu, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.SD

arXiv:2608.08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks. Despite this progress, their long-context inference remains severely bottlenecked by prohibitive KV cache memory demands...

📖 Read original article


115. FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents ​

Author: Dongyi Lv, Fushun E, Aichen Cai, Liang Huang, Ya Zhang, Qiuyu Ding, Canhui Wu, Zhi Wang, Yuesong Zhang, Jiaqi Wang, Nan Duan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08570v1 Announce Type: new Abstract: Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering tasks, retaining those that pass the tests, and fine-tuning on the successful rollouts. However, even strong code age...

📖 Read original article


116. SDDBMs: Soft Denoising Diffusion Bridge Models ​

Author: Shiyi Qi, Kun He, Mingmou Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08594v2 Announce Type: new Abstract: Diffusion bridge models leverage Doob's (h)-transform to construct stochastic transports between arbitrary endpoint distributions, and have shown strong potential in image-to-image translation and restoration. However, most existing bridge models rel...

📖 Read original article


117. Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents ​

Author: Gabriele La Malfa, Lakmal Meegahapola, Edyta Bogucka, Jie M. Zhang, Michael Luck, Elizabeth Black, Daniele Quercia
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.MA

arXiv:2608.08601v1 Announce Type: new Abstract: To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies focus on broad risks and do not capture job-specific risks introduced by agents. To address this gap, we make three...

📖 Read original article


118. ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration ​

Author: Guo Chen, Ziwen Li, Reed Li, Yu Lu, Haibo Shi, Bingbing Xu, Junjie Huang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08605v2 Announce Type: new Abstract: Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common basis for evaluation across methods. Outcome-only benchmarks discard collaborations, whereas LLM-as-Ju...

📖 Read original article


119. Walking through Discussions: A Mobile Visual Analytics System for In-Situ Group Discussion Analysis ​

Author: Yiping Sun, Ziyao Kang, Wei Zeng, Minli Wu, Jiazhi Xia
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.08617v1 Announce Type: new Abstract: Group discussion-based teaching is widely used to foster collaborative learning, yet teachers in physical classrooms often struggle to simultaneously monitor multiple groups and quickly diagnose a target group before intervening. Existing visual analyt...

📖 Read original article


120. Business Arena: Benchmarking LLM Agents in a Realistic Marketplace ​

Author: Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading l...

📖 Read original article


121. MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning ​

Author: Haotian Wang, Lian Yan, Xingzhi Yao, Fanshu Meng, Ye He, Jingchi Jiang, Yi Guan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08623v1 Announce Type: new Abstract: In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated using a tolerance-based reward. However, this strategy suffers from challenges such as difficulty in th...

📖 Read original article


122. UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models ​

Author: Lei Xin, Bin Gu, Peize Li, Zitong Wang, Jianbo Zhao, Changjiang Jiang, Yanyue Xie, Chao Huang, Xuyang Zhao, Zunhai Su, Fanhu Zeng, Zhenglun Kong
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08627v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes over its full expert bank. We study a deployment problem: convert that checkpoint to a smaller stand...

📖 Read original article


123. A QUBO-Inspired Computational Framework for Airport Landside Bottleneck Diagnosis and Dynamic Dispatch Optimization ​

Author: Wuming Lei, Xiaobin Li, Mingyan Sun, Jianing Long, Yulin Tong, Yanbin Gao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08632v1 Announce Type: new Abstract: Airport landside traffic centers connect terminal arrivals with taxis, ride-hailing vehicles, private cars, buses, metro services, parking facilities, and terminal-area roadways. Peak arrivals can create coupled congestion across passenger queues, vehi...

📖 Read original article


124. Can Open-Weight Models Compete on Financial Text Comprehension? ​

Author: Jan Sp"orer
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, q-fin.GN

arXiv:2608.08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability on real-world financial tasks remains largely untested. We updated the Financial Touchstone benchma...

📖 Read original article


125. Smart Compaction: Predicting Compaction Utility from Lakehouse Table Metadata ​

Author: Jannic Cutura, Subash Prakash
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08639v1 Announce Type: new Abstract: Open lakehouse table formats accumulate small data files over time, which degrades query performance. Deciding when compaction is worthwhile remains threshold-driven, but which metadata features actually determine compaction utility is not well underst...

📖 Read original article


126. SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests ​

Author: Donghong Jiang, Endian Lin, Luoping Cui, Hanqing Liu, Mingjie Liu, Fan Yang, Hong Wang, Zhao Yang, Chuang Zhu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are...

📖 Read original article


127. The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task ​

Author: Marc Alier Forment, Mar'ia Jos'e Casa~n Guerrero, Francisco Jos'e Garc'ia-Pe~nalvo, Juanan Pereira
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it reaches its tools. We set out to measure the cost of tool use over the Model Context Protocol (MCP) against tool use...

📖 Read original article


128. Branch2Skill: Efficient Skill Evolution Through Reasoning Trees ​

Author: Yanwei Ren, Haotian Zhang, Likang Xiao, Jiaxing Huang, Jiayan Qiu, Baosheng Yu, Quan Chen, Liu Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08677v1 Announce Type: new Abstract: Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing incomplete or misleading behaviors. However, existing methods mainly rely on single trajectories, where early re...

📖 Read original article


129. A Structural Dynamics Graph World Model: Unified Modeling, Constrained Rollout, and Interpretable Calibration ​

Author: Wei Wang, Yaosen Chen, Han Yang, Yuegen Liu, Mingli Luo, Xinxin Jiao, Xuming Wen, Ming Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08689v1 Announce Type: new Abstract: The state evolution of a complex system arises jointly from object laws, relational propagation, domain conservation, and unmodeled error. Forcing all sources into one black box makes mechanism attribution and constraint preservation unauditable; forci...

📖 Read original article


130. EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility ​

Author: Xudong Wu, Zeqing Wu, Jiarui Zhang, Xuhao Fan, Ziang Ding, Yuming Zhuang, Mingqi Yuan, Yilun Du, Hongjie Jia, Yunfei Mu, Jiayu Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08691v1 Announce Type: new Abstract: Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity only when residents authorize a plan and the promised response is delivered. Existing benchmarks eva...

📖 Read original article


131. PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling ​

Author: Dongjie Xu, Julius, Hanchi Dong, Minghua Tang, Yuxuan Sun, Ziwei Nie, Zicheng Liu, Dujun Qing, Jiajie Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents. Current benchmarks face three structural limitations: data distributions that follow a power law leave rare scenarios underrepresented; ...

📖 Read original article


132. AI Evaluation Should Measure Verification Cost, Not Correctness Alone ​

Author: Viviana Crescitelli, Generoso Immediato, Fabio Persia, Stefania Costantini
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.08709v1 Announce Type: new Abstract: The reliability of AI generative models is typically measured by output correctness, yet in practice it depends on the effort required to verify those outputs. We argue that current evaluation metrics overlook a critical failure mode: Verification-Cost...

📖 Read original article


133. FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models ​

Author: Kaili Zheng, Kaiwen Wang, Xun Zhu, Qingyuan Yang, Chenyi Guo, Ji Wu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08736v1 Announce Type: new Abstract: Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) in this setting remain underexplored. Existing benchmarks rely on action-specific annotation schemes ...

📖 Read original article


134. Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models ​

Author: Yifan Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response burden. We formulate conversational administration as an ordinal label-recovery problem: the system actively elicits...

📖 Read original article


135. SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification ​

Author: Wenyao Cui, Huaping Zhang, Yongyi Huang, Qiuchi Li, Jian Xu, Cheng-Lin Liu, Chunxiao Gao, Juan Wang, Baohua Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08786v1 Announce Type: new Abstract: Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when final answers are correct. Most existing ``verification'' signals are not diagnostic: answer matching observes only ...

📖 Read original article


136. Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs ​

Author: Kyeongyoon Lee, Hongyeob Kim, Youngeun Kim, Sungeun Hong
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.MM, cs.SD

arXiv:2608.08794v1 Announce Type: new Abstract: Omni-modal LLMs jointly process audio, video, and text, but long multimodal sequences incur substantial prefill and KV-cache costs. Existing omni-modal compression methods primarily focus on pre-LLM token reduction, leaving modality-specific compressio...

📖 Read original article


137. Improving Generalization Robustness of Multimodal RLVR ​

Author: Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng, Chenrui Zhou, Lama Moukheiber, Yixing Ma, Bin Xu, Jiajun Song, Zhenglin Wan, Wangbo Zhao, Jiasheng Tang, Bohan Zhuang, Fan Wang, Yang You
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08802v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment...

📖 Read original article


138. Three Generations of Healthcare IT: From the Digital Record to the Computable Care Process ​

Author: Alexander Apartsin, Yehudit Aperstein
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08806v1 Announce Type: new Abstract: Objective. Healthcare IT is usually organized by the technologies it adopts. We instead organize it by the unit of information a system makes computable, and describe a computational layer whose object is patient-specific clinical intent. Approach. We ...

📖 Read original article


139. Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models ​

Author: Abdalla Doleh, Toni Somers, Ratna Babu Chinnam
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and biased. We developed an automated pipeline that uses LLms to generate structured decision scenarios a...

📖 Read original article


Author: Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval ...

📖 Read original article


141. Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents ​

Author: Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen, Bo-Han Feng, Yun-Man Hsu, Hsiang Hsieh, Yu-Jung Lin, Yue-Ling Wu, Jia-Kai Dong, An-Yu Cheng, Yu-Han Huang, Lok-Lam Ieong, Kuan-Yu Chen, Ming-Douo Tchouang, Shao-Hua Sun, Che Lin, Jian-Jiun Ding, Hung-yi Lee
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.HC, cs.MA

arXiv:2608.08852v1 Announce Type: new Abstract: AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fit a specified learner, which education calls Pedagogical Content Knowledge (PCK), has not been bench...

📖 Read original article


142. Theory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration ​

Author: David M. Markowitz, Timothy R. Levine
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.08881v1 Announce Type: new Abstract: The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments were made relative to baseline models. Across 700 statements drawn from five published deception datas...

📖 Read original article


143. AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups ​

Author: Jack Stark, Srinath Saikrishnan, Vikram Seenivasan, Bernie Boscoe, Andrew Lizarraga, Tuan Do
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08883v1 Announce Type: new Abstract: Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. However, using proprietary commercial AI systems raises concerns about transparency, reproducibility ...

📖 Read original article


144. Full-bandwidth transformer ​

Author: Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each token broad horizontal access to the past, but the vertical feedback channel between decoding steps...

📖 Read original article


145. LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing ​

Author: Juncheng Dong, Ding Tong, Ishan Gupta, Yuyan Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08889v1 Announce Type: new Abstract: Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As Large Language Models (LLMs) are deployed as autonomous verifiers of safety and quality guidelines, they fa...

📖 Read original article


146. From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings ​

Author: Bernes Lorier Atabonfack, Zion Kongbi Nfo, Ahmed Tahiru Issah, Tolulope Olusuyi, Clemence Ingabire, Mohammed Hardi Abdul Baaki, Mawuli Deku, Abdulrazaq Zubair, Alyasaa Anas, Raymond Confidence, Maruf Adewole, Udunna C. Anazodo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08896v1 Announce Type: new Abstract: Imaging device downtime is a major barrier to healthcare delivery in low- and middle-income countries (LMICs), often driven by limited access to specialized biomedical engineering support. We present a multi-modality medical equipment maintenance quest...

📖 Read original article


147. Decoding Phenotypes: A Framework for Fusing Genomic Language Models and Neuroimaging ​

Author: Tianli Tao, Ziyang Wang, Emma Robinson, Rachel Sparks, Le Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.08926v1 Announce Type: new Abstract: Neuroimaging and genetic testing are two important clinical references for nervous system diseases, offering complementary diagnostic information. However, integrating genomic and neuroimaging data for precise disease diagnosis is challenging due to cr...

📖 Read original article


148. Integrated Multimodal AI System for Retrieval-Augmented Reasoning, Object Sensing, and Damage Analysis ​

Author: Kalelo Dukuray, Israel Pina, Evan Perez, Erika Ardiles-Cruz, Jie Wei
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08935v1 Announce Type: new Abstract: This work presents a unified multimodal AI system for damage assessment that integrates retrieval-augmented generation (RAG) models, thermal spectrum perception, vision foundation model pipelines, and exploratory wireless signal sensing. A RAG componen...

📖 Read original article


149. Not an A11y: How Android Accessibility Exposes Mobile AI Agents to Indirect Prompt Injection ​

Author: Rahul Deivasigamani, Sayeda Faatin Alvi, Derqui Andrea, Kaushal Punjabi, Stjepan Picek
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08939v1 Announce Type: new Abstract: The rise of autonomous AI agents represents a major paradigm shift in how users interact with mobile devices. Frameworks such as MobileRun and Mobile-Use can autonomously navigate Android applications and execute complex multi-step tasks. To interpret ...

📖 Read original article


150. Depth-Aware Implicit Neural Representation Priors for 3D Gravity Inversion ​

Author: Le'on Suarez-Rodriguez, Paul Goyes-Pe~nafiel, Javier Torres-Quintero, Henry Arguello
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08959v1 Announce Type: new Abstract: Gravimetry images subsurface density contrasts associated with geological structures, geothermal systems, and intrusive bodies. Recovering a three-dimensional density model from gravity observations is highly ill-posed because of its non-uniqueness, li...

📖 Read original article


151. Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression ​

Author: Cheng Fan, Junyi Zhou, Tingzhang Luo, RongJian Xu, Qiyanhui Lu, Mingjian Zhu, Hanting Chen, Jianyuan Guo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.08960v1 Announce Type: new Abstract: Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates a marked capa...

📖 Read original article


152. CoRe-UIE: Rethinking Coexisting and Region-wise Degradation for Underwater Image Enhancement ​

Author: Weifeng Kong, Chenghao Xu, Lin Chen, Ziheng Cao, Guanying Huo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.08965v1 Announce Type: new Abstract: Underwater images often suffer from diverse and coexisting degradations, including color distortion, scattering haze, texture attenuation, and uneven illumination. These degradations vary across regions and may coexist locally, making conventional unif...

📖 Read original article


153. Context Is Not Authority: Structured Runtime Governance for Financial Market Agents ​

Author: Rui Tang, Qiangqiang Liu, Yichi Zhang, Youwei Wang, Xi Chen, Chen Dong
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, stat.ML

arXiv:2608.09025v1 Announce Type: new Abstract: Financial agents can turn correct context into an unauthorized effect: a customer-facing commitment, trade, or deployed policy. We present SAGE-Fin, a finance-specific authority-handoff contract that makes the proposed effect, not merely its text, the ...

📖 Read original article


154. PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs ​

Author: Ponkrit Kaewsawee, Chaklam Silpasuwanchai, Chutiporn Anutariya
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.DB, cs.LO

arXiv:2608.09028v1 Announce Type: new Abstract: Institutional policies stay in natural language while the systems that check compliance demand machine-readable constraints. Bridging that gap is still done by hand. PolicyKG closes the loop. It is an LLM pipeline that reads a policy PDF, classifies ea...

📖 Read original article


155. DualCert: A Solver for the Traveling Salesman Problem with Constraint-Coupled Learning ​

Author: Yancheng Song, Yongzhi Qi, Wei Qi, Zuo-Jun Max Shen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09042v1 Announce Type: new Abstract: Large traveling salesman problem (TSP) instances require a solver to allocate limited computation while preserving the validity of its outputs. Existing neural--operations-research (OR) hybrids predict guidance without requiring learned transitions to ...

📖 Read original article


156. A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition ​

Author: Stefanos Gkikas, Yang Guo, Guangliang Li, Raul Fernandez Rojas, Giorgos Giannakakis, Randy Gomez
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SD

arXiv:2608.09088v1 Announce Type: new Abstract: Mixed emotions represent a clinically relevant but still underexplored target for automatic emotion recognition. EEG provides millisecond-level access to neural activity, yet most EEG pipelines analyze the signal through a single temporal window, there...

📖 Read original article


157. Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways ​

Author: Shuyi Miao, Wangjie Qiu, Pengyang Shao, Canran Xiao, Fei Shen, Zhiming Zheng, Tat-Seng Chua
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09095v1 Announce Type: new Abstract: Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mechanistic interpretability studies on multilingual safety are largely conf...

📖 Read original article


158. Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models ​

Author: Xuanchen Li, Haitao Li, Yujia Zhou, Qingyi Pan, Heng Wang, Yiqun Liu, Min Zhang, Qingyao Ai
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09109v1 Announce Type: new Abstract: User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with different scopes of generalization. We introduce SLIFT, a selective self-learning framework built on a task-relat...

📖 Read original article


159. RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement ​

Author: Ziheng Jia, Jiaying Qian, Zicheng Zhang, Xiaorong Zhu, Lancheng Gao, Xiongkuo Min
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09111v1 Announce Type: new Abstract: AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation models~(AIVGMs) have become increasingly difficult to discern using conventio...

📖 Read original article


160. Motif 3: Technical Report ​

Author: Junghwan Lim, Joon Son Chung, Sungmin Lee, Wai Ting Cheung, Gihun Cho, Minsu Ha, Sangho Kang, Beomgyu Kim, Dongseok Kim, Jangwoong Kim, Taehyun Kim, Taewhan Kim, Jeesoo Lee, Jeongdoo Lee, Junhyeok Lee, Dongpin Oh, Hyeyeon Cho, Dahye Choi, Jaeheui Her, Hanbin Jung, Changjin Kang, Minjae Kim, Youngrok Kim, Hyukjin Kweon, Hongjoo Lee, Yeongjae Park, Bokki Ryu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09119v1 Announce Type: new Abstract: We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity ...

📖 Read original article


161. MELLON - Multimodal Enhanced LLM for Online Navigation ​

Author: Ruiyu Li, Haoyang Cai, Zhitong Guo, Tong Hu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09121v1 Announce Type: new Abstract: Web navigation agents are capable of addressing various types of tasks on different websites. Current baselines on web navigation are either unimodal or lack strong reasoning abilities given multimodal inputs. Focusing on the WebShop benchmark, a real-...

📖 Read original article


162. RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning ​

Author: Jinkun Hou, Zhuo Liu, Huimin Ren, Hongsheng Xin, Pan Zhou, Kun Zhan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09123v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a single correct generation trajectory. Existing rubric-based reinforcement learning (RL) methods compr...

📖 Read original article


163. ChronoState: Hidden Elapsed-Time Conditioning for Temporal-State Action Selection in Frozen-Backbone Language Models ​

Author: Sam Siavoshian, Omar Ramadan, Amir K. Saeed, Benjamin A. Johnson, Amin Mohamed El-Amin Diab, Benjamin M. Rodriguez
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09124v1 Announce Type: new Abstract: Temporal decisions in language-model systems often depend on both symbolic task state and elapsed wall-clock time, such as cache expiration, job completion, quota resets, deadlines, or stale sessions. We study whether elapsed time can be supplied as a ...

📖 Read original article


164. TRACE: TRajectory Attribution for Automated Context Engineering ​

Author: Yikai Zhao, Pradeep Kumar Misra, Saurabh Pandey
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.09153v1 Announce Type: new Abstract: Production AI agents fail when their context sources -- system prompts, knowledge bases, tool descriptions, and procedural skills -- contain errors or gaps. Current maintenance relies on manual log review and ad-hoc debugging, creating a scalability bo...

📖 Read original article


165. CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment ​

Author: Bingcan Guo, Eryue Xu, Jijie Zhou, Zhiping Zhang, Tianshi Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09164v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms. However, a gap remains in eliciting such nuanced preferences to evaluate alignment in realistic set...

📖 Read original article


166. From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents ​

Author: Liang He, Jingbo Wen, Hongyu Gu, Hao Li, Haoyu Wang, Yixiong Chen, Kangning Cui, Xilu Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09168v1 Announce Type: new Abstract: Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval due to the increasing skill libraries, retrieving a plausible skill bundle...

📖 Read original article


167. Agentic Router: An Execution-Grounded Continual Learning Approach With Memory ​

Author: Yuxuan Chen, Rongpeng Li, Zhifeng Zhao, Yuntao Liu, Xing Xu, Honggang Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09184v1 Announce Type: new Abstract: Large language model (LLM) agents provide a promising interface for command-line-based network operations, but a plausible command may still fail or introduce operational risk after execution. Existing approaches mainly focus on command generation or f...

📖 Read original article


Author: Tanel Tammet
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2608.09190v1 Announce Type: new Abstract: GK is a query-directed first-order prover that extends ordinary resolution-based proof search with explicit positive and negative claims, numerical confidence values, and prioritized default rules with exceptions. It works directly with non-ground clau...

📖 Read original article


169. Signature-Guided Capacity Occupancy for Dense Expert Merging ​

Author: Lingching Tung, Chi-Jui Kim, Beicheng Xu, Yuchen Wang, Bin Cui
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09201v1 Announce Type: new Abstract: Dense expert merging combines domain-specialized language models into one single checkpoint, typically by admitting task-vector support in weight space. However, this admission is governed by three decisions that existing methods answer only partially:...

📖 Read original article


170. CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving ​

Author: Junyao Wang, Yulin Xu, Yu Li, Pramod Khargonekar, Mohammad Abdullah Al Faruque
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09202v1 Announce Type: new Abstract: Modern autonomous vehicles are equipped with multiple sensors, such as cameras, LiDAR, and radar, for comprehensive environmental perception. However, robust cross-modal feature fusion remains a critical challenge, as the reliability of each sensor var...

📖 Read original article


171. Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models ​

Author: Puneet Mathur, Manan Suri, Dinesh Manocha
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09227v1 Announce Type: new Abstract: Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive. While recent token compression methods attempt to alleviate this burden, compressing moda...

📖 Read original article


172. SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge ​

Author: Yuanchi Zhu, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Xinqi Yang, Hebao Zhu, Bokai Zhao, Tianyu Liang, Ziliang Wang, Faqiang Qian, Yunli Yang, Weiyang Shi, Qibing Ren
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance, identify hazardous interactions, explain potential accident mechanisms, and recommend preventive act...

📖 Read original article


173. An Explainable GNN Framework for Component-Level Anomaly Diagnosis ​

Author: Sena Ozgunay (IMT, ANITI, LAAS-DISCO, LAAS, Comue de Toulouse), Louise Trav'e-Massuy`es (LAAS-DISCO, Comue de Toulouse, ANITI), Jean-Michel Loubes (IMT, REGALIA), Raul Sena Ferreira (LAAS)
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09246v1 Announce Type: new Abstract: Industrial processes are complex systems composed of multiple interacting sensors that generate multivariate time series (MTS). Detecting anomalies in such systems is critical for reliability and safety, yet understanding their origin is equally import...

📖 Read original article


174. Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution ​

Author: Bohan Lin, Hejia Geng, Xinyi Xie, Heng Zhou, Qinghua Xing, Bo Liu, Chen Zhang, Yudong Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09248v2 Announce Type: new Abstract: Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirely on text-level signals such as task descriptions, verbal reflections, and experience-derived rules, while the mo...

📖 Read original article


175. SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance ​

Author: You Lu, Xinyu Huang, Bihuan Chen, Xin Peng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable procedural knowledge, agents may still execute them unreliably. Even when an agent has demonstrated th...

📖 Read original article


176. Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline ​

Author: Morris Lee
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questions the warehouse cannot answer, deprecated columns after a schema change, and queries that execute ...

📖 Read original article


177. Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation ​

Author: Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Anurag Koul, Zeyu Liu, Shafiq Joty
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.09263v1 Announce Type: new Abstract: Outcome verifiers score completed reasoning traces but do not assign credit to intermediate tokens. Privileged self-distillation attempts to fill this gap by rescoring a model's own rollout with training-only information. A token likelihood change, how...

📖 Read original article


178. Entropy-based Code Adversarial Translation for Real-world Repository Migration ​

Author: Yushun Tang, Yisen Cao, Zhicheng Chen, Lin Peng, Junkang Mao, Fengyi Song, Yantao Jia
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.SE

arXiv:2608.09273v2 Announce Type: new Abstract: LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produces a runnable application because long-horizon translation challenges LLM-based agents' ability to maintain repo...

📖 Read original article


179. P$^{3}$: Joint Program-and-Proof Planning for Verified Code Generation ​

Author: Zenan Li, Ziran Yang, Peiyang Song, Zhaoyu Li, Kaiyu Yang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.PL

arXiv:2608.09277v1 Announce Type: new Abstract: Verified code generation asks a large language model (LLM) to generate both an executable program and a machine-checkable proof that the program meets a formal specification, promising software that is correct by construction. The de facto workflow dec...

📖 Read original article


180. MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence ​

Author: Chenxu Du, Kang An, Tengyue Wang, Zhongyu Yang, Xinqi Yang, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian, Yunli Yang, Qibing Ren
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual eviden...

📖 Read original article


181. ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons ​

Author: Adrian Li, Kelong Mao, Yudong Guo, Heming Xia, Xinwei Yang, Lirui Luo, Jace Wong, Pu Yao, Sulong Xu, Simiu Gu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise in device setup, meal preparation, event planning, and group takeout ordering, requiring joint reas...

📖 Read original article


182. CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation ​

Author: Harmanjot Singh, Abhra Dubey, Jorge Alejandro Amador Herrera
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG, cs.RO

arXiv:2608.09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, support controlled edits, match a reference structural response under a declared analysis, and connect ...

📖 Read original article


183. Linearized 2-Simplicial Attention ​

Author: Aritra Das, Dhruman Gupta, Debayan Gupta
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09307v1 Announce Type: new Abstract: We present a linearized form of 2-simplicial attention by rewriting the trilinear score as an inner product between a composite query and a key, so that the sum over one token axis takes the same form as ordinary softmax attention. We then approximate ...

📖 Read original article


184. ASPaeroFlow: Decomposition Heuristics for Joint Air Traffic Flow & Capacity Management ​

Author: Alexander Beiser, Markus Hecher, Nysret Musliu, Georg Trausmuth, Stefan Woltran
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09315v1 Announce Type: new Abstract: While mathematical models act as vital decision support systems for operational Air Traffic Flow and Capacity Management (ATFCM), existing approaches isolate Air Traffic Flow Management (ATFM) from Dynamic Airspace Configuration (DAC). This separation ...

📖 Read original article


185. CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning ​

Author: Ambuj Mehrish, Sebastiano Vascon
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09324v1 Announce Type: new Abstract: On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding those that match the majority vote over $N$ sampled answers. That vote discards a correct answer whene...

📖 Read original article


186. GeoPhysAdapter: Scale-Matched Geophysical Adaptation for Cross-Domain Landslide Mapping with Vision Foundation Models ​

Author: Zhihang Liu, Mei-Po Kwan, Jinlin Wu, Hao Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.09325v1 Announce Type: new Abstract: Newly triggered landslides rarely carry immediate annotations, so cross-domain transferability determines the value of landslide mapping for emergency response and regional risk assessment. Vision foundation models have strengthened representational tr...

📖 Read original article


187. Control-Oriented Scenario Tree Construction through Reinforcement Learning ​

Author: Fabio Pavirani, Bert Claessens, Pierre Pinson, Chris Develder
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SY, eess.SY

arXiv:2608.09335v1 Announce Type: new Abstract: Multistage stochastic model predictive control (MPC) handles uncertainty by optimizing over a scenario tree, a finite branching approximation of future outcomes constructed from sampled forecasts. To build such a tree, conventional methods focus on mat...

📖 Read original article


188. LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling ​

Author: Jinbo Li, Chuanhao Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09343v1 Announce Type: new Abstract: Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most methods treat the simulator as a black box: aggregate scores rank candidates without revealing why they fail or which policy logic should change. We p...

📖 Read original article


189. CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits ​

Author: Xinqi Yang, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian, Yunli Yang, Qibing Ren
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.09374v1 Announce Type: new Abstract: Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, select a physical model, formulate coupled equations, propagate intermediate quantities, and preserve u...

📖 Read original article


190. OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks ​

Author: Siqi Wang, Xinlin Li, Zhenglin Li, Li Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09380v1 Announce Type: new Abstract: Long-horizon complex tasks require agents to repeatedly observe states, formulate plans, invoke tools, verify results, and recover from failures in continuously changing environments. However, such control experience often remains confined to a single ...

📖 Read original article


191. KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models ​

Author: Chen Qiu, Ziwu Liu, Chao Fei, Guozhong Li, Panos Kalnis
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09412v1 Announce Type: new Abstract: KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct executions fail nor why. We present KVDiagnosis, a diagnostic dataset and benchmark with three contributions. First, a 25-method taxonomy groups me...

📖 Read original article


192. Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models ​

Author: Zhi Zeng, Cheng Zhang, Zesheng Yang, Rendong Pi, Jiaying Wu, Di Zhang, Zihan Ma, Guodong Li, Zhou Yang, Yu Xiang, Yifei Zheng, Minnan Luo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09435v1 Announce Type: new Abstract: Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language m...

📖 Read original article


193. Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity ​

Author: Zihan Wang, Anglin Liu, Rongyi Wang, Dantong Li, Yi Lu, Siqing Yuan, Hongxia Xu, Zhongtian Long, Jintai Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on conditions, medications, and geriatric risks that users may omit. We introduce ATLAS, a coupled gra...

📖 Read original article


194. From Prompt to Harness: Coderlet from Scratch ​

Author: Mengfan Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09480v1 Announce Type: new Abstract: A model alone does not determine how a programming agent acts. What the model sees, how actions enter the environment, how feedback returns, and how one run affects the next all depend on how the harness is organized. Minimal examples usually show only...

📖 Read original article


195. Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents ​

Author: Neel Tushar Shah, Manglam Kartik, Akshat Karkar
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09485v1 Announce Type: new Abstract: Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission, false consensus, and manipulative framing. We argue that Cooperative AI evaluations should separate...

📖 Read original article


196. Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification ​

Author: Karim Zaghw, Andrew Pashea, Marc Pritsch, Wouter Nuijten, Karl Friston, Lancelot Da Costa
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.09512v1 Announce Type: new Abstract: Active inference offers a unified framework for perception, learning, and action, but scaling discrete active-inference models to rich spatial and temporal domains remains difficult. Renormalising generative models (RGMs) address this challenge by comp...

📖 Read original article


197. One Adapter Pair per Model: A Universal Activation Interface for Language Models ​

Author: Su-Hyeon Kim, Jiwan Mun, Yo-Sub Han
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09521v1 Announce Type: new Abstract: Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to be rebuilt or rediscovered for each new language model. We present a Universal Activation Bus, a fra...

📖 Read original article


198. verdi: retrieval is not transfer for continual world model optimization ​

Author: Junyu Wu, Shiqin Nie, Youyi Kou, Baohua Yin, Guocai Yao, Qingyu Chen, Jingheng Ma, Shiji Zhou, Hongyong Song, Mingchen Zhuge, Sen Cui, Changshui Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09537v1 Announce Type: new Abstract: Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence. However, optimizing a pretrained world model toward a user-specified objective remains difficult: each campaign typically rediscovers optimization...

📖 Read original article


199. Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents ​

Author: Tianjun Pan, Yuan Li, Hongda Wang, Linbo Jin, Mengfei Song, Lei Gao, Qiming Shi, Shaokang Fu, Jiarong Zhao, Chengyu Wang, Chengfu Huo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09555v1 Announce Type: new Abstract: External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Yet their effectiveness depends not only on skill quality, but also on whether the policy can translate the provid...

📖 Read original article


200. The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games ​

Author: Fatemeh Seyedin, Adrian Weller, Jinhyuk Yun, Mahmoudreza Babaei
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move from individual tools to participants in multi-agent organizations, an important question arises: do th...

📖 Read original article


201. ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization ​

Author: Hao Sui, Simeng Qin, Jie Liao, Xiaojun Jia, Bing Chen, Yang Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09577v1 Announce Type: new Abstract: Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisoned skill can persistently compromise every agent that installs it. However, existing skill attacks either fire on ...

📖 Read original article


202. CoRCi: Cross-Reconstruction of Coherent Interests Modeling in Cross-Domain Sequential Recommendation ​

Author: Qingtian Bian, Tieying Li, Marcus de Carvalho, Jiaxing Xu, Hui Fang, Yiping Ke
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09580v1 Announce Type: new Abstract: Cross-Domain Sequential Recommendation (CDSR) aims to alleviate data sparsity by transferring dynamic user interests across related domains. A key challenge lies in effectively bridging these domains. In single-domain modeling, models cannot distinguis...

📖 Read original article


203. ICM Out! Better Tournament Strategy from Computed Continuations, vs. Solvers and LLMs ​

Author: Boning Li, Longbo Huang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.GT, cs.MA

arXiv:2608.09586v1 Announce Type: new Abstract: The Independent Chip Model (ICM) converts tournament chips into reference prize equity, and policies are routinely constructed against those values. Because ICM reads only stack sizes, it omits action order, blind obligations, and seat rotation, and it...

📖 Read original article


204. From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization ​

Author: Achille Jacquemond, Yuma Ichikawa, Akira Sakai
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09595v1 Announce Type: new Abstract: Compressing large language models to two bits or fewer is increasingly feasible through block-wise post-training quantization; cross-block variants reconstruct neighboring Transformer blocks within a moving window. In the fixed two-block setting studie...

📖 Read original article


Author: Youssef A. Elhagrasy, Ian Hill, Andr'e Ivanov
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY

arXiv:2608.09622v1 Announce Type: new Abstract: Reliability qualification of advanced semiconductor devices requires sequential stress decisions that balance characterization objectives against multiple competing failure mechanisms. Current practice relies on static test plans derived from populatio...

📖 Read original article


206. Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines? ​

Author: Hui Xue, Fan Yang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09629v1 Announce Type: new Abstract: Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop. We ask whether this task-specific procedure remains necessary when...

📖 Read original article


207. Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics ​

Author: Yen-Shan Chen, Yu Chian Duan, Chih-En Kuo, Jian-Bin Wu, Yun-Nung Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY, cs.GT

arXiv:2608.09638v1 Announce Type: new Abstract: Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight. We present Avalon-ToM-Bench, a ...

📖 Read original article


208. Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching ​

Author: Yuke Li, Xuehan Hou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09654v1 Announce Type: new Abstract: GUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operate directly on screenshots. The core task, GUI grounding, requires translating abstract user instructions into pre...

📖 Read original article


209. Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models ​

Author: Shulin Tian, Ziqi Huang, Fan Zhang, Hongyuan Zhu, Yu Qiao, Ziwei Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09666v1 Announce Type: new Abstract: Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands sampling hundreds or thousands of images or videos, which is computationally expensive. Existing evaluation meth...

📖 Read original article


210. Adaptive Semantic Capacity Allocation for Parallel Generative Recommendation ​

Author: Chenxi Li, Yuchen Lu, Xu Yang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09685v1 Announce Type: new Abstract: Autoregressive semantic ID recommenders are constrained by expensive beam-search decoding, which limits the practical length of item identifiers. Parallel generation methods alleviate this bottleneck by predicting all semantic ID tokens simultaneously,...

📖 Read original article


211. Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models ​

Author: Kevin Murphy
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09696v2 Announce Type: new Abstract: Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, causal model, not a curve fit; and learning such a model requires \emph{experiments}, because passive data leaves ...

📖 Read original article


212. Matryoshka Language Model Suites ​

Author: Nathan Godey, Yoav Artzi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.09703v1 Announce Type: new Abstract: Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained en...

📖 Read original article


213. Second-Order Muon Done Right: A Principled Marriage of Spectral Geometry and Curvature ​

Author: Tong Che
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09763v1 Announce Type: new Abstract: Muon's polar update is exact for an unweighted spectral geometry. We introduce GO-MUON, which uses a matched data-dependent geometry and reuses it across several optimization steps. Conditioned on any positive-definite left and right maps, its raw upda...

📖 Read original article


214. AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting ​

Author: Fan Yang, Nan Chen, Yijie Dong, Yuchen Zhang, Wei Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.LG

arXiv:2608.09775v1 Announce Type: new Abstract: Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels differ in periodicity and distribution drift, while their concentration trajectories contain both mul...

📖 Read original article


215. CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation ​

Author: Yaoning Yu, Kai-Min Chang, Ye Yu, Yi-Chia Wang, Haojing Luo, Haohan Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SI

arXiv:2608.09790v2 Announce Type: new Abstract: Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these discussions requires more than just generating individual comments, the generated threads should also match how r...

📖 Read original article


216. Mismatch Matters: On-Policy Distillation Beyond Token Agreement ​

Author: Zichao Yu, Chengzhi Yu, Shengze Xu, Yujin Han, Bingqing Jiang, Xu Wang, Difan Zou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.09836v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect token agreement with the teacher des...

📖 Read original article


217. CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems ​

Author: Aimilios Hadjiliasi, Louis Nisiotis
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09848v1 Announce Type: new Abstract: The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a challenge, even with today's advancements in technology. Existing architectures are often focused on ...

📖 Read original article


218. Agentic Auto-Research is Fuzz Testing ​

Author: Yifeng He, Jicheng Wang, Yinzhe Zhao, Jiachen Liu, Hao Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.09855v1 Announce Type: new Abstract: Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the proposer and ranking more samples with a learned judge or human reviewers. We argue that this generate-and-rank p...

📖 Read original article


219. Towards Expert-level Medical AI for Real-time Video Consultations ​

Author: Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Li'evin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemengway, Sunny Virmani, David Racz, Carey Radebaugh, Jo"elle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Po-Hsuan Cameron Chen, Mike Schaekermann, Anil Palepu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV

arXiv:2608.09861v1 Announce Type: new Abstract: Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions...

📖 Read original article


220. ArchAgent v2: A Case Study with the Data Prefetching Championship ​

Author: Abraham Gonzalez, Raghav Gupta, Akanksha Jain, Hanna Alam, Alexander Novikov, Po-Sen Huang, Matej Balog, Marvin Eisenberger, Sergey Shirobokov, Ng^an V~u, Hank Levy, Borivoje Nikoli'c, Sagar Karandikar, Martin Dixon, Parthasarathy Ranganathan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.AR

arXiv:2608.09874v1 Announce Type: new Abstract: Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarchitecture discovery remains challenging due to vast search spaces, strict hardware budgets, and long simulation t...

📖 Read original article


221. SHE: Trajectory-driven Safety Harness Evolution for LLM Agents ​

Author: Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CV

arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harness as a fixed deploym...

📖 Read original article


222. DSLE: A Learning Environment for Dark Souls Boss Encounters ​

Author: Derin Gezgin, Jim O'Connor, Tanner Goodwin, Gary B. Parker
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.NE

arXiv:2608.09902v1 Announce Type: new Abstract: We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounters of Dark Souls: Remastered as game-playing agent benchmarks through a Gymnasium-style interface. DSLE combines real-time combat, high-...

📖 Read original article


223. GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis ​

Author: Alban Puech, Matteo Mazzonelli, Tamara R. Govindasamy, Mangaliso Mngomezulu, H'ector Maeso-Garc'ia, Thomas Tolhurst, Javad Bayazi, Ali Moeini, Naomi Simumba, Celia Cintas, David Nelischer, Romeo Kienzler, Jonas Weiss, Anna Varbella, Florian D"orfler, Gabriela Hug, Martin Mevissen, Juan Bernab'e-Moreno, Fran\c{c}ois Mirall`es, Hendrik F. Hamann, Etienne Vos, Thomas Brunschwiler
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.09921v1 Announce Type: new Abstract: Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from engineering domains such as power system analysis, where strict physical consistency must be enforced. We present GENCO (GEometric Neur...

📖 Read original article


224. From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents ​

Author: Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai
Published: 8/11/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.05235v1 Announce Type: cross Abstract: Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions. Yet a completed trajectory is not automatically evidence: generated ar...

📖 Read original article


225. Coordinated incentives in AI-generated misinformation governance ​

Author: Qin Li, Gui Zhang, Minyu Feng, Matjaz Perc, Attila Szolnoki
Published: 8/11/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI

arXiv:2608.07070v1 Announce Type: cross Abstract: With the rapid diffusion of AI-generated content, AI-driven misinformation is becoming increasingly pervasive and difficult to govern, undermining information credibility and social trust. This study models the strategic interdependence among a gover...

📖 Read original article


226. Application of Artificial Intelligence for Fraudulent Banking Operations Recognition ​

Author: Bohdan Mytnyk, Oleksandr Tkachyk, Nataliya Shakhovska, Solomiia Fedushko, Yuriy Syerov
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, cs.CR, cs.CY

arXiv:2608.07471v1 Announce Type: cross Abstract: This study considers the task of applying artificial intelligence to recognize bank fraud. In recent years, due to the COVID19 pandemic, bank fraud has become even more common due to the massive transition of many operations to online platforms and t...

📖 Read original article


227. Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use ​

Author: Yizhu Gao, Zhongzhou Chen, Min Li, Xiaoming Zhai
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.07475v1 Announce Type: cross Abstract: Generative Artificial Intelligence (GenAI) presents a governance challenge for STEM assessment. Unrestricted access can enable task outsourcing that undermines the validity of traditional assessments, while blanket prohibitions are difficult to enfor...

📖 Read original article


228. Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources ​

Author: Sundaraparipurnan Narayanan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.07477v1 Announce Type: cross Abstract: This thesis examines the fairness of Automated Machine Learning (AutoML) tools in human resource hiring systems through the combined lenses of regulation, business strategy, and Human-Computer Interaction (HCI). It argues that fairness is no longer m...

📖 Read original article


229. Cross-Model Humor Preference Modeling with Cards Against Humanity ​

Author: Victor Winter, Farhan Lakhany
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.07481v1 Announce Type: cross Abstract: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style task. Two models - GPT-4o as Czar and Claude Opus-4.5 as Player - are evaluated on a binary humor-s...

📖 Read original article


230. Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents ​

Author: Yingying Guo, Zhuoxuan Ju, Ruibo Ming, Ruicheng Feng, Jinjin Gu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.07490v1 Announce Type: cross Abstract: Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players learn from repeated interaction. We study experience-sensitive game learning: how gameplay experience changes t...

📖 Read original article


231. How to Ask the AI: A User Perspective Survey for Large Language Model Prompting ​

Author: Yiqun Zhang, Yunfan Zhang, Mingjie Zhao, Sen Feng, Yiu-ming Cheung
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as plan a three-day Vienna trip'', solve the attached mathematical problem'...

📖 Read original article


232. EmoPatient: An Emotion-Directed Patient Simulator for Realistic Palliative Care Communication Training ​

Author: Yining Wu, Tianshu Du, Jinrui Fang, Chi Zhang, Sonal Admane, Ying Ding
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.07495v1 Announce Type: cross Abstract: Effective communication during palliative care discussions is a critical clinical skill, yet training clinicians to manage complex patient emotions remains challenging. Large language model (LLM)-based patient simulators provide a scalable approach f...

📖 Read original article


233. Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation ​

Author: Ljubisa Bojic, Ljiljana Matic, Joerg Matthes, Milan Cabarkapa, Bojana Dinic, Jue Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.LG, cs.MA

arXiv:2608.07498v1 Announce Type: cross Abstract: Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing. A central open question is whether persona-prompted LLMs can simulat...

📖 Read original article


234. Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients ​

Author: Jiading Zhu, Xinyu Cindy Wang, Thomas Nguyen, Yan Qing Lee, Osnat C. Melamed, Peter Selby, Jonathan Rose
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL

arXiv:2608.07499v1 Announce Type: cross Abstract: The development and benchmarking of Large Language Model (LLM)-based Motivational Interviewing (MI) counsellors now often rely on LLM-based simulated clients. Prior work on simulated clients, however, has not aligned with the specific tasks fundament...

📖 Read original article


235. Harnessing Abundance: A Generativity Perspective on Human-GenAI Collaboration ​

Author: Yoram M Kalman, Yun Wan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.07500v1 Announce Type: cross Abstract: Research on human-GenAI collaboration yields conflicting findings: GenAI can enhance creativity yet reduce collective diversity, with uneven benefits across skill levels. Rather than treating these as contradictions, we argue they reflect a core feat...

📖 Read original article


236. Innovating with Generative AI: A Human Bottleneck Framework ​

Author: Julian De Freitas, Ayelet Israeli, Gideon Nave, Artem Timoshenko, Olivier Toubia
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, econ.GN, q-fin.EC

arXiv:2608.07504v1 Announce Type: cross Abstract: We propose a human bottleneck perspective for understanding how generative AI transforms the innovation process. The central premise is that many constraints traditionally plaguing the innovation process are cognitive and social in origin, rooted in ...

📖 Read original article


237. JaleesBench: Are AI Assistants Good Spiritual Company? ​

Author: M. Waleed Kadous (iaser.ai, Faith Family Technology Network), Benjamin Olsen (Faith Family Technology Network)
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL

arXiv:2608.07508v1 Announce Type: cross Abstract: Large language models are already advisors to millions of people of faith who bring them real decisions. The pressing question for a person of faith is not what a model knows or professes but what its counsel does to the person who receives it. We in...

📖 Read original article


238. PIVOT: Preference-based Intervention Vectors for Pedagogical Tutor Steering ​

Author: Fares Fawzi, Jiaxu Zhao, Tanya Nazaretsky, Tanja K"aser
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.07509v1 Announce Type: cross Abstract: LLMs are increasingly used for conversational tutoring, but effective tutoring requires more than correct answers. Tutors must choose when to scaffold reasoning, hint, give feedback, explain, or invite reflection. Existing prompting and training meth...

📖 Read original article


239. How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare ​

Author: Dorothee Amelung, Andrew M. Bean, Sabine C. Herpertz, Felix H. Krones, Guy Parsons, Adam Mahdi, Isabella Schneider
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL, cs.CY

arXiv:2608.07511v1 Announce Type: cross Abstract: Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (LLMs) have been proposed for tasks such as triaging patients, report drafting or translating medical jargon to s...

📖 Read original article


240. EMMR: Emotion-Mediated Multimodal Reasoning for Personality Assessment in Asynchronous Video Interviews ​

Author: Dongsheng Hu, Tianyi Zhang, Chuang Liu, Yuan Zong Yong Li, Wenming Zheng, Xiu-xiu Zhan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.07512v1 Announce Type: cross Abstract: Asynchronous Video Interviews (AVIs) have become increasingly popular for personality assessment. Recent large language models (LLMs) have shown potential for personality assessment from transcribed interview responses. However, text-centered methods...

📖 Read original article


241. Representation Matters in Longitudinal Affective Computing ​

Author: Igor Matias, Maximilian Haas, Eric J. Daza, Matthias Kliegel, Katarzyna Wac
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.07518v1 Announce Type: cross Abstract: Longitudinal, in-the-wild, wearable sensing yields day-level physiology, sleep, activity, and environmental streams, whereas affect and cognition are labeled only episodically (per waves). We recast this cadence mismatch as a temporal representation ...

📖 Read original article


242. KumbhDoot: A Scale-Ready, LLM-Bounded Architecture for Mass-Gathering Public-Service Assistants ​

Author: Saurabh Sakalkar, Abhishek Singh, Ramesh Raskar
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.07520v1 Announce Type: cross Abstract: Mass religious gatherings such as the Kumbh Mela concentrate tens of millions of people into a single region over a few weeks, producing intense, repetitive, multilingual, and safety-critical demand for information. The default response, a conversati...

📖 Read original article


243. From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations ​

Author: Hongfei Yan, Jiangkai Xiong, Yiqing Li, Chong Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.PL

arXiv:2608.07523v1 Announce Type: cross Abstract: Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions large language models from benchmark evaluation targets to auxiliary evidence sources for interpreting exam difficu...

📖 Read original article


244. Unified Hallucination Fuzzing for Multimodal Large Language Models ​

Author: Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.07525v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical...

📖 Read original article


245. DocAtlas: Long-Document Understanding as Mutable-State Interaction ​

Author: Hongchen Wei, Yuanzhe Wang, Bei Liu, Yifan Yang, Qi Dai, Kai Qiu, Yunsheng Li, Dongdong Chen, Chong Luo, Zhenzhong Chen, Baining Guo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.07527v1 Announce Type: cross Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually select evidence from a static index before generation, while recent agentic...

📖 Read original article


246. WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management ​

Author: Yi Zhang, Hongyang Wang, Zheng Hao Leong, Zihao Wu, Kaijun Lin, Zhixing Pan, Qixun Huangfu, Wei Ren, Wenyan Wu, Fangyun Wang, Wenting Yu, Hengyu Lin, Muling Yang, Zongguo Wen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.07529v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than professional decisions un...

📖 Read original article


247. Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards ​

Author: Cheng Ruoxi, Ma Haoxuan, Zhang Hongyi, Zhang Junming, Duan Ranjie, Xia Qiaolin, Wang Hao, Lu Yu, Shi Haibo, Ma Xingjun
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.07531v1 Announce Type: cross Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations ...

📖 Read original article


248. Ultraconstructive Model Theory via Bounded Adversarial Finite Structures ​

Author: Mirco A. Mannucci
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LO, cs.AI

arXiv:2608.07534v1 Announce Type: cross Abstract: Ultraconstructive Model Theory (UCMT) replaces idealized satisfaction, at finite compu- tational scale, by bounded adversarial survival. A finite partial structure is tested by an Opponent (Devil) drawing legal challenges from a bounded attack surfac...

📖 Read original article


249. Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards ​

Author: Xi Li, Shu Zhao, Xiaohan Zou, Fei Zhao, Fuxiao Liu, Yusen Zhang, Cheng Han, Yushun Dong, Jiaqi Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY

arXiv:2608.07535v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this architectural shift reshapes the safety landscape of machine learning. In...

📖 Read original article


250. An evolutionary model of animats with VLM-based subjective evaluation ​

Author: Shota Miyazaki, Takaya Arita, Reiji Suzuki
Published: 8/11/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CL, cs.HC, cs.MA

arXiv:2608.07537v1 Announce Type: cross Abstract: In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language Model (VLM) into the fitness evaluation and selection processes of a genetic algorithm. As the target of evolution, we employ virtual soft ro...

📖 Read original article


251. NeuroPilot: An Agent-Driven Smart Pipeline for Processing, Quality Control, and Managing Neuroimages ​

Author: Yiyao Chen, Yucheng Li, Jungong Tong, Shaoqi Wang, Kunhao Zhou, Ziquan Wei, Monica Murea, Marissa DiPiero, Tingting Dan, Guorong Wu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MA

arXiv:2608.07541v1 Announce Type: cross Abstract: Transforming raw neuroimage archives into analysis-ready derivatives relies on three brittle stages: data standardization, modality-specific preprocessing, and quality control (QC). While individual neuroimaging tools are well developed, their orches...

📖 Read original article


252. Performance of large language models in the optical diagnosis of colorectal polyps ​

Author: Joshua C. Vences, William T. Tran, Nikko Gimpaya, Catharine M. Walsh, Rishad J. Khan, Robert Bechara, Asher C. Wiggins, Celine N. Rousan, Kaitlyn V. G. L. Morgado, Angie Ibrahim, Kevin H. M. Kuo, Daniel von Renteln, Alexander Hann, Dennis L. Shung, Michael A. Scaffidi, Charles M'enard, Joshua Landy, Samir C. Grover
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07543v1 Announce Type: cross Abstract: Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis. We aimed to evaluate the diagnostic accu...

📖 Read original article


253. MOSAIC: Adversarial Co-evolution of Specialist Heuristics and Problem Instances for LLM-based Automated Heuristic Design ​

Author: Oguzhan Gungordu, Siheng Xiong, Faramarz Fekri
Published: 8/11/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2608.07544v1 Announce Type: cross Abstract: Automated heuristic design (AHD) with large language models (LLMs) has produced strong heuristics for combinatorial optimization problems (COPs). Yet existing frameworks optimize for average performance on a small fixed dataset and steer the search w...

📖 Read original article


254. DarwinX: Evolving Agent Harnesses Through Natural Selection ​

Author: Yifan Zhang, Yutong Dai, Juntao Tan, Luyu Yang, Rishi Mullur, Thai Hoang, Zhiyuan Hu, James Zhu, Phil Mui, Silvio Savarese, Ran Xu, Zeyuan Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.LG, cs.SE

arXiv:2608.07545v1 Announce Type: cross Abstract: An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other ta...

📖 Read original article


255. Generalizing deep reinforcement learning across cable-driven parallel robot configurations with actuator-level policies ​

Author: Abir Bouaouda (CRAN, UIR), Mohamed Boutayeb (CRAN, UIR), Fran\c{c}ois Charpillet (LARSEN), Dominique Martinez (LORIA, ISM), R'emi Pannequin (CRAN)
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.07546v1 Announce Type: cross Abstract: Cable-driven parallel robots (CDPRs) present diverse configurations and complex control challenges, which can be addressed by deep reinforcement learning (DRL) by learning their nonlinear dynamics. However, DRL methods often require extensive trainin...

📖 Read original article


256. Learning an Interior Layout Policy in a Domain Specific Language Action Space ​

Author: Yuhao Lu, Weichen Zhang, Wenyi Xiao, Haohui Chen, Yiyun Fei
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07547v1 Announce Type: cross Abstract: Indoor scene layout generation is a challenging task in interior design. Existing methods often oversimplify the task by reducing room conditions to coarse 3D bounding boxes and neglecting structural elements such as doors and windows. More fundament...

📖 Read original article


257. P2Voxel: Pyramid Pivot Voxelization for 3D Mesh Tokenization ​

Author: Zhenhong Sun, Haozhe Liu, Yifu Wang, Xibin Song, Senbo Wang, Huadong Mo, Daoyi Dong, Hongdong Li, Pan Ji
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07549v1 Announce Type: cross Abstract: Triangle meshes provide explicit and accurate surface geometry, yet their irregular topology connectivity makes 3D mesh tokenization a geometric sampling problem: how to sample and organize geometric evidence into compact, structured and learnable to...

📖 Read original article


258. MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures ​

Author: Zhuoning Xu, Xiucheng Zhang, Hanjun Luo, Yingbin Jin, Yinpeng Dong, Hanan Salam
Published: 8/11/2026, 4:00:00 AM
Categories: cs.MA, cs.AI

arXiv:2608.07556v2 Announce Type: cross Abstract: Multi-agent systems (MAS) decompose long-horizon tasks across supervisors and subagents, but delegated goals do not necessarily carry their original authorization boundaries. Existing safety benchmarks mainly study adversarial compromise, while work ...

📖 Read original article


259. AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization ​

Author: Peng Xu, Chengcheng Wang, Shaohua Wan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV

arXiv:2608.07557v1 Announce Type: cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist end-to-end paradigms show great promise but typically rely on massive language models containing billi...

📖 Read original article


260. Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth ​

Author: Sho Mitarai, Hikaru Kayo, Hisashi Ozaki, Yuichiro Imai, Megumi Nakao
Published: 8/11/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV

arXiv:2608.07564v1 Announce Type: cross Abstract: In digital dentistry and oral surgery, the registration of jawbone CT and intraoral scanner (IOS) data is essential for integrating internal bone structure with high-resolution dental surface geometry. However, this registration is challenging becaus...

📖 Read original article


261. What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems ​

Author: Zhijing Zhang, Jinpeng Yu, Xin Song, Bingnan Li, Chuyue Li, Changhui Du, Xiaolin Fang, Jiaming Liu, Ruihua Huang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07565v1 Announce Type: cross Abstract: Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactions, leaving image-creation conversations underexplored. In image-creation tasks, useful follow-up ed...

📖 Read original article


262. Temporal Generalization in fNIRS-Based Autism Classification: A Cross-Time-Window Transfer Benchmark ​

Author: Marios Petrov, Sahana Vinayak, Targol Bakhtiarvand, Moses Smith Guddah, Adham Atyabi, Frederick Shic, Kevin A. Pelphrey
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07567v1 Announce Type: cross Abstract: Functional near-infrared spectroscopy (fNIRS) is a promising modality for autism spectrum disorder (ASD) classification, yet existing approaches assume temporally aligned evaluation. In practice, the optimal observation window varies across subjects ...

📖 Read original article


263. Latent-Frequency Validity: Fast Spectral Editing with Screened Video-VAE Transfer Operators ​

Author: Bowen Xue, Jiafeng Xiong, Xin Quan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.07569v1 Announce Type: cross Abstract: Direct spectral editing in video-VAE latents can control noise, flicker, smoothness, and frequency content without a decode--filter--reencode pass. However, video VAEs may redistribute pixel-space frequency bands across latent channels, and latent ed...

📖 Read original article


264. COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping ​

Author: Rui Yang, Wei Zhou, Dingyong Gou, Xiaohui Cui, Cong Li, Yinyin Gong, Yipo Huang, Jiliang Zhao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07570v1 Announce Type: cross Abstract: Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-explain methods largely treat explanation as post-hoc text generation and overlook composition, a ke...

📖 Read original article


265. BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference ​

Author: Jinlong Yang, Jinke Wu, Lizilin, Yao Zhou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07572v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their massive computational overhead, temporal feature caching has been proposed to bypass redundant computations. Howev...

📖 Read original article


266. Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects ​

Author: Felix Schaller
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO

arXiv:2608.07577v1 Announce Type: cross Abstract: A closed-set detector for autonomous driving must assign every object one of a fixed set of labels. On an object outside that set (a horse-drawn carriage, road debris, livestock on a rural road) it can only force a confident but wrong specific label ...

📖 Read original article


267. Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real ​

Author: Abdullah Naeem, Anav Katwal, Ayon Dey, Noman Khan, Md Tamjidul Hoque
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07579v1 Announce Type: cross Abstract: The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only. We use two RGB-only routes...

📖 Read original article


268. Multi-Branch Policy Optimization for Multimodal Large Language Models ​

Author: Shuai Lyu, Yuning Gong, Ruiling Gao, Xiaoran Shang, Zhonghong Ou, Ping Zong, Yifan Zhu, Yuan Sun, Yang Qin, Peng Hu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07581v1 Announce Type: cross Abstract: Group-based reinforcement learning methods for multimodal large language models typically rely on trajectory-level credit assignment that applies a single advantage to all tokens in a response. However, multimodal reasoning involves substantially hig...

📖 Read original article


269. Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning ​

Author: Kadharmoideen Fadurudeen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.IR

arXiv:2608.07593v1 Announce Type: cross Abstract: Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Existing weather-aware food and point-of-interest recommenders, however, typically treat weather generi...

📖 Read original article


270. Scaling Inherently Interpretable Language Models ​

Author: Guide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail, Giang Nguyen, Isaac Plant, Muawiz Chaudhary, Nathaniel Monson, Saqib Azim, Zhichen Guo, Julius Adebayo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.07594v1 Announce Type: cross Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than rev...

📖 Read original article


271. Enhanced Real-Time 6-DOF Extended Reality Catheter Tracking for Evaluating Potential Improvement in Efficiency, Precision, and Depth Perception for Cardiac Interventions ​

Author: Mohsen Annabestani, Sandhya Sriram, Andrew Kuzemczak, S. Chiu Wong, Alexandros Sigaras, Bobak Mosadegh
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.HC, eess.IV

arXiv:2608.07606v1 Announce Type: cross Abstract: Despite advances in 3D ultrasound, most percutaneous cardiac interventions still rely on 2D visualization, limiting depth perception and spatial understanding. To address this challenge, we developed an Extended Reality (XR)-based platform that enabl...

📖 Read original article


272. Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Screening Assays ​

Author: Xiaohua Douglas Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: stat.AP, cs.AI, q-bio.QM, stat.ML

arXiv:2608.07609v1 Announce Type: cross Abstract: High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary screens typically use only a single replicate per test substance. This sparsity makes conventional machine-lea...

📖 Read original article


273. PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue ​

Author: Shibo Wang, Zicheng Zhang, Libo Wang, Junfeng Ma
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM

arXiv:2608.07631v1 Announce Type: cross Abstract: LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue state faster than clients can play it, subsequent user speech may be interpreted based on content th...

📖 Read original article


274. Adversarial Attacks on Deep OCR Systems ​

Author: Wenbo Sun, Hongzong LI, Yanyun Wang, Jiahao MA, Shuxin Zhuang, Rong Feng, Shiqin Tang, Zi Liang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CV

arXiv:2608.07636v1 Announce Type: cross Abstract: Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its increased complexity may introduce new security vulnerabilities. In this...

📖 Read original article


275. SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment ​

Author: Chaofan Meng, Yuhang Zheng, Yingnan Zhou, Sihan Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07639v1 Announce Type: cross Abstract: Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill selection. Recent Agent Skill research has increasingly examined Agent Skill consistency detection. E...

📖 Read original article


276. Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding ​

Author: Yeeun Choi, Youngbeom Yoo, Joon-Young Lee, Hyolim Kang, Seon Joo Kim
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2608.07663v1 Announce Type: cross Abstract: When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necessitates a two-stage paradigm: query-agnostic memory construction follow...

📖 Read original article


277. CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting ​

Author: Quang Minh Dinh, Tuan Kiet Doan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2608.07693v1 Announce Type: cross Abstract: Generative traffic video forecasting aims to synthesize long-horizon, temporally coherent future videos of traffic scenes from a short observation history and textual descriptions. In this paper, we present CosmosAlign, a generative traffic video for...

📖 Read original article


278. SpikeWorld: Fast-State Adaptation for Frozen Spiking World Models ​

Author: Ziqiao Yu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07712v1 Announce Type: cross Abstract: A predictive model receives a self-supervised signal whenever the consequence of an action is observed. Using that signal after deployment is difficult when dynamics and semantics share parameters: freezing prevents adaptation, whereas weight updates...

📖 Read original article


279. LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs ​

Author: Liad Gerstman, Aditya Dhakal, Dejan Milojicic, Avi Mendelson
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.AR, cs.LG, cs.PF

arXiv:2608.07733v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems. However, as graph sizes grow, storing and processing them entirely on a single-node CPU-GPU system...

📖 Read original article


280. Complete, Scalable, and Robust Prioritized Planning for Multi-Robot Ordered Storage and Retrieval at Maximum Capacity ​

Author: William Zhang, Tzvika Geft, Jingjin Yu, Kostas Bekris
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.MA

arXiv:2608.07734v1 Announce Type: cross Abstract: Automated warehouses face a fundamental trade-off between maximizing storage density and achieving high retrieval throughput. While puzzle-based storage (PBS) architectures increase capacity by eliminating aisles, coordinating multiple robots in thes...

📖 Read original article


281. LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks ​

Author: Saed Moradi, Benyamin Ghojogh, M. Hadi Sepanj, Yimin Yang, Ashirbani Saha
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.07749v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resources, but a single low-rank update can constrain all task-specific changes to one narrow parameter subspace. This ...

📖 Read original article


282. Multi-Task Consistency-based Detection of Adversarial Attacks ​

Author: Cong Chen, Jean-Philippe Monteuuis, Jonathan Petit
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07750v1 Announce Type: cross Abstract: Deep Neural Networks (DNNs) have found successful deployment in numerous vision perception systems. However, their susceptibility to adversarial attacks has prompted concerns regarding their practical applications, specifically in the context of auto...

📖 Read original article


283. CFD-Guided Detection of Concept Drift in Multimodal Physiologic Signals ​

Author: Farouk Ganiyu Adewumi, Timothy Oladunni, Rochak Ghimire, Kosisochukwu Ogbuanya, Sanaa Reeves, Sandy Akoy
Published: 8/11/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2608.07759v1 Announce Type: cross Abstract: Cardiovascular AI models can classify clean elec- trocardiogram (ECG) signals, but real wearable signals change because of motion, breathing, posture, sensor contact, and true clinical deterioration. This paper asks when a model should keep its predi...

📖 Read original article


284. The Anatomy of a Prompt Injection: A Component Model for Structured Analysis ​

Author: Jeremy McHugh
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.07808v1 Announce Type: cross Abstract: Four years after prompt injection was first identified in 2022, attacks are still predominantly documented as verbatim strings rather than structured exploits, despite advancing agent capabilities and threat actors embedding injections to subvert AI-...

📖 Read original article


285. Shape Mutating Expert Compression:LorExperts and BTExperts ​

Author: Inesh Chakrabarti, Sourjya Roy, Bowen Bao, Thiago Crepaldi, Spandan Tiwari, Ashish Sirasao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07814v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many expert weight matrices. Expert pruning (e.g., REAP) and merging reduce cost but sacrifice accuracy and...

📖 Read original article


286. Distilling CT Foundation Models into Editable Concept Bottlenecks for Lung Nodule Malignancy Prediction ​

Author: Fakrul Islam Tushar, Stephen Adamo, Geoffrey D. Rubin
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.07857v1 Announce Type: cross Abstract: Foundation models provide transferable CT representations, but predictions based directly on these embeddings are difficult to interpret. We developed concept bottleneck models that map two frozen CT foundation-model representations to eight radiolog...

📖 Read original article


287. Vision-Language Grounding as Bidirectional Concept Correspondence ​

Author: Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL

arXiv:2608.07886v1 Announce Type: cross Abstract: Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region. This ...

📖 Read original article


288. Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models ​

Author: Ali Janati, Kaoutar El Maghraoui, Xinyi Luo, Wenyuan Shen, Owen Zou, Yankai Mao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07890v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert. Recent theory shows that pruning experts with the smallest router-norm changes during fine-tuning can preserve accur...

📖 Read original article


289. Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety ​

Author: Hannah Cha, Neha Shukla, Solon Barocas, Alexandra Chouldechova, Eugenia Kim, Jennifer Wortman Vaughan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2608.07902v1 Announce Type: cross Abstract: Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes situations. However, existing child safety evaluations of AI lack grounding in real-world harms that ...

📖 Read original article


290. Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation: Guarantees and Empirical Limits ​

Author: Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.07913v1 Announce Type: cross Abstract: Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate for federated, differentially private, adaptively monitored retrieval-augmented generation. Clients release only ...

📖 Read original article


291. Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention ​

Author: Kasun Dewage, Marianna Pensky, Suranadi De Silva, T. H. Bandara
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.07921v1 Announce Type: cross Abstract: We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bulk and a set of spectral outliers. We validate this decomposition causally: zeroing the MP-identifie...

📖 Read original article


292. Second Order Drifting Models ​

Author: Drake Brown, Yuhao Huang, Shih-Hsin Wang, Bao Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NA, math.NA

arXiv:2608.07924v1 Announce Type: cross Abstract: Drifting models are a recent class of one-step generative models that evolve the model distribution during training using a predefined sample-based drift field. Although they avoid iterative inference, their kernel-based drift fields induce frequency...

📖 Read original article


293. ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB ​

Author: Yifan Wu, Yuhan Li, Zhenhua Wang, Ke Chen, Lidan Shou, Zonghao Chen, Liang Lin, Huan Li, Gang Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.DC

arXiv:2608.07945v1 Announce Type: cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge. Our analysi...

📖 Read original article


294. Metadata Reconstruction from Values Alone: Recovering Column Semantics in Undocumented Warehouses ​

Author: Mike Helwig
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2608.07946v1 Announce Type: cross Abstract: Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean. Production warehouses are the inverse: cryptic identifiers, partial or absent documentation. We address the problem they pose first: recovering what columns and...

📖 Read original article


295. Persistent Semantic Entities in Tool-Augmented LLM Systems ​

Author: Zhaohui Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR, cs.SE

arXiv:2608.07952v1 Announce Type: cross Abstract: Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries---largely invisible to standard debugging. We formalize this as Persistent Semantic Entities (PSEs): c...

📖 Read original article


296. EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference ​

Author: Yize Wu, Ke Gao, Ling Li, Yanjun Wu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07964v1 Announce Type: cross Abstract: Load Balancing has emerged as a critical problem in expert-parallel distributed inference of Mixture-of-Experts (MoE) models. As routing distributions are typically skewed across experts, devices hosting lighter-loaded experts must idle to wait for t...

📖 Read original article


297. Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions ​

Author: Chenrui Fan, Yize Cheng, Ming Li, Yongyuan Liang, Tianyi Zhou, Soheil Feizi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.07968v2 Announce Type: cross Abstract: Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency constraint, models must ...

📖 Read original article


298. Verication-driven closed-loop multi-agent large language modelframework for code-compliant structural design ​

Author: Jianbin Luo, Weibin Lin, Yiran Lin, Qing Wei, Wei Guo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.07978v1 Announce Type: cross Abstract: Multi-agent large language model(LLM)systems are applied to structural design,yet most use one-shot generation and cannot verify their output,leaving themill-suited to safety-critical tasks.Rather than trusting LLM self-correction,thisframework injec...

📖 Read original article


299. Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation ​

Author: Yichun Yeh, Yiheng Li, Xiaobo Hu, Zhen Lei, Yang Yang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08009v1 Announce Type: cross Abstract: Fake news increasingly relies on cross-modal image-text forgeries, making transparent and verifiable reasoning chains an urgent need for Detecting and Grounding Multi-Modal Media Manipulation (DGM4). Existing methods produce black-box detection resul...

📖 Read original article


300. Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models ​

Author: Jianqi Zhang, Xingyu Zhang, Zeen Song, Changwen Zheng, Fanjiang Xu, Wenwen Qiang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.08010v1 Announce Type: cross Abstract: Time series forecasting (TSF) plays an important role in a wide range of real-world applications. Recently, time series foundation models (TSFMs), pretrained on large-scale datasets, have demonstrated strong generalization capabilities and emerged as...

📖 Read original article


301. Evidence-RL: Towards Evidence-intensive Visual Reasoning ​

Author: Haojie Huang, Xinlei Yu, Chengming Xu, Zhangquan Chen, Cheng Yang, Qingdong He, Yu Yang, Jiangning Zhang, Xiaobin Hu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08021v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or...

📖 Read original article


302. Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States ​

Author: Zakhar Mrykhin, Valentin Malykh
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08024v1 Announce Type: cross Abstract: Large language models (LLMs) can generate fluent and useful responses but remain prone to hallucinations. We introduce Prompt Embedding Probes (PEP), a white-box method for answer-level hallucination detection from the hidden states of a frozen LLM. ...

📖 Read original article


303. DA-NBV: A Direction-Aware Next-Best-View Planner for Efficient 3D Reconstruction of Ships at Sea ​

Author: Jiaming Chen, Juntao Yang, Zhentao Zou, Qi Ming, Yi Yu, Zhihang Zhong, Xue Yang, Xue Jiang, Yue Zhou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.08025v1 Announce Type: cross Abstract: Accurate 3D reconstruction of ships at sea is important for maritime supervision, damage assessment, and autonomous maritime operations. Although 3D reconstruction has advanced considerably, high-quality data acquisition still largely relies on manua...

📖 Read original article


304. Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families ​

Author: Alizishaan Khatri, Dun Li Chan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2608.08029v1 Announce Type: cross Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-8B) detect harmful prompts at F1 competitive with guard models 1000x larger, using one probe per ben...

📖 Read original article


305. Tools to Explain Neural Networks for Power System Dynamics ​

Author: Petros Ellinas, Johanna Vorwerk, Spyros Chatzivasileiadis
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CE, cs.AI, cs.SY, eess.SY

arXiv:2608.08048v1 Announce Type: cross Abstract: This paper presents, for the first time in power systems literature to our knowledge, analytical tools to explain the training performance of machine learning surrogate models for power system dynamics. Power system simulations are increasingly chall...

📖 Read original article


306. DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects ​

Author: Yi Shu, Tianyu Peng, Yingzhuo Deng, Wen Yang, Jun Lin, Changming Xie, Xinyu Yu, Jiajun Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08067v1 Announce Type: cross Abstract: Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the semantic representati...

📖 Read original article


307. HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection ​

Author: Alireza Joonbakhsh (Shiraz University), Arda Canser Adal{\i} (Utrecht University), Slinger Jansen (Utrecht University), Farshad Khunjush (Shiraz University), Siamak Farshidi (Wageningen University,Research)
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.08069v1 Announce Type: cross Abstract: Foundation models are increasingly reused as software components, making model selection a critical software-engineering decision. Current model hubs primarily support discovery through popularity metrics, often neglecting functional capabilities, op...

📖 Read original article


308. Effect of Abstractions and Prompting Strategies on LLM-Guided High-Performance Optimizations ​

Author: Ji\v{r}'i Klepl, Maty'a\v{s} Brabec, Martin Kruli\v{s}
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2608.08085v1 Announce Type: cross Abstract: Code performance optimization is a vital aspect of modern software development, as it enables faster response times and reduced resource usage. These optimizations require a deep understanding of low-level hardware details and the intricacies of para...

📖 Read original article


309. Adaptive Symmetry Discovery for Dynamical System Identification ​

Author: Behrooz Tahmasebi, Melanie Weber
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.DS

arXiv:2608.08091v1 Announce Type: cross Abstract: Dynamical systems model trajectory data generated by fixed underlying dynamics, with applications ranging from biology to physics. Especially in scientific settings, dynamical systems are not generic but often exhibit symmetries imposed by physical l...

📖 Read original article


310. Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection ​

Author: Kaysarul Anas Apurba, Md. Hasibul Hasan, Mahedee Zaman Moon, Sk. Md. Mizanur Rahman, Atsuo Inomata
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2608.08100v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enables large language models to classify network flows and generate human-readable incident reports by retrieving semantically similar historical traffic from a vector knowledge base. However, the retrieval layer...

📖 Read original article


311. NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs ​

Author: Jiayue Jin, Jingwei Zhang, Chen Wang, Jing Liu, Longteng Guo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08107v1 Announce Type: cross Abstract: Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence acquired during pretraining. In this work, we investigate this phenomenon from the perspective of internal adapta...

📖 Read original article


312. Hierarchical Multi-Task Federated Learning in VANETs ​

Author: M. Saeid HaghighiFard, Sinem Coleri
Published: 8/11/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.DC, cs.LG, cs.NI, cs.SY

arXiv:2608.08111v1 Announce Type: cross Abstract: Vehicular Ad hoc Networks (VANETs) increasingly rely on federated learning (FL) to enable collaborative intelligence without sharing raw sensory data. However, most existing vehicular FL frameworks assume that all vehicles train a single global model...

📖 Read original article


313. Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario ​

Author: Satoshi Matsuoka
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.MA

arXiv:2608.08131v1 Announce Type: cross Abstract: In the fictional Order 66, catastrophe does not arise from a powerful command alone: a trusted population is preconditioned, a short directive activates the concealed condition, and protective authority turns against the system. This paper translates...

📖 Read original article


314. Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching ​

Author: Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model...

📖 Read original article


315. DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment and Multi-Task Learning in Oncology ​

Author: Junfei Ling (Institute of Medical Robotics, Shanghai Jiao Tong University), Bangzheng Pu (Institute of Medical Robotics, Shanghai Jiao Tong University), Bingsen Xue (Institute of Medical Robotics, Shanghai Jiao Tong University), Tianle Li (Institute of Data Science, The University of Hong Kong), Ruying Hu (Oriental Pan-Vascular Devices Innovation College, University of Shanghai for Science and Technology), Cheng Jin (Institute of Medical Robotics, Shanghai Jiao Tong University)
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.08148v1 Announce Type: cross Abstract: Attention mechanisms have been widely utilized in modern deep learning, and many existing multi-omics models inherit their conventional use to allow unrestricted bidirectional interactions. However, the fundamental logic of life is directional. Exist...

📖 Read original article


316. Exact Zarankiewicz Values On Two Finite Frontier Slices ​

Author: Koyar Afrasyab
Published: 8/11/2026, 4:00:00 AM
Categories: math.CO, cs.AI

arXiv:2608.08154v1 Announce Type: cross Abstract: The Zarankiewicz number Z(m,n,s,t) is the maximum number of edges in a bipartite graph with parts of orders m and n containing no copy of Ks,t. We give one combined, certificate-based computer-assisted proof for two finite slices and a corrected neig...

📖 Read original article


317. Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives ​

Author: Yingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam, Runnan Wang, Xuebo Liu, Yulong Chen, Yue Zhang, Derek F. Wong
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08160v1 Announce Type: cross Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaining long-horizon logi...

📖 Read original article


318. STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs ​

Author: Nuthakki Siva Gopala Krishna, Kanishka Jain
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08164v1 Announce Type: cross Abstract: Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller student model while signi...

📖 Read original article


319. $\texttt{DisMorph}$: learning to disentangle technical distortions from true biological change ​

Author: Jingru Fu, Kathleen E. Larson, Douglas N. Greve, Bruce Fischl, Malte Hoffmann
Published: 8/11/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, q-bio.QM

arXiv:2608.08173v1 Announce Type: cross Abstract: Longitudinal MRI enables sensitive measurement of structural brain change for studying aging and neurodegenerative disease. Deformable image registration is a key tool for estimating such change by computing a dense deformation that captures geometri...

📖 Read original article


320. A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization ​

Author: Praveen Kumar Katwe, Rakesh Chandra Balabantaray, Kali Prasad Vittala, Naman Kabadi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08180v1 Announce Type: cross Abstract: Abstractive text summarization systems frequently generate fluent yet unfaithful summaries by fabricating or distorting relationships between entities and events. Such relation-level hallucinations undermine the reliability of generated summaries, pa...

📖 Read original article


321. Biologically Informed Representation Learning for Robust Cross-Center Generalization of MALDI-TOF Mass Spectrometry ​

Author: Alejandro L. Garc'ia-Navarro, Carlos Sevilla-Salcedo, Bel'en Rodr'iguez-S'anchez, Vanessa G'omez-Verdejo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.08182v1 Announce Type: cross Abstract: Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial identification and antimicrobial resistance prediction. However, their deployment across institutions remains li...

📖 Read original article


322. Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification ​

Author: Yutong Wu, Xiaofan Bai, Shixin Li, Pingyi Hu, Ziqi Zhou, Zilong Wang, Xiaojing Ma, Songfeng Lu, Yuhong Li, Jin Xuan, Yi Wang, Dongmei Zhang, Bin Benjamin Zhu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.08195v1 Announce Type: cross Abstract: Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because deployed LLMs are commonly exposed only through query APIs, ownership verification must often rely o...

📖 Read original article


323. FreSH: Frequency-Segmented Hierarchical Multi-Expert Framework for Multivariate Time Series Classification ​

Author: Pingping Liu, Muyao Wang, Zijian Zhang, Tongshun Zhang, Hao Miao, Guorui Xie, Qingliang Li, Qiuzhan Zhou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.08207v1 Announce Type: cross Abstract: Multivariate Time Series Classification (MTSC) demands models that can effectively capture complex temporal patterns across multiple scales while remaining computationally efficient. However, existing approaches generally struggle to reconcile fine-g...

📖 Read original article


324. VTO: Visual Tool Orchestration for Video Anomaly Detection ​

Author: Rui Wang, Yeteng Wu, Xianling Zhang, Mengshi Qi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08219v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learning approaches are fundamentally limited by poor generalization across diverse scenarios. While mult...

📖 Read original article


325. Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via Proxy Representations ​

Author: Michael Levit, Josh Ledgard, Haoyu Dong, Vishwas Suryanarayanan, Eyal Kolman, Sharon Tan, Qiang Gan, Vishal Chowdhary
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.HC

arXiv:2608.08245v1 Announce Type: cross Abstract: LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making it difficult to obtain any representative evaluation dataset or to track the ongoing evolution of production t...

📖 Read original article


326. On the Robustness of LLMs' Internal Representation of Code Correctness ​

Author: Francisco Ribeiro, Sohaila Abdulsattar, Renata Gonzalez, Mahmoud Kassem, Sarah Nadi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2608.08266v1 Announce Type: cross Abstract: Code generated by modern language models often reads naturally. Yet, it also often fails to implement what was asked. This should be no surprise, as research shows the models' own confidence signals are poorly calibrated with actual correctness. A pr...

📖 Read original article


327. Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations? ​

Author: Osvaldo Quinjica, Eric Bennett, Xinchen Yang, Andrew Schonebaum, Marine Carpuat
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08283v1 Announce Type: cross Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics dev...

📖 Read original article


328. Frequency-Domain Dual-Branch Fusion for Medical Visual Question Answering ​

Author: Yusra Tariq, Rakesh Chandra Joshi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08307v1 Announce Type: cross Abstract: Medical Visual Question Answering (VQA) requires aligning subtle visual evidence, including lesion texture, boundary sharpness, and diffuse density changes, with clinical language. Existing multimodal fusion approaches operating in the spatial domain...

📖 Read original article


329. Open-World Semantic Segmentation with Sensitivity Modeling ​

Author: Anastasios Romanos Varvarigos, Nikos Giakoumoglou, Tania Stathaki
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08308v1 Announce Type: cross Abstract: Modern vision systems must operate in "open-world" settings, where models must recognize known categories and detect unseen or anomalous content. Conventional semantic segmentation models operate under a "closed-world" assumption, often producing ove...

📖 Read original article


330. Three Necessary Principles for Self-Supervised Visual Representation Learning ​

Author: Nikos Giakoumoglou, Paschalis Giakoumoglou, Tania Stathaki
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.08309v1 Announce Type: cross Abstract: We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational non-deg...

📖 Read original article


331. Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution ​

Author: Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev, Nikita Dragunov, Roman Yampolskiy, Andrei Kuznetsov
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.08311v2 Announce Type: cross Abstract: We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evo...

📖 Read original article


332. PRISM: A Predictive Protocol for Permutation Optimization via Landscape Diagnostics ​

Author: Blessings Mambwe
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC

arXiv:2608.08344v1 Announce Type: cross Abstract: Permutation optimization arises whenever the components of a system are fixed but their ordering affects performance. We introduce PRISM, a predictive protocol for permutation optimization that measures a fitness landscape before selecting a search s...

📖 Read original article


333. Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books ​

Author: Karim Benharrak, Oriol Nieto, Bryan Wang, Zeyu Jin, Amy Pavel
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.MM

arXiv:2608.08349v1 Announce Type: cross Abstract: Audio dramas weave dialogue, sound effects, and music into immersive stories. Creators often adapt books into audio dramas, but this process remains labor-intensive, requiring them to interpret source material, author scripts, generate audio assets, ...

📖 Read original article


334. From Product Search to Preference Articulation: The Economics of Agentic Commerce ​

Author: Lingxiu Dong, Kaiwen Luo, Fasheng Xu
Published: 8/11/2026, 4:00:00 AM
Categories: econ.TH, cs.AI, cs.GT

arXiv:2608.08395v1 Announce Type: cross Abstract: Generative AI is shifting digital commerce from browsing toward agentic search, in which consumers delegate product discovery to AI agents. We compare manual search, which accurately evaluates a limited product set, with agentic search, which screens...

📖 Read original article


335. Does a Toehold Make a Bidder Bolder? Preemption and Multiplicity in Multi-Round Takeover Auctions ​

Author: Zain Naboulsi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.LG

arXiv:2608.08407v1 Announce Type: cross Abstract: A bidder can quietly buy a stake in a company before making an offer for it. That stake, a toehold, is supposed to pay for itself twice: it makes the bidder willing to bid harder, and it frightens rivals into staying out of the fight. The first effec...

📖 Read original article


336. Abstracted Away: Resisting Alienation and Ungrounded Abstraction in AI Research Communities ​

Author: Vyoma Raman, Isabel O. Gallegos, Neha Srivathsa
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2608.08408v1 Announce Type: cross Abstract: Logics of abstraction in computational AI research often push important forms of knowledge and reflection aside: dominant standards of legitimacy separate from lived experience of harm; the goals of work misalign with the practices that operationaliz...

📖 Read original article


337. Human-Guided Causal Knowledge Injection for Virtual Cells ​

Author: Pengcheng Wang, Changjian Chen, Zhuo Tang, You Wu, Long Wang, Feng Yu, Kenli Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.08430v1 Announce Type: cross Abstract: Virtual cells employ machine learning models to simulate and predict cellular behaviors, serving as a critical computational framework for investigating health and disease. Injecting causal graphs into virtual cells can improve the interpretability, ...

📖 Read original article


338. FSTC-Encoder: Feature--Spatial--Temporal Correlation Learning for Generalizable RF Sensing ​

Author: Jing Wang, Zhu Wang, Changlong Cheng, Yifan Guo, Yin Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.08439v1 Announce Type: cross Abstract: Heterogeneous RF sensing differs substantially in feature structure, spatial layout, and temporal scale, making existing models difficult to reuse across devices, environments, and RF modalities. We propose FSTC-Encoder, which unifies heterogeneous R...

📖 Read original article


339. Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction ​

Author: Miki Ueno
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.08443v1 Announce Type: cross Abstract: Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other parts of a relational microculture. Recent work has also examined how humans and conversational AI negotiate an...

📖 Read original article


340. Hidden Language Consistency Phenomena in Reasoning LLMs ​

Author: Muhammad Ali Shafique, Kelly Marchisio
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08447v1 Announce Type: cross Abstract: Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended language while reasoning and responding. This omission conceals important multilingual behaviors that eme...

📖 Read original article


341. Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization ​

Author: Haojie Yu, Ziyou Jiang, Junjie Wang, Mingyang Li, Yuekai Huang, Jie Huang, Qing Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08451v1 Announce Type: cross Abstract: Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reasoning Chain (ORC) of recurring topics, harm language indicators, severity hierarchies, and type char...

📖 Read original article


342. Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction ​

Author: Zhuowen Liang, Zhengxuan Zhang, Jiayang Wang, Jiazhuo Chen, Nan Tang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DB

arXiv:2608.08459v1 Announce Type: cross Abstract: Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domains such as finance, healthcare, education, transportation, and enterprise operations, downstream work...

📖 Read original article


343. Halpern Iteration Achieves $\tilde{\mathcal{O}}(\epsilon^{-1/p})$ $p$th-Order Oracle Complexity for Monotone Variational Inequalities ​

Author: Lesi Chen, Xinliang Zhang, Hengyu Wang, Chengchang Liu, Yongchao Chen, Jingzhao Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: math.OC, cs.AI

arXiv:2608.08463v1 Announce Type: cross Abstract: We study second- and higher-order methods for solving smooth monotone variational inequalities (MVI). Monteiro and Svaiter (SIAM J. Optim., 2012) showed that a second-order method, NPE, converges at the rate of $\mathcal{O}(T^{-1.5})$. For convex-con...

📖 Read original article


344. SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills ​

Author: Xinze Chen, Chi Zhang, Ping Ji, Yimin Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.08468v1 Announce Type: cross Abstract: Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their security properties remain under-explored. We present \textsc{SkillsMetric}, a five-stage static analysis framework th...

📖 Read original article


345. SuperNeuroMAT: An Efficient Matrix-based Simulator for Spiking Neural Networks ​

Author: Prasanna Date, Kevin Zhu, Shruti Kulkarni, Ashish Gautam, Chathika Gunaratne, Robert Patton, Tyler Nitzsche, Ian Mulet, Zachary Johnson-Scott, Addison Helms, Duncan Rowden, Simon Weston, Maryam Parsa, Catherine Schuman, Thomas Potok
Published: 8/11/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.CE, cs.ET, cs.LG

arXiv:2608.08479v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) offer a promising pathway to energy-efficient AI and brain-inspired computing. However, their widespread adoption is hindered by a lack of fast, accessible, and versatile simulation frameworks. In this paper, we introdu...

📖 Read original article


346. A Combined Feature-Based Framework for Disguise and Spoofing Detection in Face Recognition Systems ​

Author: Sangiya Pararajasingham
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CR

arXiv:2608.08521v2 Announce Type: cross Abstract: Face recognition systems face two distinct, commonly-separated failure modes: spoofing, where an impostor presents a photograph or video of an authorized user, and disguise, where a legitimate user is rejected because their appearance differs from th...

📖 Read original article


347. On-Device Multi-Species Malaria Detection with Uncertainty-Calibrated Slide-Level Aggregation ​

Author: Idaya Seidu, Ahmed Tahiru Issah, Charles B. Delahunt, Carine Mukamakuza
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08566v1 Announce Type: cross Abstract: Malaria remains a leading cause of mortality in resource-limited settings, where expert microscopists are scarce. Automated diagnosis based on microscopy images thus has strong potential to improve care delivery. But for an algorithm to deploy, a nec...

📖 Read original article


348. CDGC-Net: 3D Medical Image Segmentation with Cooperative Dual-Scale Self-Attention and Grouped Channel Modeling ​

Author: Zheyang Jing, Qin Lu, Jianwang Li, Yujie Yang, Chen Yi, Shaofeng Jiang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08575v1 Announce Type: cross Abstract: Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail. Existing methods often model global and local features in separate modules or feature levels and perform channel recalibration...

📖 Read original article


349. Population-Scalable Multi-Agent World Modeling ​

Author: Renjie Zhao, Yuxiang Wu, Mingyu Zhang, Jiaxin Li, Sisi Li, Yimin Sheng, Tianxi Tan, Zhenkai Zhang, Jianyi Zhu, Yong-Lu Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.08600v1 Announce Type: cross Abstract: World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent environments introduces a fundamental scalability challenge. Existing methods generally assume a fixed number o...

📖 Read original article


350. Mitigating Gender Bias in English to Romanian Machine Translation ​

Author: Ioana Grigore, Sergiu Nisioi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08606v1 Announce Type: cross Abstract: Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-neutral language like English to a gendered target language such as Romanian. This bias results in translations that default to mascul...

📖 Read original article


351. REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering ​

Author: Caijun Yan, Yang Zhou, Meixing Shi, Haoran Sun, Yichen Li, Yuxiang Cai, Yankai Jiang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08612v1 Announce Type: cross Abstract: Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering. However, existing methods typically rely on rigid, fixed-length temporal chunking (e.g., 10s) and static offline mem...

📖 Read original article


352. RAG-Based Auto-Configuration for Industrial Fieldbus Devices ​

Author: Aadil Gani Ganie, Saad Ezzini, Naveed Farooz Marazi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CL

arXiv:2608.08618v1 Announce Type: cross Abstract: Industrial device commissioning requires engineers to manually extract hundreds of protocol-specific parameters from heterogeneous PDF manuals and transcribe them into supervisory control systems, a time-intensive, error-prone workflow. This paper pr...

📖 Read original article


353. Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach ​

Author: Tong Bao, Yi Zhao, Heng Zhang, Chengzhi Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.DL, cs.IR

arXiv:2608.08636v1 Announce Type: cross Abstract: Scientific named entity recognition (SciNER) plays a crucial role in information extraction and knowledge discovery from scientific texts. Recently, large language models (LLMs) have demonstrated the capacity to achieve competitive SciNER performance...

📖 Read original article


354. CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents ​

Author: Yuqian Zhang, Yao Shi, Kexin Huang, Botian Jiang, Zhe Xu, Yiwei Zhao, Min Liang, Shuang Chen, Xipeng Qiu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL

arXiv:2608.08638v1 Announce Type: cross Abstract: Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools. All TTS systems require faithful linguistic rendering, consistent speaker identity, and low-latency response. Yet compact streaming syste...

📖 Read original article


355. LegoLM: Structured Weight Sharing for Large Language Models ​

Author: Joseph Bingham
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.08652v1 Announce Type: cross Abstract: We present \LegoLM{}, a structured weight-sharing compression framework for large language models grounded in a systematic study of why global weight sharing fails and how to fix it. We identify two distinct failure modes. Distributional mismatch: fo...

📖 Read original article


356. UniSpace: Unified Visual Representation and Scalable Multimodal Modeling ​

Author: Jinbo Yan, Limeng Qiao, Jie Qin, Junyan He, Feize Wu, Guanglu Wan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08676v1 Announce Type: cross Abstract: Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine-grained visual details, leading to poor pixel reconstruction and limi...

📖 Read original article


357. RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation ​

Author: Dongjie Xu, Kai Qian, Julius, Weijie Shi, Yuxuan Sun, Minghua Tang, Fenglei Jin, Hanchi Dong, Jiajie Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.08684v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. Existing methods rely on proxies such as layer depth, attention statistics, or representation change. These prox...

📖 Read original article


358. Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation ​

Author: Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jeremias Traub, Tassilo Wald, Constantin Ulrich, Klaus Maier-Hein
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08713v1 Announce Type: cross Abstract: Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges. Modern foundation vision encoders (VEs) can produce tens of thousands of ...

📖 Read original article


359. LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization ​

Author: Zexun Lin, Yuan Feng, Junlin Lv, Kevin S. Zhou, Xike Xie
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08721v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model inference by drafting multiple tokens for parallel verification, with efficiency critically determined by the speculative length selected at each decoding round. Existing dynamic speculation metho...

📖 Read original article


360. Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure ​

Author: V'ictor Gallego
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.08722v1 Announce Type: cross Abstract: Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely in two GPU-kernel-optimization suites with held-out generalization gates: Metal-Sci (10 scientific-c...

📖 Read original article


361. PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation ​

Author: Yangyang Feng, Zhuoyan Feng, Junlan Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.08726v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts. Yet each rollout also reveals how the student's response unfolds and whether it succeeds, student-specific hindsigh...

📖 Read original article


362. TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases ​

Author: Gia-Han Truong, Khang Nguyen Quoc, Luyl-Da Quach
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08727v1 Announce Type: cross Abstract: To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding. TomaMMU comprises 28,808 high-quality images spanning...

📖 Read original article


363. Can We Optimize the Performance-Carbon Emission Break-Even Point?: The Quest for Greener LLMs ​

Author: Sourav Das, Tanmay Joshi, Kripabandhu Ghosh
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CE, cs.ET, cs.LG

arXiv:2608.08744v1 Announce Type: cross Abstract: The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially exceeds the one-time cost of fine-tuning. Yet most efficiency interventions target either pre-training scale o...

📖 Read original article


364. Eco-SoC: A Sustainable VLSI Architecture for Energy-Proportional Artificial Intelligence ​

Author: Jatin Chopra
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AR, cs.AI

arXiv:2608.08761v1 Announce Type: cross Abstract: In an era defined by escalating climate change and the pervasive deployment of edge intelligence, the environmental cost of semiconductor manufacturing and operation has reached a critical threshold. As Deep Learning (DL) accelerators dominate System...

📖 Read original article


365. Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast ​

Author: Jiaxin Guo, Yanwei Yue, Xuanbo Fan, Chunyu Yang, Yan Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.08764v1 Announce Type: cross Abstract: On-policy self-distillation improves language-model reasoning by querying a teacher on states actually visited by the student. Recent methods create a powerful information asymmetry by exposing the teacher to privileged context, yet they fundamentall...

📖 Read original article


366. 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents ​

Author: Kenta Watanabe, Atsuyuki Miyai, Mizuki Takenawa, Kiyoharu Aizawa, Toshihiko Yamasaki
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.08814v1 Announce Type: cross Abstract: We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or compl...

📖 Read original article


367. Hybrid Neural-Classical Correction for Frozen Time Series Foundation Models: A Comprehensive Ablation Study on High-Frequency Stock Prediction ​

Author: Kasun Dewage, Suranadi De Silva, Shankhadeep Mondal
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-fin.ST

arXiv:2608.08825v1 Announce Type: cross Abstract: Foundation models for time series forecasting demonstrate impressive zero-shot generalization but often underperform on specialized domains such as high-frequency finance. We present a comprehensive study of hybrid neural-classical correction for ada...

📖 Read original article


368. Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models ​

Author: Muhammad Faishal Adly Nelwan, Alfan Farizki Wicaksono
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08829v1 Announce Type: cross Abstract: Activation steering edits the behaviour of a frozen language model by adding a learned vector to its residual stream, and current practice fixes the injection layers globally per task. We argue that the best layers are an instance-level decision, and...

📖 Read original article


369. Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data ​

Author: Anushka Roy, Jyotirmoy Singh, Shreea Bose, Chittaranjan Hota
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.08859v1 Announce Type: cross Abstract: Wireless Body Area Networks (WBANs) generate multivariate physiological time series that are highly nonstationary and must often be processed under strict computational and memory constraints. A critical yet underexplored challenge in this setting is...

📖 Read original article


370. DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference ​

Author: Asaad Althoubi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.PF

arXiv:2608.08878v1 Announce Type: cross Abstract: Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequence length, creating a severe memory bottleneck for long-context inference. Existing heuristic evicti...

📖 Read original article


371. Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol ​

Author: Christoph Trattner
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.08882v1 Announce Type: cross Abstract: AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after using such a tool, what can the user still do on their own? I call this epistemic transfer. It refers to the ef...

📖 Read original article


372. A New Approach to Characterising Optimisation Problems Using Programmatic Representation and Complexity Measures ​

Author: Marcus Gallagher, Katherine M. Malan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2608.08898v1 Announce Type: cross Abstract: Characterising optimisation problem instances is a fundamental part of understanding the behaviour and performance of different algorithms as well as providing information for algorithm selection and configuration. In this paper we propose a novel ap...

📖 Read original article


373. From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability ​

Author: Alexander Hackett, Arnaud Denis-Remillard, Axel Cassou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.08904v1 Announce Type: cross Abstract: How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric understanding, from every dec...

📖 Read original article


374. ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision ​

Author: Delin Mao, Chenghao Sun, Jingwei Song, Chishui Chen, Linfeng Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08907v1 Announce Type: cross Abstract: Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then-RL recipe creates a different supervision misalignment at each stage. SFT is expected to teach how ...

📖 Read original article


375. Toward CT-Equivalent Image Quality in Low-Dose Radiotherapy Planning: Conditional Diffusion-Based CBCT-to-CT Synthesis and the Impact of CBCT Input Representation ​

Author: Alzahra Altalib, Chunhui Li, Christopher Hamill Taylor, Sankar Pillai, Alessandro Perelli
Published: 8/11/2026, 4:00:00 AM
Categories: physics.med-ph, cs.AI

arXiv:2608.08919v1 Announce Type: cross Abstract: During standard radiotherapy planning, repeated CT acquisitions are often required for patient registration, verification, and adaptive planning, resulting in increased cumulative X-ray dose. To mitigate this, low-dose cone-beam CT (CBCT) is routinel...

📖 Read original article


376. From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving ​

Author: Chaitanya Shinde, Hadi Hajieghrary, Miguel Hurtado
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.SY, eess.SY

arXiv:2608.08941v1 Announce Type: cross Abstract: Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe what the ADS must demonstrably do once deployed within that domain. This gap between operating conditio...

📖 Read original article


377. Do AI Forecast Ensembles Sample the Correct Conditional Distribution? ​

Author: Lucas J. Howard, Elizabeth A. Barnes
Published: 8/11/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.AI, cs.LG, stat.AP, stat.ML

arXiv:2608.08954v1 Announce Type: cross Abstract: Ensemble forecasting aims to sample the conditional distribution of outcomes; whether AI forecast ensembles do this correctly in a joint sense remains largely untested. We train a diffusion model for probabilistic subseasonal coastal sea level foreca...

📖 Read original article


378. Idea Search: Guiding Tree Search with Ideas to Explore Diverse Scientific Methods ​

Author: Xuefei Julie Wang, Hao Cui, Michael P. Brenner, Subhashini Venugopalan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.GN, q-bio.QM

arXiv:2608.08958v1 Announce Type: cross Abstract: Tree Search-based test-time scaling of LLMs is a powerful tool for automated scientific coding. However, pure Tree Search sometimes struggles with systematic exploration, becoming trapped in local optima, or unproductive loops, especially in the vast...

📖 Read original article


379. Fourier Self-Supervision for Fine-Grained Generalized Category Discovery ​

Author: Sarah Rastegar, Mina Ghadimi Atigh, Pascal Mettes, Yuki M. Asano, Cees G. M. Snoek
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.08963v1 Announce Type: cross Abstract: Generalized Category Discovery aims to recognize known categories while identifying novel ones within unlabeled data. Existing methods, typically based on self-supervision and contrastive learning, often struggle to capture fine-grained distinctions,...

📖 Read original article


380. GALA: Graph-Augmented LLM Agents for Root Cause Analysis and Incident Response in Microservices ​

Author: Yifang Tian, Yaming Liu, Zichun Chong, Zihang Huang, Yiran Li, Hans-Arno Jacobsen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.08968v1 Announce Type: cross Abstract: Microservice root cause analysis (RCA) requires correlating failures across heterogeneous telemetry within complex service dependency graphs. Existing methods often rely on a single telemetry modality; recent LLM-based approaches can suffer from unco...

📖 Read original article


381. How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review ​

Author: Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08975v1 Announce Type: cross Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary a...

📖 Read original article


382. Detecting Clear Contact Lenses for Iris Recognition: A Two-Stage Mask-Guided Attention Approach ​

Author: Parisa Farmanifard, Arun Ross
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08977v1 Announce Type: cross Abstract: This work focuses on the impact and detection of clear contact lenses in the context of iris recognition. While the detection of cosmetic or patterned contact lenses has been extensively studied under the presentation attack detection (PAD) paradigm,...

📖 Read original article


383. How Far Do Foundation Models Transfer to Infant Signals? A Cross-Dataset Transfer Audit with a Unified Need Ontology ​

Author: Wu Hangyu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.08989v1 Announce Type: cross Abstract: Public infant cry corpora are small, label-incompatible, and almost always evaluated one corpus at a time. We ask what this practice hides and what fixes it. Across four cry corpora screened by a multi-level leakage audit (byte-level and embedding-le...

📖 Read original article


384. Guardian Crawler: Retrieval-First Knowledge Discovery with Bounded LLM Augmentation for Noisy Web Intelligence ​

Author: Joshua Castillo, Santosh Nukavarapu, Ravi Mukkamala
Published: 8/11/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL, cs.LG

arXiv:2608.08994v1 Announce Type: cross Abstract: Retrieving relevant evidence from noisy web data is challenging, particularly in sensitive domains containing incomplete reports, heterogeneous language, and irrelevant content. We present Guardian Crawler, a reproducible retrieval-first testbed for ...

📖 Read original article


385. Multi-agent discovery of practical quantum LDPC codes ​

Author: Dongheng Qian, Tianyi Li
Published: 8/11/2026, 4:00:00 AM
Categories: quant-ph, cs.AI

arXiv:2608.08996v1 Announce Type: cross Abstract: Quantum low-density parity-check (qLDPC) codes can encode multiple logical qubits using sparse parity checks, yet searching for useful finite-length instances remains a challenging design problem because code performance must be optimized while satis...

📖 Read original article


386. DeepFreqMark: End-To-End Learnable Frequency-Domain Watermarking with Spherical Attack Simulation for Latent Diffusion Models ​

Author: Chen-Hsiu Huang, Mario K"oppen, Ja-Ling Wu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.08999v1 Announce Type: cross Abstract: The proliferation of AI-generated images produced by Latent Diffusion Models (LDMs) has raised critical concerns regarding copyright infringement and misinformation. Although existing frequency-domain watermarking methods embed handcrafted geometric ...

📖 Read original article


387. SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs ​

Author: Shiwei Gan, Xiao Liu, Yafeng Yin, Zhiwei Jiang, Bowen Guo, Lie Xie, Sanglu Lu, Hongkai Wen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09006v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Translation (GFSLT) remains a challenge. In this paper, we investigate how to effectively adapt LLMs to...

📖 Read original article


388. How People Evaluate AI-, Expert-, and Peer-Style Financial Advice ​

Author: Aryan Ramchandra Kapadia, Eshwar Chandrasekharan, Koustuv Saha
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CL, cs.CY, cs.SI

arXiv:2608.09019v1 Announce Type: cross Abstract: As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice. We conducted a preregistered vignette experiment (N = 285) in...

📖 Read original article


389. MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation ​

Author: Shuyu Li, Kejun Zhang, Jiahe Lei, Shulei Ji, Zihao Wang, Jiaxing Yu, Wanying Wu, Lei Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.MM

arXiv:2608.09035v1 Announce Type: cross Abstract: Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation. To a...

📖 Read original article


390. Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization ​

Author: Hyangsuk Min, Hwanjun Song
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.09043v1 Announce Type: cross Abstract: Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context to be interpreted on its own. We formalize this setting as streaming dialogue summarization, where a system must summarize a current...

📖 Read original article


391. Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production ​

Author: Xiao Liu, Shiwei Gan, Yafeng Yin, Jiaxin Yin, Bowen Guo, Yaqi Sun, Zhiwei Jiang, Lei Xie
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.MM

arXiv:2608.09045v1 Announce Type: cross Abstract: Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolated sign language recognition (ISLR), continuous sign language recognition (CSLR), and sign language t...

📖 Read original article


392. Triple Expert Learning from Noisy Labels for Semi-Supervised Vision Foundation Model Adaptation ​

Author: Xuanyu Liu, Zheng Fang, Hongyang He, Yundi Hong, Daizong Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09052v1 Announce Type: cross Abstract: Semi-supervised adaptation of vision foundation models (VFMs) commonly freezes the pretrained backbone and updates lightweight modules such as LoRA. However, pseudo-labels have mixed reliability, and a single LoRA adapter must absorb reliable, ambigu...

📖 Read original article


393. Two-Step MV-DeepONet: Probabilistic Operator Learning for Uncertainty Propagation Driven by Random Input Fields ​

Author: Yupei Nie, Lei Wang, Jiasen Liu
Published: 8/11/2026, 4:00:00 AM
Categories: math.NA, cs.AI, cs.NA

arXiv:2608.09071v1 Announce Type: cross Abstract: Forward uncertainty propagation in complex physical systems can induce structured covariance across field-valued outputs. For a probabilistic surrogate, the total predictive covariance comprises the covariance of conditional means across input realiz...

📖 Read original article


394. A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents ​

Author: Xin Zhou, Chun Yong Chong, Kisub Kim, Yun Peng, Rui Shu, Zihan Wu, Xu Han, Guowen Yuan, Zeyang Zhuang, Jounghoon Kim, Jeongjin Ju, Seongmin Ju, Taein Yoon, David Lo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs. Yet existing repository-level benchmarks typically evaluate only whether the final patch passes tes...

📖 Read original article


395. When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information ​

Author: Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam, Quan Z. Sheng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC, cs.LG

arXiv:2608.09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability under uncertainty remains poorly understood which raises critical concerns for deployment in high-sta...

📖 Read original article


396. TLDChoiceNet: Quantitatively Choosing a Transfer Learning Dataset ​

Author: Jing Ning, James D. Braza
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09091v1 Announce Type: cross Abstract: Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer learn upon massive datasets like ImageNet , CIFAR-100, or COCO . Qualitatively, it seems a transfer learning dat...

📖 Read original article


397. The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora ​

Author: E. M. Freeburg
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.09093v1 Announce Type: cross Abstract: How a document's arrangement is written down, its notation, is a training variable that no dataset card records. The field has established that text-extraction choices change model behaviour, and has never once measured the notation of what those cho...

📖 Read original article


398. Visual Distortion Detection in UGC Images Using Large Multimodal Models ​

Author: Ziheng Jia, Yingji Liang, Jiaying Qian, Xiongkuo Min
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09122v1 Announce Type: cross Abstract: The localized depiction of perceptual quality has long been a crucial, yet underexplored, challenge in image quality assessment (IQA). Existing approaches based on large multimodal models (LMMs) predominantly rely on text-driven supervised fine-tunin...

📖 Read original article


399. Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments ​

Author: Keyu He, Xuhui Zhou, Maarten Sap
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.MA

arXiv:2608.09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social interaction offers no objec...

📖 Read original article


400. MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning ​

Author: Hanye Zhao, Muning Wen, Yong Yu, Weinan Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09130v1 Announce Type: cross Abstract: Allocating limited computation among concurrent learning tasks is difficult when each task must reach a target loss before a deadline but its required training effort is unknown. Existing approaches combine online loss prediction with adaptive resour...

📖 Read original article


401. When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution ​

Author: Yu Shi, Yuyao Zhang, Yu-wing Tai
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09133v1 Announce Type: cross Abstract: Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to the LR observation remains challenging. In particular, we observe that diffusion transformers (DiTs) built on la...

📖 Read original article


402. SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning ​

Author: David D. Yuan, Tony Z. Zhao, Kaylee Burns, Chelsea Finn
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.09138v2 Announce Type: cross Abstract: While learned robotic policies hold promise for advancing generalizable manipulation, their practical deployment is often hindered by suboptimal execution speeds. Imitation learning policies are inherently limited by hardware constraints and the spee...

📖 Read original article


403. From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs ​

Author: Yuanhe Zhang, Weiliu Wang, Jie Ren, Liang Lin, Zhenhong Zhou, Haoran Gao, Kun Wang, Chen Li, Li Sun, Sen Su
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2608.09158v1 Announce Type: cross Abstract: Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. Ho...

📖 Read original article


404. Tabular Numeric Stretch Transformation ​

Author: Zihao Ye, Juyong Kim, Johnna Sundberg, Burak Varici, Pradeep Ravikumar
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09162v1 Announce Type: cross Abstract: Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distributions, scales, and statistical properties. Although recent advances have improved how models learn from tabular ...

📖 Read original article


405. Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression ​

Author: Jingbo Wen, Liang He, Mingyu Cao, Haoyu Wang, Minxuan Hu, Kangning Cui, Xilu Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09176v1 Announce Type: cross Abstract: Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed compute budget, implicitly assuming that all errors carry equal cost....

📖 Read original article


406. Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction ​

Author: Jingxian Xu, Yuhao Huang, Rusi Chen, Yanfeng Zhou, Dong Ni
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2608.09182v1 Announce Type: cross Abstract: Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Existing localization methods have advanced, among which multi-stage refinement is a superior solution. Although thi...

📖 Read original article


407. SiriusDeliver: Automating Data Warehouse Delivery at Tencent ​

Author: Haining Xie, Xiaokai Zhou, Jiaming Yang, Siqi Shen, Ziwei Wang, Yifeng Zheng, Tengyue Xu, Yipeng Shi, Zefang Zong, Yang Li, Peng Chen, Jie Jiang, Debiao He, Xiao Yan, Jiawei Jiang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.SE

arXiv:2608.09185v1 Announce Type: cross Abstract: Enterprise data warehouses (DWs) support business-critical analytics, but warehouse task delivery remains a complicated production process involving context retrieval, workflow configuration, code generation, platform submission, and failure diagnosi...

📖 Read original article


408. FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learning ​

Author: Van Truong Vo, Khoa Nguyen, Taehong Kim
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, cs.DC

arXiv:2608.09208v1 Announce Type: cross Abstract: Decentralized intelligence systems with heterogeneous devices and limited coordination increasingly rely on decentralized federated learning (DFL). However, DFL suffers from convergence inefficiency under data heterogeneity due to the use of a unifor...

📖 Read original article


409. AkasicDB: Demonstrating Omni RAG with a Unified Vector-Graph-Relational DBMS ​

Author: Geonho Lee, Jeongho Park, Donghyoung Han, Min-Soo Kim
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2608.09214v1 Announce Type: cross Abstract: Recent Retrieval-Augmented Generation (RAG) systems increasingly combine vector retrieval with structured knowledge, such as Graph RAG and Filtered vector search. However, existing database architectures struggle to support such complex RAG workflows...

📖 Read original article


410. Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training ​

Author: Ting Zhou, Zhenqing Ling, Daoyuan Chen, Qianli Shen, Yilun Huang, Ying Shen, Yaliang Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09217v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization. Exi...

📖 Read original article


411. FedTVD: Balancing Data Quality and Quantity for Robust Federated Learning ​

Author: Radwan Selo, Majid Kundroo, Taehong Kim
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, cs.DC

arXiv:2608.09221v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative model training across distributed client devices while preserving data privacy. However, FL faces significant challenges due to data heterogeneity, particularly in terms of label distribution skewness and...

📖 Read original article


412. Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference ​

Author: Tejasvi C. Addagada
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.09225v1 Announce Type: cross Abstract: The key-value (KV) cache is the primary throughput optimization in modern large language model (LLM) inference, enabling prefix reuse across requests. In multi-tenant deployments this cache is shared across tenants, creating a timing side channel: an...

📖 Read original article


413. RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation ​

Author: Yuhan Li, Fangao Zeng, Sicong Kang, Mengfei Xu, Hao Zhou, Wei Li, Pipei Huang, Bingbing Ni
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09226v1 Announce Type: cross Abstract: Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains du...

📖 Read original article


414. Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation ​

Author: Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar, Junpei Komiyama
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09228v1 Announce Type: cross Abstract: On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student's trajectory. However, this interpretation conflates two e...

📖 Read original article


415. Multimodal Federated Learning under Dual-Axis Modality Missingness ​

Author: Adiba Orzikulova, Jaehyun Kwak, Jaemin Shin, Yunqi Guo, Xiaomin Ouyang, Guoliang Xing, Steven Euijong Whang, Sung-Ju Lee
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09240v1 Announce Type: cross Abstract: Multimodal federated learning (FL) supports collaborative modeling in privacy-sensitive health-sensing and medical settings, but realistic deployments often exhibit dual-axis modality missingness: clients have different modality sets, and individual ...

📖 Read original article


416. MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts ​

Author: Peiwen Li, Shiyang Zhang, Yangtian Zhang, Sizhuang He, David van Dijk, Rex Ying
Published: 8/11/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CL, cs.LG

arXiv:2608.09251v1 Announce Type: cross Abstract: Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptation for diverse subtasks, resul...

📖 Read original article


417. SafeQL: Search-based Refinement for Safe and Efficient LLM-based Text-to-SQL ​

Author: Geonho Lee, Min-Soo Kim
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2608.09260v1 Announce Type: cross Abstract: Large language models (LLMs) have advanced Text-to-SQL by enabling natural language interfaces to databases without task-specific fine-tuning. However, existing LLM-based systems remain unreliable, often generating SQL queries that are invalid under ...

📖 Read original article


418. Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations ​

Author: Weijie Liang, Yuanfeng Song, Xing Chen, Caleb Chen Cao, Sirui Han, Yike Guo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2608.09268v1 Announce Type: cross Abstract: Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understanding. We study whether this representation can serve as operational context for agentic coding, where an agent...

📖 Read original article


419. GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views ​

Author: Jiahui Cui, Yan Zhao, Kan Wei, Enze Zhu, Peirong Zhang, Lei Wang, Yiru Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.IR, cs.MM

arXiv:2608.09270v1 Announce Type: cross Abstract: Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field of view and overhead perspective of drone scenarios impose dual challenges on vision-language understanding. At...

📖 Read original article


420. SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation ​

Author: Jefferson Hernandez, Jaywon Koo, Zilin Xiao, Chen Wei, Vicente Ordonez
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09271v1 Announce Type: cross Abstract: Group-based reinforcement learning objectives such as GRPO can allocate learning signal poorly across prompt difficulty: under binary rewards, group normalization induces a divergent weighting on easy prompts. We introduce Softmax Advantage Group Est...

📖 Read original article


421. Software Engineering for and with GUI Agent ​

Author: Shengcheng Yu, Yuchen Ling, Junyang Xing, Quan Zhou, Chunrong Fang, Zhenyu Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2608.09278v1 Announce Type: cross Abstract: GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity of the field. GUI agents remain technically brittle, incompletely engineered, and insufficiently va...

📖 Read original article


422. GLocFM: A Geometry-Aware Foundation Model for 3D Indoor Wireless Localization ​

Author: Chenghong Bian, Chaozheng Wen, Hongze Chen, Jun Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: eess.SP, cs.AI

arXiv:2608.09285v1 Announce Type: cross Abstract: Learning-based wireless localizers often fail to utilize geometric information about the propagation environment, limiting their ability to exploit non-line-of-sight (NLoS) propagation and generalize across scenes. To bridge this gap, we propose GLoc...

📖 Read original article


423. VeinCast: Physics-Guided Dynamic Field Graphs with Graph-Conditioned Fusion for Global Medium-Range Weather Forecasting ​

Author: Zhisheng Chen, Jinhan Li, Yuxuan Li, Yuan Gao, Hao Wu, Zheng Lu, Jinlong Du, Kun Wang, Bo An
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09286v1 Announce Type: cross Abstract: Global medium-range weather forecasting requires modeling structured yet state-dependent interactions among heterogeneous atmospheric fields. Existing data-driven models largely learn these interactions implicitly, whereas equation-level physical con...

📖 Read original article


424. UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation ​

Author: Xuewan He, Tong Chu, Zihan Cheng, Yuchen Su, Qianxin Xia, Guoming Lu, Jielei Wang, Wen Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09287v1 Announce Type: cross Abstract: Data-Free Knowledge Distillation (DFKD) transfers knowledge from a pretrained teacher model to a compact student model by synthesizing semantically informative data, eliminating the need for access to the original training dataset. Existing DFKD meth...

📖 Read original article


425. DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation ​

Author: Wei Zhou, Wanyi Ning, Yinshang Guo, Qianxiao Fang, Haitao Qian, Yingpeng Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2608.09288v1 Announce Type: cross Abstract: Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions. Existing approaches usually fuse visual features directly in...

📖 Read original article


426. WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation ​

Author: Peterson Co, Sicheng Hu, Chunxuan Jiao, Hongyang Cheng, Yulin Luo, Yijie Xu, Sixiang Chen, Zhongxia Zhao, Zihao Wang, DaFeng Chi, Peidong Liu, YuTong Chen, Henghua Liu, Zhihao Yuan, Huizhu Jia, Yuzheng Zhuang, Tianle Zhang, Liang Lin, Huajie Tan, Shanghang Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.09298v1 Announce Type: cross Abstract: Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this promise requires precise action-conditioned transitions rather than merely...

📖 Read original article


427. RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction ​

Author: Ambuj Mehrish, Sebastiano Vascon
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2608.09331v1 Announce Type: cross Abstract: Brain-to-audio reconstruction is limited by \emph{prior domination}: when a pretrained generator is conditioned on a weak neural signal, it produces realistic but stimulus-inaccurate audio. We introduce RAG-Audio, which decodes fMRI into a semantic a...

📖 Read original article


428. Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute ​

Author: Nikita Kozodoi, Zainab Afolabi, Jack Butler
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.09351v1 Announce Type: cross Abstract: Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deployment. Self-consistency is one of the established approaches, which spends this budget entirely on t...

📖 Read original article


429. Deep Learning based Detection of Fishing Vessels and Fishing Monitoring using Nightlight Images ​

Author: Shantakar Mohanty, Prasun Kumar Gupta, Raian Vargas Maretto
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV, stat.AP

arXiv:2608.09360v1 Announce Type: cross Abstract: The demand for maritime surveillance has given rise to the need for monitoring fishing vessel activities, particularly in addressing the challenge of "dark vessels" that operate without Automatic Identification System (AIS) transmission. This study p...

📖 Read original article


430. FeedbackTrack: Visual-Cortex-Inspired Cross-Frame Feedback for Transformer Tracking ​

Author: Yueyang Cang, Xiaoteng Zhang, Zhiyuan Ning, Yuchen He, Li Shi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09369v1 Announce Type: cross Abstract: Visual object tracking requires effective temporal integration, yet most Transformer trackers still rely on predominantly feed-forward feature extraction. Existing temporal mechanisms typically update templates, prompts, queries, or prediction states...

📖 Read original article


431. Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation ​

Author: Farzan Farnia, Hossein Goli, Amin Gohari
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2608.09385v2 Announce Type: cross Abstract: Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. We introduce Imag...

📖 Read original article


Author: Rose Cymbler, Daniel Guez, Laurent Fabre
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR

arXiv:2608.09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one. Standard legal RAG treats the corpus as static; we arg...

📖 Read original article


433. Monotonicity-Guided Bottom-Up Petri Net Discovery: The SPECpp Framework ​

Author: Leah Tacke genannt Unterberg, Lisa L. Mannel, Wil M. P. van der Aalst
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2608.09398v1 Announce Type: cross Abstract: Process discovery is one of the central challenges in process mining. Petri nets are particularly attractive because simple local constructs can express complex behavior, including concurrency. While their global behavior may be difficult to analyze,...

📖 Read original article


434. Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models ​

Author: Rustem Ozakar, Eyup Gedikli
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09400v1 Announce Type: cross Abstract: Research regarding the sign language recognition mostly relies on RGB images, whileas sign language datasets that provide depth images are limited. Point clouds obtained from depth images can be used for sign language recognition with neural networks...

📖 Read original article


435. LITEWAY: LIghtweight HAR via Temporal Efficient highWAY ​

Author: Dominique Nshimyimana, Vitor Fortes Rey, Mengxi Liu, Bo Zhou, Paul Lukowicz
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.HC

arXiv:2608.09421v1 Announce Type: cross Abstract: Wearable human activity recognition (HAR) remains challenging due to the computational and energy constraints of deep learning models on resource-limited devices. Existing lightweight approaches often rely on recurrent architectures (e.g., GRU and LS...

📖 Read original article


436. ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models ​

Author: R'ois'in Luo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.09432v1 Announce Type: cross Abstract: Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order. Existing architectures address this limitation by explicitly incorporating...

📖 Read original article


437. How Simple Can It Get? From Interpretable Equations to Readable Rules for Financial Decision Making ​

Author: Adia Lumadjeng, Ilker Birbil, Erman Acar
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09433v1 Announce Type: cross Abstract: In regulated domains such as finance, a model that cannot be explained cannot be deployed, yet many interpretable classifiers defeat their own purpose by producing formulas with dozens of features that no regulator could read. We take the reverse dir...

📖 Read original article


438. WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training ​

Author: Zehao Chen, Gongxun Li, Tianxiang Ai, Yifei Li, Zixuan Huang, Wang Zhou, Tao Huang, Fuzhen Zhuang, Xianglong Liu, Jianxin Li, Deqing Wang, Yikun Ban
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09447v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch of offline distillation. The same feedback loop can nevertheless be unstable: each update changes bot...

📖 Read original article


439. Learning to Modulate, Not to Cycle: Soft Actor---Critic Recovers Inverter-Style Heat-Pump Control ​

Author: Faizan Ahmed, Aniket Dixit, James Brusey
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09453v1 Announce Type: cross Abstract: On--off cycling is the main cause of compressor wear in residential heat pumps, yet reinforcement learning (RL) controllers for buildings typically optimise only energy cost and thermal comfort, ignoring how much the learned policy cycles. We add a l...

📖 Read original article


440. RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation ​

Author: Boxiong Wang, Hui Kang, Geng Sun, Jiahui Li, Chao Yu, Daxin Tian
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09467v1 Announce Type: cross Abstract: Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-V...

📖 Read original article


441. MixFormer: Linear Transformer with Mixture of Memory Experts ​

Author: Yu Guo, Lei Duan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09468v1 Announce Type: cross Abstract: State Space Models (SSMs), as a mainstream research direction of linear Transformers, aim to achieve higher efficiency than standard Transformers in long-context modeling. However, existing SSMs suffer from limited input adaptivity and constrained me...

📖 Read original article


442. ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents ​

Author: Hongwei Yao, Yiming Liu, Meihui Chen, Jieling Chen, Zikun Chen, Yiling He, Wangze Ni, Cong Wang, Kui Ren
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.09476v1 Announce Type: cross Abstract: Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBench, a self-evolving benchmark that evaluates such behavior risk from ...

📖 Read original article


443. Beyond Uniform Restoration: Empowering All-in-One Restoration with Pixel-Level Multimodal Guidance ​

Author: Chunxiao Liu, Wei Liu, Anbin Xiong, Erli Meng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09482v1 Announce Type: cross Abstract: All-in-one image restoration is a unified low-level vision task that aims to effectively recover high-quality images from inputs degraded by various types and levels of corruption using a single model. Recent works have achieved remarkable progress b...

📖 Read original article


444. Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning ​

Author: Yuting Liu, Wei Wu, Jianzhe Zhao, Guibing Guo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.09507v1 Announce Type: cross Abstract: Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full preference summary...

📖 Read original article


445. Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts ​

Author: Kevin Thomas, Milosz Kasprzyk, Reuel C Igbokwe Onuigbo, Elliott Pert, Cameron Tovey, Jo~ao A. Leite, Olesya Razuvayevskaya, Carolina Scarton
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.SI

arXiv:2608.09510v1 Announce Type: cross Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measuring detector performance on f...

📖 Read original article


446. STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework ​

Author: Hanlin Jiang, Jionghao Huang, Shaofei Li, Bojia Yu, Peng Jiang, Yuxin Ren, Ning Jia, Yao Guo, Ding Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.09524v1 Announce Type: cross Abstract: Incident response planning is critical for restoring compromised software systems after cyberattacks. Common practice relies on expert-driven playbooks that encode fixed response procedures, but these static workflows struggle to adapt to evolving in...

📖 Read original article


447. RangeFactory: Scalable Construction of Multi-Hop Cyber Ranges ​

Author: Hanlin Jiang, Puyi Wang, Jiandong Jin, Shaofei Li, Zhan Shen, Pengli Wang, Ziming Wang, Yifeng Cai, Ning Jia, Yuxin Ren, Peng Jiang, Yao Guo, Ding Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.09526v1 Announce Type: cross Abstract: Real-world cyberattacks often require sustained progress across multiple hosts and network segments, making multi-hop cyber ranges essential infrastructure for studying and improving LLM agents' ability to sustain complete attack chains. Prior work h...

📖 Read original article


448. Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries ​

Author: Matthew Russo, Yash Agarwal, Tianyu Li, Zhuohan Gu, Michael Cafarella, Omar Khattab, Tim Kraska, Samuel Madden
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DB, cs.AI

arXiv:2608.09532v1 Announce Type: cross Abstract: Enterprises increasingly seek to query data lakes using natural language via AI-driven tools like semantic operators or deep research agents. However, the latter operates as an opaque black box, hiding its intermediate reasoning and data retrieval st...

📖 Read original article


449. TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability ​

Author: Vincent Cohen-Addad, Dimitris Paparas, Ernest van Wijland, Max Springer, Julien Canitrot-Paradis, Honghao Lin, David Woodruff, Adarsh Kumarappan, Rajesh Jayaram, Rudrajit Das, Lalit Jain, Ola Svensson, Silvio Lattanzi, Mislav Balunovic, Theophane Weber, Vahab Mirrokni
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.09538v1 Announce Type: cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top theoretical computer scien...

📖 Read original article


450. Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs ​

Author: Hongli Shen, Shaopeng Fu, Qinbo Zhang, Jian Li, Di Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2608.09542v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent methods align LRMs using direct refusals or safety rationales, yet often focus on prompt patterns ra...

📖 Read original article


451. ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models ​

Author: Yilin Jiang, Xiaorong Zhu, Fei Tan, Zicheng Zhang, Kaiyi Huang, Yang Yu, Zexuan Fei, Yiming Luo, Keqian Li, Hao Hao, Guangtao Zhai, Aimin Zhou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2608.09548v2 Announce Type: cross Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe...

📖 Read original article


452. From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation ​

Author: Zeyuan Ma, Jiaxin Chen, Di Huang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09564v1 Announce Type: cross Abstract: UAV vision-language navigation (UAV-VLN) focuses on enabling an aerial agent to follow natural-language instructions in open 3D environments from egocentric visual observations. Current approaches suffer from three coupled issues: weak grounding of i...

📖 Read original article


453. Distributed Optimization with Streaming Data: A Temporal Weighting Perspective ​

Author: Muhammad Faraz Ul Abrar, Nicol`o Michelusi, Erik G. Larsson
Published: 8/11/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG, cs.SY, eess.SY, math.OC

arXiv:2608.09565v1 Announce Type: cross Abstract: Optimization theory is a widely used tool for intelligent decision-making. While classical optimization deals with fixed, time-invariant objective functions, many modern applications operate in dynamic environments where data arrive sequentially, and...

📖 Read original article


454. MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection ​

Author: Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2608.09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherw...

📖 Read original article


455. Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluation ​

Author: Yifei Xue, Yuanchen Fei, Hao Zhang, Chenzhi Nie, Tie ji, Yizhen Lao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09594v1 Announce Type: cross Abstract: Recently, AI-driven video generation has attracted considerable attention. This surge increases the demand for reliable video quality assessment (VQA) metrics to evaluate AI-generated content (AIGC) videos and guide model optimization. Existing studi...

📖 Read original article


456. LEED: Local Embedding Evolution Distance for over-smoothing estimation and virtual node selection in GNN ​

Author: Killian Cressant, Pedro B. Velloso
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09596v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) suffer from two fundamental limitations: over-smoothing, where node representations become indistinguishable with depth, and over-squashing, where long-range information is compressed through limited message-passing chann...

📖 Read original article


457. TSPORec: Token Selection via Preference Optimization for LLM-Based Sequential Recommendation ​

Author: Wenqiao Zhu, Chao Xu, Haipang Wu, Ji Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.IR, cs.AI

arXiv:2608.09605v1 Announce Type: cross Abstract: Large Language Models (LLMs) have emerged as powerful tools for improving recommendation systems. The effectiveness of LLMs arises from their ability to harness rich textual information and their capacity to model heterogeneous user preferences based...

📖 Read original article


458. Structure-Enhanced Features and Quality-Aware Dynamic Anchor Scoring for Robust Lane Detection ​

Author: Weize Cai, Yongqi Dong, Zhida Shao, Yichen Liu, Zixin Fu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV, eess.SP

arXiv:2608.09610v1 Announce Type: cross Abstract: Lane detection requires recovering thin, elongated, and frequently occluded lane structures under challenging driving conditions. While anchor-based detectors provide efficient candidate generation, their performance is limited by two coupled issues:...

📖 Read original article


459. Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks ​

Author: Mingyu Luo, Ming Deng, Zilang Qiu, Yiming Cheng, Ci Tao, Xue Tan, Sijin Sun, Yangfu Li, Ping Chen, Jun Dai, Xiaoyan Sun
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CR

arXiv:2608.09624v1 Announce Type: cross Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. H...

📖 Read original article


460. NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation ​

Author: Haiyang Yan, Jinyue Guo, Yanchao Zhang, Bingqing Wang, Zhenchen Li, Jing Liu, Jiazheng Liu, Linlin Li, Hua Han
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09636v1 Announce Type: cross Abstract: Accurate 3D neuron segmentation in fluorescence microscopy is critical for neuroscience. However, the sparse and elongated morphology of neurons poses significant challenges to existing segmentation methods. These methods struggle to preserve both lo...

📖 Read original article


461. DUET: A Diversity-Quality Duet of Distillation Experts for Two-Step Video Generation ​

Author: Zian Li, Litong Gong, Borui Liao, Pengfei Liu, Xinyu Wang, Xinyuan Wei, Yifan Gao, Tiezheng Ge, Muhan Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09637v1 Announce Type: cross Abstract: Diffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practical deployment. Few-step distillation alleviates this cost, yet exposes a quality--diversity trade-off between it...

📖 Read original article


462. Predictive safety filter enhanced curriculum learning control for efficient vehicle dynamics controller ​

Author: Baocong Zhang, Siliang Lu, Chenyang Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.09653v1 Announce Type: cross Abstract: Recent advances in learning-based control have enabled impressive achievements in solving complex control problems in various domains. However, since learning-based control may not be able to realize safety-guaranties, it is of great importance to en...

📖 Read original article


463. Confusion-Geometry Rebalancing for Long-Tailed Adversarial Training ​

Author: Mengnan Zhao, Geyong Min, Lihe Zhang, Tianhang Zheng, Jie Cui
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09688v1 Announce Type: cross Abstract: Adversarial training under long tailed distributions suffers from a dual imbalance: the class imbalance skews the training objective toward head classes, and the adversarial inner maximization may further amplify this bias. Existing methods mitigate ...

📖 Read original article


464. Evaluating Generative Time-Series Models on Data with Point Masses ​

Author: Jian Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09692v1 Announce Type: cross Abstract: Many of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not rain, no ride is requested, no part is ordered. We report what happens when such data is evaluated carefully. Fi...

📖 Read original article


465. How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans ​

Author: Hasan Mahmud, Khawaja Abaid Ullah, Mohammad Javad Khojasteh, Jamison Heard, Prabu David
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2608.09717v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction from theory-grounded...

📖 Read original article


466. ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners ​

Author: Puyu Zeng, Simeng Qin, Jingzhi Li, Ju Jia, Zheli Liu, Xiaojun Jia
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.09732v1 Announce Type: cross Abstract: Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners, we find that current defenses mainly inspect individual skills, leaving risks from cross-skill composition ins...

📖 Read original article


467. Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach ​

Author: Xinyi Xu, Bingnan Xiao, Shuang Qin, Gang Feng, Tony Q. S. Quek
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC

arXiv:2608.09742v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an efficient way to fine-tune large models in federated learning paradigm. Inspired by the asymmetric roles of the ...

📖 Read original article


468. SR-OPSD: Self-Referenced On-Policy Self-Distillation ​

Author: Zhuo Sun, Entong Li, Yanlong Zhao, Xiaoyuan Cheng, Wenxuan Yuan, Kaiyu Li, Che Liu, Huihang Liu, Harrison Bo Hua Zhu, Li Zeng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.09745v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, providing a useful complement to reinforcement learning with sparse outcome rewards. However, the self-te...

📖 Read original article


469. Defining Decentralization: An Ontological Perspective ​

Author: Jakub Kacper Szel\k{a}g, Aydin Abadi, Mohammad Naseri
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG, cs.LO, cs.SY, eess.SY

arXiv:2608.09748v1 Announce Type: cross Abstract: Decentralization as a concept in computer science has existed for over half a century. Despite its fundamental role across domains such as security, distributed computing, artificial intelligence, cloud infrastructures, and Internet of Things (IoT) a...

📖 Read original article


470. MoNo: Multiscale Optimal Transport Neural Operator for Solving PDEs on General Geometries ​

Author: Zijiang Yang, Xiaomeng Wu, Dongmei Fu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09764v1 Announce Type: cross Abstract: Transformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatial observations into compact latent tokens and learning physical interactions in latent spaces. However, we reve...

📖 Read original article


471. Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness ​

Author: Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu, David Tan, Doreen Osmelak, Ona de Gibert, Ariun-Erdene Tumurchuluun, Ashok Urlana, Fedor Sizov, Hale Sirin, Jesujoba Alabi, Karrar Talib Abed, Mateusz Klimaszewski, Nikolay Bogoychev, Niyati Bafna, Patricia Schmidtova, Preksha Manjunath Shanbhag, Sherrie Shen, Vilem Zouhar, Vivek Iyer, Yasser Hamidullah, Yusser Al Ghussin, Zheng Zhao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.09766v1 Announce Type: cross Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural consider...

📖 Read original article


472. KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs ​

Author: Ghanshyam Verma, Simanta Sarkar, Devishree Pillai, Hotaka Shiokawa, Yourong Xu, Fiona Veazey, Peter Hubbert, Hui Su, Paul Buitelaar
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.09779v1 Announce Type: cross Abstract: Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly in domain-specific contexts where general-purpose LLMs and RAG tend to underperform. We hypothesize ...

📖 Read original article


473. Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization ​

Author: Dinh Tan Nguyen, Quang-Hien Kha, Le-Hoang Nguyen, Minh-Toan Dinh, Xuan-Huy Nguyen, Dac Phu Ho, Cao Truong Tran, Sai Ho Ling, Lan T Ho-Pham, Liem Pham, Nguyen Quoc Khanh Le
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09801v1 Announce Type: cross Abstract: Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this setting using a multi-task DETR framework, where shared representations support both image-level malignancy predictio...

📖 Read original article


474. Parameter Exploration for RLVR via Variational Learning ​

Author: Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2608.09805v1 Announce Type: cross Abstract: Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performa...

📖 Read original article


475. MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation ​

Author: Haoyu Yang, Meixing Shi, Zengjie Chen, Haoran Sun, Haitao Leng, Xiaoming Shi, Yuxiang Cai, Yankai Jiang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09818v1 Announce Type: cross Abstract: Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack precise localization, whereas medical segmenters typically rely on explic...

📖 Read original article


476. Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation ​

Author: Yubo Jiang, Fengying Xie, Zhiguo Jiang, Haopeng Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.09826v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong, which account for 63.0-68.0% of groups in our experiments. We propose SKALD (Skill-Anchored Latent Distillati...

📖 Read original article


477. Multi-Agent AI Safety as an Institutional Design Problem ​

Author: Abdullah X
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA

arXiv:2608.09828v1 Announce Type: cross Abstract: AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent work already shows that deployment rules can change collective behavior. Here we ask which parts of an...

📖 Read original article


478. Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy ​

Author: Rohan Bhagra, Mahantesh Halapannavar, Uddhav Bhattarai
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.09857v1 Announce Type: cross Abstract: Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility actions planning models propose. Like general-purpo...

📖 Read original article


479. Stealing Reasoning Traces from Proprietary LLM APIs ​

Author: Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, Maksym Andriushchenko
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the c...

📖 Read original article


480. Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains ​

Author: Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.09873v1 Announce Type: cross Abstract: We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Sc...

📖 Read original article


481. Energy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion Planning ​

Author: Yapeng Liu, Yuanzhao Zhai, Bo Ding, Huaimin Wang, Lin Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2608.09876v1 Announce Type: cross Abstract: Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must strictly conform to real-world execution dynamics. While latent world models offer a promising approach by predicting these dynamics,...

📖 Read original article


482. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning ​

Author: Bj"orn Engdahl, Adrian Kosowski, Jan Chorowski, Zuzanna Stamirowska, Przemys{\l}aw Uzna'nski, Junlin Jiang, Rohan Phadke, Remigiusz Kinas, Richard Zhong
Published: 8/11/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.LG, stat.ML

arXiv:2608.09888v1 Announce Type: cross Abstract: We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuously update the model's recurrent memory; the model then solves a query through iterative computation...

📖 Read original article


483. Fusion Training for Mathematical Generalization in Large Language Models ​

Author: Congfeng Cao, Pengyu Zhang, Jelke Bloem
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.09893v1 Announce Type: cross Abstract: Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and a thinking mode within a single model. However, its training dynamics, including the \emph{data rat...

📖 Read original article


484. From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch ​

Author: Laurens Samson, Iva Gornishka, Gossa L^o, Yuki M. Asano, Sennay Ghebreab
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2608.09925v1 Announce Type: cross Abstract: Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip o...

📖 Read original article


485. Multimodal Model Diffing for Feature Discovery and Control ​

Author: Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr, Christian Schroeder de Witt, Constantin Venhoff, Ronald Clark
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG

arXiv:2608.09928v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomp...

📖 Read original article


486. Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions ​

Author: Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, cs.CL

arXiv:2608.09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speec...

📖 Read original article


487. Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory ​

Author: Gordon Dai, Weijia Zhang, Jinhan Li, Siqi Yang, Chidera Onochie lbe, Srihas Rao, Arthur Caetano, Misha Sra
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CY, cs.HC, cs.MA

arXiv:2406.14373v3 Announce Type: replace Abstract: The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational social science research at scale. Building upon prior explorations of LLM agent design, our work introduces a simu...

📖 Read original article


488. LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning? ​

Author: Kexian Tang, Junyao Gao, Yanhong Zeng, Haodong Duan, Yanan Sun, Zhening Xing, Wenran Liu, Kai Chen, Kaifeng Lyu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2503.19990v4 Announce Type: replace Abstract: Many real-world applications of spatial intelligence, such as robotic control, autonomous driving, and automated assembly, require spatial reasoning across multiple sequential steps. However, the extent to which current Multimodal Large Language Mo...

📖 Read original article


489. How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning ​

Author: Hongyi James Cai, Junlin Wang, Xiaoyin Chen, Bhuwan Dhingra
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2505.24273v2 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) suggest that reinforcement learning (RL) effectively internalizes search strategies, yielding significant improvements on challenging reasoning tasks through extended chains of thought. While back...

📖 Read original article


490. EgoBrain: Synergizing Minds and Eyes For Human Action Understanding ​

Author: Nie Lin, Yansen Wang, Dongqi Han, Weibang Jiang, Jingyuan Li, Ryosuke Furuta, Yoichi Sato, Dongsheng Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2506.01353v3 Announce Type: replace Abstract: The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has shown tremendous promise in decoding human cognition and behavior from neural signals. In particular, the rise of...

📖 Read original article


491. ACEvo: Adversarial Co-Evolution of Problem Distributions and Solvers for Combinatorial Optimization ​

Author: Ruibo Duan, Yuxin Liu, Haoran Ye, Xinyao Dong, Zhiqiang Xu, Chenglin Fan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2506.02594v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to synthesize heuristic programs, yet most existing pipelines optimize solvers against fixed benchmark distributions. This static setup can obscure solver weaknesses and limit understanding of how ...

📖 Read original article


492. Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response ​

Author: Yangqing Zheng, Shunqi Mao, Dingxin Zhang, Weidong Cai
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2506.07223v2 Announce Type: replace Abstract: Large language models (LLMs) have substantially improved the planning capabilities of embodied agents, enabling their deployment in dynamic and safety-critical environments. However, these settings expose a critical limitation: inference latency. D...

📖 Read original article


493. Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents ​

Author: Thassilo M. Schiepanski, Nicholas Pi"el
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC

arXiv:2508.04412v3 Announce Type: replace Abstract: The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serialised user interface (UI) state, an LLM is expected to suggest input actions that incrementally solve the give...

📖 Read original article


494. Probabilistic Circuits for Knowledge Graph Completion with Reduced Rule Sets ​

Author: Jaikrishna Manojkumar Patil, Nathaniel Lee, Al Mehdi Saadat Chowdhury, YooJung Choi, Paulo Shakarian
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LO

arXiv:2508.06706v2 Announce Type: replace Abstract: Rule-based methods for knowledge graph completion provide explainable results, but often require tens of thousands of rules to achieve competitive performance. Although individual predictions may use only a few rules, reasoning over an entire datas...

📖 Read original article


495. From Mimicry to True Intelligence (TI) -- A New Paradigm for Artificial General Intelligence ​

Author: Meltem Subasioglu, Nevzat Subasioglu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2509.14474v3 Announce Type: replace Abstract: The debate around Artificial General Intelligence (AGI) remains open due to two fundamentally different goals: replicating human-level performance versus replicating human-like cognitive processes. We argue that performance-based definitions are in...

📖 Read original article


496. ToolUniverse: An open platform for democratizing AI scientists ​

Author: Shanghua Gao, Richard Zhu, Pengwei Sui, Zhenglun Kong, Sufian Aldogom, Yepeng Huang, Ayush Noori, Reza Shamji, Krishna Parvataneni, Theodoros Tsiligkaridis, Marinka Zitnik
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2509.23426v3 Announce Type: replace Abstract: AI scientists are emerging computational systems that serve as collaborative partners in discovery. These systems remain difficult to build because they are bespoke, tied to rigid workflows, and lack shared environments that unify tools, data, and ...

📖 Read original article


497. TempoBench: Reasoning Execution Without Causal Attribution Is Just Simulation ​

Author: Nikolaus Holzer, William Fishell, Baishakhi Ray, Mark Santolucito
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.FL

arXiv:2510.27544v3 Announce Type: replace Abstract: Current training paradigms, optimized for long-horizon reasoning trace execution, have made Large Language Models (LLMs) excel at pattern matching and forward simulation of reasoning, but underperform at counterfactual causal understanding and reas...

📖 Read original article


498. The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation ​

Author: Tim R. Davidson, Adam Fourney, Saleema Amershi, Robert West, Eric Horvitz, Ece Kamar
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2511.02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independently developed agents with different information, privileges, and tools. The success of these systems w...

📖 Read original article


499. Intelligence Foundation Model: A New Perspective to Approach Artificial General Intelligence ​

Author: Borui Cai, Yao Zhao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.10119v4 Announce Type: replace Abstract: We propose a new perspective for approaching artificial general intelligence (AGI) through an intelligence foundation model (IFM). Unlike existing foundation models (FMs), which specialize in pattern learning within specific domains such as languag...

📖 Read original article


500. The Belief-Desire-Intention Ontology for modelling mental reality and agency ​

Author: Sara Zuppiroli, Carmelo Fabio Longo, Anna Sofia Lippolis, Rocco Paolillo, Lorenzo Giammei, Miguel Ceriani, Francesco Poggi, Antonio Zinilli, Andrea Giovanni Nuzzolese
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.17162v2 Announce Type: replace Abstract: The Belief-Desire-Intention (BDI) model is a cornerstone for representing rational agency in artificial intelligence and cognitive sciences. Yet, its integration into structured, semantically interoperable knowledge representations remains limited....

📖 Read original article


501. M$^3$Prune: Hierarchical Communication Graph Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation ​

Author: Weizi Shao, Taolin Zhang, Zijie Zhou, Chen Chen, Chengyu Wang, Xiaofeng He
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.19969v2 Announce Type: replace Abstract: Recent advancements in multi-modal retrieval-augmented generation (mRAG), which enhance multi-modal large language models (MLLMs) with external knowledge, have demonstrated that the collective intelligence of multiple agents can significantly outpe...

📖 Read original article


502. Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering ​

Author: Zijian Fu, Changsheng Lv, Xianlin Zhang, Mengshi Qi, Huadong Ma
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2511.23304v3 Announce Type: replace Abstract: In this paper, we propose a novel Multi-Modal Scene Graph with Kolmogorov-Arnold Expert Network for Audio-Visual Question Answering (SHRIKE). The task aims to mimic human reasoning by extracting and fusing information from audio-visual scenes, with...

📖 Read original article


503. Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets ​

Author: Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2512.01045v2 Announce Type: replace Abstract: Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets, yet the construction of such datasets often remains labor-intensive, weakly traceable, and difficult to conf...

📖 Read original article


504. Agentic AI for Clustering, Relationship Discovery, and Semantic Trading in Prediction Markets ​

Author: Agostino Capponi, Alfio Gliozzo, Brian Zhu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2512.02436v2 Announce Type: replace Abstract: Prediction markets allow users to trade on outcomes of real-world events, but are prone to fragmentation with overlapping questions, implicit equivalences, and hidden contradictions across markets. We present an agentic AI (AAI) pipeline that auton...

📖 Read original article


505. Neuronal Attention Circuit (NAC) for Representation Learning ​

Author: Waleed Razzaq, Yun-Bo Zhao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2512.10282v4 Announce Type: replace Abstract: Attention improves representation learning over RNNs, but its discrete nature limits continuous-time (CT) modeling. We introduce Neuronal Attention Circuit (NAC), a novel, biologically inspired CT-attention mechanism that reformulates attention log...

📖 Read original article


506. Multi-Granular Node Pruning for Causal Circuit Discovery ​

Author: Muhammad Umair Haider, Hammad Rizwan, Hassan Sajjad, A. B. Siddique
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2512.10903v3 Announce Type: replace Abstract: Circuit discovery aims to identify minimal subnetworks that are responsible for specific behaviors in large language models (LLMs). Existing approaches primarily rely on iterative edge pruning, which is computationally expensive and limited to coar...

📖 Read original article


507. SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation ​

Author: Seoyeon Kim, Jaehyung Kim
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2601.09974v2 Announce Type: replace Abstract: Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant over time. However, real-world interactions are dynamic, where user interests continuously evolve, posing a ...

📖 Read original article


508. Position: Certifiable State Integrity Should Be Built from Local Validity, Not Global Scale ​

Author: Enzo Nicol'as Spotorno, Joao R. Campos, Ant^onio Augusto Medeiros Fr"ohlich
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2601.21249v2 Announce Type: replace Abstract: Breakthroughs in language and vision have motivated increasingly general foundation models for time series and physical dynamics, where evidence is promising but less mature. In safety-critical Cyber-Physical Systems (CPS), globally parameterized d...

📖 Read original article


509. AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts ​

Author: Libin Qiu, Zhirong Gao, Junfu Chen, Yuhang Ye, Liangyu Li, Weizhi Huang, Xiaobo Xue, Wenkai Qiu, Shuo Tang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2601.22758v2 Announce Type: replace Abstract: Large language model agents repeatedly encounter related tasks, yet systems that learn from trajectories commit every lesson to one predefined artifact form. A local constraint, a reusable procedure, and a delegated objective require different amou...

📖 Read original article


510. El Agente Gr\'afico: A Semantic Execution Runtime for Scientific Agents ​

Author: Jiaru Bai, Abdulrahman Aldossary, Thomas Swanick, Marcel M"uller, Yeonghun Kang, Changhyeok Choi, Naruki Yoshikawa, Zijian Zhang, Jin Won Lee, Tsz Wai Ko, Aiwei Yin, Mohammad Ghazi Vakili, Chris Crebolder, Varinia Bernales, Al'an Aspuru-Guzik
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SE, physics.chem-ph

arXiv:2602.17902v2 Announce Type: replace Abstract: Large language models (LLMs) can plan scientific workflows and generate code, but these capabilities do not specify how scientific state is validated, transferred and recorded across heterogeneous computational and experimental operations. Here we ...

📖 Read original article


Author: Hadar Peer, Carlos Hernandez, Sven Koenig, Ariel Felner, Oren Salzman
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.24084v2 Announce Type: replace Abstract: Empirical evaluation in multi-objective search (MOS) has historically suffered from fragmentation, relying on heterogeneous problem instances with incompatible objective definitions that make cross-study comparisons difficult. This standardization ...

📖 Read original article


512. A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling ​

Author: Kirill Skobelev, Eric Fithian, Yegor Baranovski, Jack Cook, Sandeep Angara, Shauna Otto, Zhuang-Fang Yi, John Zhu, Neeraj Mainkar, Margaux Masson-Forsythe, Daniel A. Donoho, X. Y. Han
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2603.27341v4 Announce Type: replace Abstract: Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but surgical benchmarks in particular are often missing from prominent medical benchmark suites. Since surgery r...

📖 Read original article


513. MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models ​

Author: Han Wang, Yifan Sun, Brian Ko, Mann Talati, Jiawen Gong, Zimeng Li, Naicheng Yu, Xucheng Yu, Wei Shen, Vedant Jolly, Huan Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2603.28590v3 Announce Type: replace Abstract: Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mismatch occurs, the CoT no longer faithfully reflects the actual reasons (i.e., decision-critical fact...

📖 Read original article


514. SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents ​

Author: Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Nathaniel Gorski, Jianxin Sun, Guoxi Liu, Helgi I. Ingolfsson, David Lenz, Hanqi Guo, Hongfeng Yu, Teja Leburu, Michael Molash, Bei Wang, Tom Peterka, Chaoli Wang, Shusen Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.GR, cs.HC

arXiv:2603.29139v4 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visualization (SciVis) tasks. Despite rapid progress, the community lacks a principled and reproducible ben...

📖 Read original article


515. Self-Routing: Parameter-Free Expert Routing from Hidden States ​

Author: Jama Hussein Mohamud, Drew Wagner, Mirco Ravanelli
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.00421v3 Announce Type: replace Abstract: Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a learned router to map hidden states to expert assignments. In this work, we ask whether a dedicated learned rout...

📖 Read original article


516. AIVV: Neuro-Symbolic LLM Agent-Integrated Verification and Validation for Trustworthy Autonomous Systems ​

Author: Jiyong Kwon, Ujin Jeon, Sooji Lee, Guang Lin
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.02478v2 Announce Type: replace Abstract: Deep learning models excel at detecting anomaly patterns in normal data. However, they do not provide a direct solution for anomaly classification and scalability across diverse control systems, frequently failing to distinguish genuine faults from...

📖 Read original article


517. A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges ​

Author: Chenchen Kuai, Jiwan Jiang, Zihao Zhu, Hao Wang, Keshu Wu, Zihao Li, Yunlong Zhang, Chenxi Liu, Zhengzhong Tu, Zhiwen Fan, Yang Zhou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2604.07650v2 Announce Type: replace Abstract: The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Shared pretraining data, distillation, and alignment pipelines can induce hidden behavioral dependencies, or la...

📖 Read original article


518. SkillClaw: Let Skills Evolve Collectively with Agentic Evolver ​

Author: Ziyu Ma, Shidong Yang, Yuxiang Ji, Xucong Wang, Yong Wang, Yiming Hu, Tongwen Huang, Xiangxiang Chu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2604.08377v2 Announce Type: replace Abstract: Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static after deployment. As a result, similar workflows, tool usage patterns, and failure modes are repeatedly redi...

📖 Read original article


519. DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? ​

Author: Young-Suk Lee, Ramon Fernandez Astudillo, Radu Florian
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.09251v3 Announce Type: replace Abstract: Deep research agents increasingly interleave web browsing with multi-step computation, yet existing benchmarks evaluate these capabilities in isolation, creating a blind spot in assessing real-world performance. We introduce DRBENCHER, a synthetic ...

📖 Read original article


520. FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks ​

Author: Yupeng Cao, Haohang Li, Weijin Liu, Wenbo Cao, Anke Xu, Lingfei Qian, Xueqing Peng, Minxue Tang, Zhiyuan Yao, Jimin Huang, K. P. Subbalakshmi, Zining Zhu, Jordan W. Suchow, Yangyang Yu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CE, cs.CL, cs.MM

arXiv:2604.10015v3 Announce Type: replace Abstract: Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks. While existing benchmarks have begun evaluating financial tool calling, they focus...

📖 Read original article


521. Collaborative Multi-Agent Scripts Generation for Enhancing Imperfect-Information Reasoning in Murder Mystery Games ​

Author: Keyang Zhong, Junlin Xie, Hefeng Wu, Haofeng Li, Guanbin Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.11741v2 Announce Type: replace Abstract: Vision-language models (VLMs) have shown impressive capabilities in perceptual tasks, yet they degrade in complex multi-hop reasoning under multiplayer game settings with imperfect and deceptive information. In this paper, we study a representative...

📖 Read original article


522. RankGuide: Tensor-Rank-Guided Routing and Steering for Efficient Reasoning ​

Author: Jiayi Tian, Yupeng Su, Ryan Solgi, Souvik Kundu, Zheng Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.16694v2 Announce Type: replace Abstract: Large reasoning models (LRMs) enhance problem-solving capabilities by generating explicit multi-step chains of thought (CoT) reasoning; however, they incur substantial inference latency and computational overhead. To mitigate this issue, recent wor...

📖 Read original article


523. Time-Series Forecasting in Safety-Critical Environments: An Open-Source Package for EU-AI-Act-Compliant Development / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein Open-Source-Paket f\"ur die KI-VO-konforme Entwicklung ​

Author: Thomas Bartz-Beielstein, Eva Bartz
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2604.23859v2 Announce Type: replace Abstract: With spotforecast2-safe we present an integrated Compliance-by-Design approach to Python-based point forecasting of time series in safety-critical environments. A review of the relevant open-source tooling shows that existing compliance solutions o...

📖 Read original article


524. ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems ​

Author: Alexander Bering
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2604.23878v3 Announce Type: replace Abstract: ZenBrain is a seven-layer, neuroscience-derived memory architecture for LLM agents that unifies fifteen mechanisms - from Two-Factor synaptic consolidation to a Simulation-Selection sleep loop - under a single MemoryCoordinator: nine foundational a...

📖 Read original article


525. FitText: Evolving Agent Tool Ecologies via Memetic Retrieval ​

Author: Kyle Zheng, Han Zhang, Renliang Sun, Chenchen Ye, Wei Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.IR, cs.LG, cs.MA

arXiv:2605.02411v4 Announce Type: replace Abstract: Efficient reasoning is not only a matter of shortening an answer trace; for tool-using agents, it also depends on whether the agent is reasoning over the right action space. As API ecosystems scale to tens of thousands of endpoints, the semantic ga...

📖 Read original article


526. C2L-Net: A Data-Driven Model for State-of-Charge Estimation of Lithium-Ion Batteries During Discharge ​

Author: Khoa Tran, Tri Le, Nhu Nguyen Gia, T. Nguyen-Thoi, Vin Nguyen-Thai, Duong Tran Anh, Hung-Cuong Trinh
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.08653v2 Announce Type: replace Abstract: Accurate state-of-charge (SOC) estimation is critical for the safe and efficient operation of lithium-ion batteries in battery management systems (BMS). Although data-driven approaches can effectively capture nonlinear battery dynamics, many existi...

📖 Read original article


527. AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization ​

Author: Hyunmin Hwang, Jaemin Kim, Choonghan Kim, Hangeol Chang, Jong Chul Ye
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.08704v3 Announce Type: replace Abstract: Multi-agent reasoning has shown promise for improving the problem-solving ability of large language models by allowing multiple agents to explore diverse reasoning paths. However, most existing multi-agent methods rely on inference-time debate or a...

📖 Read original article


528. Understanding and Mitigating Premature Confidence for Better LLM Reasoning ​

Author: Jingchu Gai, Guanning Zeng, Christina Baek, Chen Wu, J. Zico Kolter, Andrej Risteski, Aditi Raghunathan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2605.24396v2 Announce Type: replace Abstract: Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test-time compute. Improving reasoning quality directly would require process reward models, but the...

📖 Read original article


529. PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management ​

Author: Yuxuan Zhao, Sijia Chen, Ningxin Su
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, q-fin.PM

arXiv:2605.27887v3 Announce Type: replace Abstract: Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked. Existing benchmarks exhibit two gaps: they are often equity-only and ignore cross-asset correlation...

📖 Read original article


530. The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary ​

Author: Dongxin Guo, Jikun Wu, Siu Ming Yiu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2606.00376v2 Announce Type: replace Abstract: Extended chain-of-thought reasoning can degrade performance on deterministic state-tracking tasks, not solely because of preference biases but, on the evidence we present, because of information-theoretic limits in the capacity of decoder-only atte...

📖 Read original article


531. Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition ​

Author: Wanlong Fang, Tianle Zhang, Wen Tao, Alvin Chan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.00959v2 Announce Type: replace Abstract: Understanding how multimodal large language models use different modalities is important for reliable reasoning. We employ Partial Information Decomposition (PID) as a decision-level lens and introduce Sensory PID, a conditional formulation that co...

📖 Read original article


532. Human agency in initial human-AI proof formalization workflows ​

Author: Katherine M. Collins, Simon Frieder, Jonas Bayer, Jacob Loader, Jeck Lim, Peiyang Song, Fabian Zaiser, Lexin Zhou, Shanda Li, Sam Looi, Joshua B. Tenenbaum, Umang Bhatt, Adrian Weller, Jose Hernandez-Orallo, Cameron E. Freer, Valerie Chen, Ilia Sucholutsky
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.04273v2 Announce Type: replace Abstract: For centuries, human mathematicians have written proofs to substantiate their mathematical arguments; yet, the ability to automatically verify the validity of proofs has long been a challenge. Advances in AI systems' ability to generate code and en...

📖 Read original article


533. SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research ​

Author: Xiaochong Lan, Quan Chen, Kun Tao, Xinyu Tang, Tianshu Wang, Qianggang Cao, Xinyu Kong, Zujie Wen, Zhiqiang Zhang, Jun Zhou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.09730v2 Announce Type: replace Abstract: Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inherently finite. Recent work explores a paradigm where a main agent de...

📖 Read original article


534. READER: Dynamic LLM Provenance from Query-Varying Interactions ​

Author: Jiaxu Liu, Sunnan Mu, Dong Huang, Liuyin Wang, Jing Shao, Jie Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.10794v3 Announce Type: replace Abstract: Existing black-box LLM provenance methods achieve comparability by querying every candidate model with the same diagnostic prompts. In deployment, auditors inherit a different evidence stream: heterogeneous prompt-response traces that arrive increm...

📖 Read original article


535. Unbiased Canonical Set-Valued Oracles Via Lattice Theory ​

Author: Jobst Heitzig
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2606.26418v3 Announce Type: replace Abstract: An oracle that tells you the probability of some future event can change that very probability because you act on the answer. We argue that this performativity is OK as people consult oracles to be informed, and hence moved, by the answer. We worry...

📖 Read original article


536. Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation ​

Author: Han-yu Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2606.26502v3 Announce Type: replace Abstract: Large reasoning models (LRMs) spend more reasoning tokens on problems that take humans longer, suggesting sensitivity to a similar structure of difficulty. That alignment identifies which problems elicit more deliberation and leaves open how a diff...

📖 Read original article


537. NormAct: Benchmarking Embodied Agents' Proactive Compliance with Unspoken Social Norms ​

Author: Shiyun Zhao, Xinwei Song, Tianyu Guo, Xiaomeng Gao, Mingyuan Liu, Xu Han, Yuanyuan Zhang, Zhenliang Zhang, Xue Feng, Bo Dai
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2606.27826v3 Announce Type: replace Abstract: Embodied agents driven by multimodal large language models (MLLMs) can often complete everyday tasks from visual observations, but goal achievement does not establish whether they proactively respect unstated social norms. Existing benchmarks asses...

📖 Read original article


538. ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair ​

Author: Chiwang Luk, Matin Mohammad Najafi, Zhifeng Jia, Wei Yang, Xiuchang Li, Jinwei Zhu, Yang Ren, Lei Chen, Gao Cong
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.01916v4 Announce Type: replace Abstract: Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad searches, and long terminal outputs where useful evidence is mixed with irrelevant code and logs. This paper presen...

📖 Read original article


539. LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL ​

Author: Yujin Kim, Namgyu Ho, Sangmin Hwang, Joonkee Kim, Yongjin Yang, Sangmin Bae, Seungone Kim, Jaehun Jung, Se-Young Yun, Hwanjun Song
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.04412v2 Announce Type: replace Abstract: Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals. While recent methods adapt these rubrics to the evolving policy during training, the training pro...

📖 Read original article


540. Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale ​

Author: Ziting Wang, Yin Li, Zuhao Yang, Xiuchang Li, Jiale Bai, Gao Cong
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.06233v3 Announce Type: replace Abstract: LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in heterogeneous enterprise set...

📖 Read original article


541. Experience Memory Graph: One-Shot Error Correction for Agents ​

Author: Wenjun Wang, Yuchen Fang, Fengrui Liu, Zibo Liang, Kai Zheng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.13884v2 Announce Type: replace Abstract: Large Language Model (LLM) agents have shown remarkable capabilities in autonomous decision-making by generating sequential trajectories of states, actions, and observations. However, in complex, long-horizon tasks, these agents frequently suffer f...

📖 Read original article


542. Concept-Guided Spatial Regularization for World Models in Atari Pong ​

Author: Yukuan Lu, Zaishuo Xia, Weyl Lu, Yubei Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.15142v2 Announce Type: replace Abstract: World models are usually evaluated as components of model-based reinforcement learning (MBRL) systems, leaving their standalone reliability understudied. We reproduce five visual world-model agents in Atari Pong -- DreamerV3, DIAMOND, TWISTER, Simu...

📖 Read original article


543. Exact Network Surgery: Functional Invariance and Gradient Plasticity in Reactive Computational Graphs ​

Author: Abdallah Khemais (ISITCOM, University of Sousse)
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.PL

arXiv:2607.16568v2 Announce Type: replace Abstract: Function-preserving network growth techniques such as Net2Net and progressive stacking expand a model's capacity without destroying its learned function, but existing formulations either tolerate numerical perturbations or require a full rebuild of...

📖 Read original article


544. ProbSPARQL: Querying Knowledge Graphs with Multi-dimensional, Uncertain Numeric Data ​

Author: Jingcheng Wu, Ratan Bahadur Thapa, Daniel Hernandez, Hongkuan Zhou, Steffen Staab
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.18262v2 Announce Type: replace Abstract: The SFB 1574 Circular Factory is building a shared knowledge graph infrastructure for integrating data about returned products. A central challenge is that circular-factory data include numeric measurements that (i) originate from sensors or are de...

📖 Read original article


545. SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data ​

Author: Wael AbdAlmageed
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.20402v2 Announce Type: replace Abstract: In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensional inputs. Further, the predicate vocabulary, argument structure, and trusted evidence are supplied by a Knowledge Graph (KG), or...

📖 Read original article


546. From Errors to Rules: Iterative Prompt Optimization for Text Classification ​

Author: Yueying Cui, Renhao Xue, Yi Zhang, Mukul Prasad
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.20497v2 Announce Type: replace Abstract: Prompt optimization for text classification spans diverse approaches, from demonstration selection to exploration-based search to error-driven diagnosis, each with known but incompletely characterized strengths and limitations. We conduct a compreh...

📖 Read original article


547. Physical AI Governance: From Theory to Practice Across Life Cycle ​

Author: Wang Yang, Shaobo Wang, Hongxuan Liu, Xiaoran Cai, Yunyu He, Jingzong Zhou, Mengzhong Ma, Yi Yu, Rohit Sharma, Jingjing Fu, Peng Qi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.RO

arXiv:2607.22877v2 Announce Type: replace Abstract: With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact with, and act in the physical world. Unlike traditional AI, Physical AI operates under real-time sa...

📖 Read original article


548. RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning ​

Author: Xi Chen, Hongru Zhou, Shiyu Feng, Hanyu Zhou, Huahui Yi, Rongsheng Wang, Tiancheng He, Kun Wang, Pingping Liu, Qiankun Li, Sicheng Lin, Huiying Ou, Xiaohong Zheng, Tianying Zang, Zhuohang Wu, Leheng Jiang, Kexin Cao, Wenhan Zhang, ChengYi Li, Zhiyang Wang, Songlin Li, Benyou Wang, Ningbei Yin, Shaoting Zhang, Weili Fu, Jian Li, Kang Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.23290v2 Announce Type: replace Abstract: Rare diseases represent one of the most challenging settings for clinical decision-making, where heterogeneous presentations, sparse evidence and limited expertise create persistent uncertainty throughout the care pathway. Although artificial intel...

📖 Read original article


549. When benchmark inferences do not compose: Projectibility in AI evaluation ​

Author: Brett Reynolds
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG

arXiv:2607.26159v2 Announce Type: replace Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with ass...

📖 Read original article


550. SemPIC: Learning Semantic Position-Independent KV Caches ​

Author: Hui Xie, Peng Xiao, Yutong Deng, Shuoran Dou, Jian Yang, Jinyang Guo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28069v2 Announce Type: replace Abstract: Long-context retrieval and agentic workloads repeatedly reuse the same documents under changing instructions, histories, and document orders. Prefix caching cannot exploit this reuse, while position-independent caching (PIC) remains unreliable beca...

📖 Read original article


551. ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs ​

Author: Bingchen Liu, Yuanyuan Fang, Lei Liu, Guangyuan Dong, Xing Fu, Yuanyuan Gao, Shuyue Wei, Xin Li, Xiangtian Meng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28126v2 Announce Type: replace Abstract: Long-horizon steel-equipment inspection requires reasoning over heterogeneous records accumulated across repeated inspection cycles. Existing retrieval-augmented generation systems treat historical logs as a static corpus and retain records without...

📖 Read original article


552. A foundation model of numerical intelligence with cross-disciplinary generalization ​

Author: Chenghan Wu, Zongmin Yu, Liu Yang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28432v2 Announce Type: replace Abstract: Intelligence is commonly understood as the ability to acquire and apply knowledge, adapt to unfamiliar situations and solve new problems. Large language models exhibit this capacity by inferring task-relevant knowledge from textual context and appl...

📖 Read original article


553. InfoOps Bench: A live information operations safety benchmark ​

Author: Dorian Quelle, Lisa-Maria Neudert, Jonathan Bright, John Gallacher
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2607.28503v3 Announce Type: replace Abstract: In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for use by authoritarian state "information operations": intentional, coordinated activities by one ...

📖 Read original article


554. Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates ​

Author: Bohan Chen, Shivam N. Patel, Richard Hoffmann, Sam Looi, Tony Yue Yu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.00326v2 Announce Type: replace Abstract: Tool calling allows large language models (LLMs) to invoke external computation during problem solving, a useful capability in various fields including AI for mathematics. We study this setting through weighted sum-of-squares (SOS) decomposition, a...

📖 Read original article


555. SymboUQ: Symbolic Uncertainty Quantification for Spatial Reasoning in LLMs ​

Author: Dahai Yu, Lin Jiang, Rongchao Xu, Guang Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.00417v2 Announce Type: replace Abstract: Although large language models (LLMs) can produce fluent spatial reasoning traces, their intermediate relations may fail to support the final conclusion, making token-level confidence insufficient for final-answer reliability estimation. Existing f...

📖 Read original article


556. Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Models ​

Author: Yibin Dong
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.00591v2 Announce Type: replace Abstract: A calibrated stochastic world model can reveal how uncertain a future is without revealing why it branches. The same conditional future law can arise because an observation aliases physical states or because dynamics remain random after the declare...

📖 Read original article


557. Evolutionary Curriculum Learning Improves Biological Sequence Modeling ​

Author: Richard Zhu, Kento Nishi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, q-bio.BM, stat.ML

arXiv:2608.00697v2 Announce Type: replace Abstract: Variational autoencoders (VAEs) trained on multiple sequence alignments (MSAs) have emerged as powerful generative models for biological sequences, with applications ranging from disease variant prediction to functional RNA design. However, standar...

📖 Read original article


558. The Scaling Paradox in Human-AI Collaboration ​

Author: Anyan Qi, Mengxin Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, econ.GN, q-fin.EC

arXiv:2608.00818v2 Announce Type: replace Abstract: The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems scale, their capabilities tend to improve predictably. Yet, in real-world applications, AI rarely operates in i...

📖 Read original article


559. Agentic Stage-One Stellarator Optimization: Autonomous Multi-Objective Search for Finite-Beta Equilibria ​

Author: Tingjia Zhang, Zhuoran Meng, Runlai Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.01344v2 Announce Type: replace Abstract: Stage-one stellarator design searches a high-dimensional family of three-dimensional plasma boundaries and fixed-boundary MHD equilibria for configurations that jointly meet requirements on confinement, field-line topology, force balance, stability...

📖 Read original article


560. Emergence Invariance: From Symbolized Thought to Structural Control ​

Author: Yi Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.01548v2 Announce Type: replace Abstract: Language-first intelligence is constrained by which distinctions enter its symbolic record, which mappings its language--interpreter--environment complex can execute, and which possibilities can be realized with finite resources. We formalize these...

📖 Read original article


561. State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking ​

Author: Xiaohe Li, Yang Lu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.03425v2 Announce Type: replace Abstract: Transformer-based architectures have dominated sequence modeling, largely due to the expressive power of attention mechanisms. However, for a class of deterministic state tracking tasks---such as parity checking, modular counting, and parenthesis m...

📖 Read original article


562. When Efficiency Becomes Fragility: Exploiting Dynamic Routing Vulnerabilities in Adaptive UAV Tracking ​

Author: Shaofeng Liang, Runwei Guan, Wenshuo Chen, Jiemin Wu, Bowen Tian, Haozhe Jia, Kaishen Yuan, Songning Lai, Daizong Liu, Yutao Yue
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.03902v2 Announce Type: replace Abstract: Resource constraints on UAV platforms have driven a paradigm shift in aerial tracking, from pursuing performance toward balancing accuracy with efficiency. Adaptive Transformer Trackers, which leverage an input-dependent dynamic routing architectur...

📖 Read original article


563. The Transformer Revolution, Part 1: Dynamic Processing through Output-Weight Interconnections ​

Author: Marco Giunti, Fabrizia Giulia Garavaglia
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.NE

arXiv:2608.03921v2 Announce Type: replace Abstract: This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models merely reproduce statistical regularities learned in training, we argue that Transformers construct and appl...

📖 Read original article


564. Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load ​

Author: Thomas Bartz-Beielstein
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05018v2 Announce Type: replace Abstract: Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law designate as critical. Determinism, reproducibility, and auditability are engineering requirements rather tha...

📖 Read original article


565. Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks ​

Author: Boxiu Li, Zimo Wen, Yijia Fan, Chuan Wen, Fan Yang, Hangxi Guo, Jiaao Wu, Jiachen Zhang, Junxiang Lei, Mukai Li, Ruize Tang, Runjing Gu, Shibo Hu, Sihan Chen, Sufeng Guo, Wanbo Zhang, Xian Zhang, Xiaoyu Chen, Xuanhe Zhou, Xuyao Huang, Yifei Gao, Yifei Shen, Yilin Chen, Yuheng Wu, Yuzhe Zhang, Zelong Zhao, Zhijie Deng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05144v2 Announce Type: replace Abstract: Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolvi...

📖 Read original article


566. Small Foundation Models of Human Cognition and Behaviour ​

Author: Nick Oh, Fernand Gobet
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CY

arXiv:2608.05224v3 Announce Type: replace Abstract: Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We t...

📖 Read original article


567. Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows ​

Author: Nimisha Karnatak, Max Van Kleek, Nigel Shadbolt
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.HC

arXiv:2608.05602v2 Announce Type: replace Abstract: Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what condi...

📖 Read original article


568. ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion ​

Author: Jiafan Li, Mengxue Yang, Jiaqi Zhu, Liang Chang, Ying Li, Hongan Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.05833v2 Announce Type: replace Abstract: Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text a...

📖 Read original article


569. Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts ​

Author: Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio, Holger Boche
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.CL

arXiv:2608.06123v2 Announce Type: replace Abstract: Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce...

📖 Read original article


570. Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques ​

Author: Hotaka Maeda, Yikai Lu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.06609v2 Announce Type: replace Abstract: Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation. We aimed to build a near-comprehensive AIE model by predicting ...

📖 Read original article


571. bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning ​

Author: Koushik Howlader, Tirtho Roy, Md Tauhidul Islam, Wei Le
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.06727v2 Announce Type: replace Abstract: Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. Mixture-of-Recursions (MoR) improves efficiency through adaptive token-choice or expert-choice routing....

📖 Read original article


572. $A^2E$ : An End-to-End Agent Auditing Engine ​

Author: Haoning Wang, Mingxun Zhang, Chenyue Yu, Yingjun Shang, Xia Hu, Guanchu Wang, Na Zou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07346v2 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains. The fast-evolving harness ecosystem has also made rigorous capability evaluation increas...

📖 Read original article


573. CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing ​

Author: Yan Zhou, Yue Ouyang, Kaiyang Zheng, Suncheng Xiang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07424v2 Announce Type: replace Abstract: Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a chain of thought, or applying a stronger evaluator. Under a fixed inference budget, these choices compete. This paper formulates te...

📖 Read original article


574. TEPA: Revoking Stale Memories for Conflict-Robust Language Agents ​

Author: Yan Zhou, Yue Ouyang, Kaiyang Zheng, Suncheng Xiang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AI

arXiv:2608.07429v2 Announce Type: replace Abstract: Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also creates a central falsifiability problem: when the world changes, stale memories can remain retrievable and pollute the prompt. We char...

📖 Read original article


575. Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization ​

Author: Nachuan Xiao, Xiaoyin Hu, Kim-Chuan Toh
Published: 8/11/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.LG, stat.ML

arXiv:2307.10053v5 Announce Type: replace-cross Abstract: In this paper, we focus on providing convergence guarantees for stochastic subgradient methods in minimizing nonsmooth nonconvex functions. We first investigate the global stability of a general framework for stochastic subgradient methods, w...

📖 Read original article


576. Explainable Machine Learning-Based Security and Privacy Protection Framework for Internet of Medical Things Systems ​

Author: Ayoub Si-ahmed, Mohammed Ali Al-Garadi, Narhimene Boustia
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2403.09752v4 Announce Type: replace-cross Abstract: The Internet of Medical Things transcends traditional medical boundaries, enabling a transition from reactive treatment to proactive prevention. This innovative method revolutionizes healthcare by facilitating early disease detection and tail...

📖 Read original article


577. Ethical Framework for Responsible Foundational Models in Medical Imaging ​

Author: Debesh Jha, Gorkem Durak, Abhijit Das, Jasmer Sanjotra, Onkar Susladkar, Suramyaa Sarkar, Ashish Rauniyar, Nikhil Kumar Tomar, Linkai Peng, Sirui Li, Koushik Biswas, Ertugrul Aktas, Elif Keles, Matthew Antalek, Zheyuan Zhang, Bin Wang, Xin Zhu, Hongyi Pan, Deniz Seyithanoglu, Alpay Medetalibeyoglu, Vanshali Sharma, Vedat Cicek, Amir A. Rahsepar, Rutger Hendrix, A. Enis Cetin, Bulent Aydogan, Mohamed Abazeed, Frank H. Miller, Rajesh N. Keswani, Hatice Savas, Sachin Jambawalikar, Daniela P. Ladner, Amir A. Borhani, Concetto Spampinato, Michael B. Wallace, Ulas Bagci
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2406.11868v2 Announce Type: replace-cross Abstract: The emergence of foundational models represents a paradigm shift in medical imaging, offering extraordinary capabilities in disease detection, diagnosis, and treatment planning. These large-scale artificial intelligence systems, trained on ex...

📖 Read original article


578. Transformer Explainer: Learning LLM Transformers with Interactive Visual Explanation and Experimentation ​

Author: Aeree Cho, Grace C. Kim, Alexander Karpekov, Seongmin Lee, Alec Helbling, Benjamin Hoover, Zijie J. Wang, Minsuk Kahng, Duen Horng Chau
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.HC

arXiv:2408.04619v2 Announce Type: replace-cross Abstract: The Transformer architecture underpins modern large language models powering state-of-the-art text generation and AI applications. However, its complexity makes it difficult for non-experts to learn. Existing resources often lack interactivit...

📖 Read original article


579. See Me, Believe Me: Causality, Intersectionality, and Interventions Improving the Appearance of Patients ​

Author: Kenya S. Andrews, Mesrob I. Ohannessian, Elena Zheleva
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2410.01227v2 Announce Type: replace-cross Abstract: In the context of medical records, patients often experience testimonial injustice, where the textual account undermines the validity of their experiences. Past work has demonstrated that intersectionality of demographic features is crucial t...

📖 Read original article


580. A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence ​

Author: Sharon Temtsin, Diane Proudfoot, David Kaber, Christoph Bartneck
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY

arXiv:2501.17629v2 Announce Type: replace-cross Abstract: Several studies claim that large language models have passed the Turing Test and hence can "think", yet none follow Turing's original instructions precisely. Passing the test holds significance as evidence that a machine demonstrates human-li...

📖 Read original article


581. LF${}^{2}$AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models ​

Author: Santiago Cuervo, Adel Moumen, Yanis Labrak, Sameer Khurana, Antoine Laurent, Mickael Rouvier, Phil Woodland, Ricard Marxer
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, eess.AS

arXiv:2503.06211v3 Announce Type: replace-cross Abstract: Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as audio and images while effectively leveraging that knowledge remains challenging. Perceptual modalities...

📖 Read original article


582. REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation ​

Author: Puzhen Yuan, Angyuan Ma, Yunchao Yao, Huaxiu Yao, Masayoshi Tomizuka, Mingyu Ding
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CL, cs.CV

arXiv:2503.22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition. Existing methods typically rely ...

📖 Read original article


583. When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output ​

Author: Shuoming Zhang, Jiacheng Zhao, Hanyuan Dong, Ruiyuan Xu, Zhicheng Li, Yangyu Zhang, Shuaijiang Li, Yuan Wen, Chunwei Xia, Zheng Wang, Xiaobing Feng, Huimin Cui
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2503.24191v4 Announce Type: replace-cross Abstract: Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (LLMs) increasingly serve as tooling platforms through structured output APIs, but the grammar-guided ...

📖 Read original article


584. An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning ​

Author: Tianbing Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated by systems such as OpenAI's O1~\cite{o1} and DeepSeek-R1~\cite{r1}. However, widely used algorithms s...

📖 Read original article


585. TreeHop: Efficient Embedding-Level Query Rewriter ​

Author: Zhonghao Li, Kunpeng Zhang, Jinghuai Ou, Shuliang Liu, Xuming Hu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.HC, cs.LG

arXiv:2504.20114v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) systems face significant challenges in multi-hop question answering (MHQA), where complex queries require synthesizing information across multiple document chunks. Existing approaches typically rely on ite...

📖 Read original article


586. Optimal Transport for Machine Learners ​

Author: Gabriel Peyr'e
Published: 8/11/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, math.OC

arXiv:2505.06589v3 Announce Type: replace-cross Abstract: Modern machine learning repeatedly manipulates probability measures: empirical datasets, generated samples, latent distributions, class-conditional laws, particle systems, weights of wide networks and attention patterns. Optimal transport is ...

📖 Read original article


587. X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation ​

Author: Peizhen Li, Longbing Cao, Xiao-Ming Wu, Runze Yang, Xiaohan Yu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.HC

arXiv:2505.11146v4 Announce Type: replace-cross Abstract: Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition challenge due to the significant domain gap between biological facial dynamics and mechanical control spaces. While visual synthesis ...

📖 Read original article


588. SuperCoder: Assembly Program Superoptimization with Large Language Models ​

Author: Anjiang Wei, Tarun Suresh, Huanmi Tan, Yinglun Xu, Gagandeep Singh, Ke Wang, Alex Aiken
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.PF, cs.PL, cs.SE

arXiv:2505.11480v4 Announce Type: replace-cross Abstract: Superoptimization is the task of transforming a program into a faster one, and ideally the very fastest possible one, while preserving its input-output behavior. In this work, we investigate whether large language models (LLMs) can serve as s...

📖 Read original article


589. SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning ​

Author: Jiashu He, Jinxuan Fan, Bowen Jiang, Ignacio Houine, Dan Roth, Alejandro Ribeiro
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2505.15062v5 Announce Type: replace-cross Abstract: Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly available. It is essential for solving complex questions in specialized domains where retrieving comprehens...

📖 Read original article


590. Transformer-Based Neural Quantum Digital Twins for Many-Body Spectral Reconstruction and Adaptive Quantum-Annealing Schedule Design ​

Author: Jianlong Lu, Hanqiu Peng, Hongrui Zhang, Ying Chen
Published: 8/11/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.ET

arXiv:2505.15662v3 Announce Type: replace-cross Abstract: We introduce Transformer-based Neural Quantum Digital Twins (Tx-NQDTs) to reconstruct the low-energy spectral evolution of many-body quantum systems along quantum-annealing paths, including ground- and first-excited-state energies, spectral g...

📖 Read original article


591. FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records ​

Author: Vincent Jeanselme, Zilin Jing, Aparajita Kashyap, Chao Pang, Florent Pollet, Young Sang Choi, Xinzhuo Jiang, Yuta Kobayashi, Yanwei Li, Sara Matijevic, Karthik Natarajan, Shalmali Joshi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2505.16941v4 Announce Type: replace-cross Abstract: Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (iii) poor transportability. Despite methodological advances in ...

📖 Read original article


592. The Cell Must Go On: Agar.io for Continual Reinforcement Learning ​

Author: Mohamed A. Mohamed, Kateryna Nekhomiazh, Vedant Vyas, Marcos M. Jose, Andrew Patterson, Marlos C. Machado
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2505.18347v3 Announce Type: replace-cross Abstract: Continual reinforcement learning (RL) concerns agents that are expected to learn continually, rather than converge to a policy that is then fixed for evaluation. This setting is well-suited to environments that the agent perceives as changing...

📖 Read original article


593. HyperFake: Hyperspectral Reconstruction and Attention-Guided Analysis for Advanced Deepfake Detection ​

Author: Pavan C Shekar, Pawan Soni, Vivek Kanhangad
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2505.18587v2 Announce Type: replace-cross Abstract: Deepfakes pose a significant threat to digital media security, with current detection methods struggling to generalize across different manipulation techniques and datasets. While recent approaches combine CNN-based architectures with Vision ...

📖 Read original article


594. From Alignment to Synthesis Contrastive Volumetric Grounding for Text-to-CT Generation ​

Author: Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2506.00633v3 Announce Type: replace-cross Abstract: Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded in volumetric space. Existing Text-to-CT approaches condition generation on encod...

📖 Read original article


595. WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks ​

Author: Atsuyuki Miyai, Zaiying Zhao, Kazuki Egashira, Atsuki Sato, Tatsumi Sunada, Shota Onohara, Hiromasa Yamanishi, Mashiro Toyooka, Kunato Nishina, Ryoma Maeda, Kiyoharu Aizawa, Toshihiko Yamasaki
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2506.01952v2 Announce Type: replace-cross Abstract: Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent and general framework for automating web-based tasks. As these agents rapidly improve and achieve st...

📖 Read original article


596. Contamination Means Overestimation? A Fine-Grained Empirical Study in Code Intelligence ​

Author: Zhen Yang, Hongyi Lin, Yifan He, Junqi Wang, Zeyu Sun, Shuo Liu, Jie Xu, Pengpeng Wang, Zhongxing Yu, Qingyuan Liang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2506.02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread adoption of Pretrained Language Models (PLMs) and Large Language Models (LLMs) has raised concerns r...

📖 Read original article


597. Transformer Circuits Can Realize Clustering Algorithms ​

Author: Kenneth L. Clarkson, Lior Horesh, Takuya Ito, Charlotte Park, Parikshit Ram
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2506.19125v2 Announce Type: replace-cross Abstract: Although transformers are most commonly optimized as statistical sequence models, it is unclear to what extent they can implement and learn exact algorithmic computations. Here, we specify a transformer implementation from first principles th...

📖 Read original article


598. MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks ​

Author: Adrian Marius Dumitran, Theodor-Pierre Moroianu, Mihnea-Vicentiu Buca
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL, cs.LG

arXiv:2507.03162v2 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These models exhibit remarkable capabilities in code-related tasks and problem-solving, raising questions abo...

📖 Read original article


599. Dynamic gain neuromodulation attenuates the stability gap under joint training ​

Author: Alejandro Rodriguez-Garcia, Anindya Ghosh, Srikanth Ramaswamy
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-bio.NC

arXiv:2507.14056v3 Announce Type: replace-cross Abstract: Recent work in continual learning has highlighted the stability gap -- a temporary performance drop on previously learned tasks when new ones are introduced. This phenomenon reflects a mismatch between rapid adaptation and strong retention at...

📖 Read original article


600. MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks ​

Author: Sara Papi, Maike Z"ufle, Marco Gaido, Beatrice Savoldi, Danni Liu, Ioannis Douros, Luisa Bentivogli, Jan Niehues
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.SD

arXiv:2507.19634v4 Announce Type: replace-cross Abstract: Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a single framework. As these models are rapidly evolving toward general-purpose instruction following a...

📖 Read original article


601. Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings ​

Author: Yahya Badran, Christine Preisach
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG

arXiv:2508.17092v2 Announce Type: replace-cross Abstract: Knowledge Tracing (KT) aims to predict a student's future performance based on their sequence of interactions with learning content. Many KT models rely on knowledge concepts (KCs), which represent the skills required for each item. However, ...

📖 Read original article


602. Deep Residual Echo State Networks: exploring residual orthogonal connections in untrained Recurrent Neural Networks ​

Author: Matteo Pinna, Andrea Ceni, Claudio Gallicchio
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2508.21172v3 Announce Type: replace-cross Abstract: Echo State Networks (ESNs) are a particular type of untrained Recurrent Neural Networks (RNNs) within the Reservoir Computing (RC) framework, popular for their fast and efficient learning. However, traditional ESNs often struggle with long-te...

📖 Read original article


603. NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models ​

Author: Chuhan Zhang, Ye Zhang, Bowen Shi, Yuyou Gan, Tianyu Du, Shouling Ji, Dazhan Deng, Yingcai Wu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2509.03985v2 Announce Type: replace-cross Abstract: In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. However, the continuous advancement of jailbreak attack techniques, designed to bypass safety mechanisms ...

📖 Read original article


604. ATLASFusion: Aggregation Tracking with Location-Aware Sparse Fusion for Robust Spatio-Temporal Multi-View Pedestrian Tracking ​

Author: Keisuke Toida, Taigo Sakai, Takeshi Nakamura, Hiroshi Shimizu, Kazuhiro Hotta
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2509.08421v2 Announce Type: replace-cross Abstract: For multimedia spatial intelligence through time, multi-view multi-object tracking (MVMOT) suffers from persistent challenges in maintaining consistent object identities across different camera views, leading to tracking inaccuracies. A key s...

📖 Read original article


605. Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis ​

Author: Anjiang Wei, Tianran Sun, Tarun Suresh, Haoze Wu, Ke Wang, Alex Aiken
Published: 8/11/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.CL, cs.LG

arXiv:2509.21629v4 Announce Type: replace-cross Abstract: Program verification relies on loop invariants, yet automatically discovering strong invariants remains a long-standing challenge. We investigate whether large language models (LLMs) can accelerate program verification by generating useful lo...

📖 Read original article


606. Topographic Constraints Shape Brain-Like Component Structure in Auditory Models ​

Author: Haider Al-Tahan, Mayukh Deb, Jenelle Feather, N. Apurva Ratan Murty
Published: 8/11/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.CV, cs.SD

arXiv:2509.24039v2 Announce Type: replace-cross Abstract: If topography is a fundamental feature of the brain, it should influence both how neurons are arranged in space (i.e. explain brain maps) and how information is structured within the neural population. The human auditory cortex provides a str...

📖 Read original article


607. Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents ​

Author: Zhiping Zhang, Yi Evie Zhang, Freda Shi, Tianshi Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CR

arXiv:2510.04465v3 Announce Type: replace-cross Abstract: LLM agents require personal information for personalization in order to effectively act on users' behalf, but this raises privacy concerns that can discourage data sharing, limiting both the autonomy levels at which agents can operate and the...

📖 Read original article


608. Deep Generative Model for Human Mobility Behavior ​

Author: Ye Hong, Yatao Zhang, Konrad Schindler, Martin Raubal
Published: 8/11/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.AI, cs.SI

arXiv:2510.06473v4 Announce Type: replace-cross Abstract: Understanding and modeling human mobility is central to challenges in transport planning, sustainable urban design, and public health. Despite decades of effort, simulating individual mobility remains challenging because of its complex, conte...

📖 Read original article


609. OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference ​

Author: Yuzhe Gu, Xiyu Liang, Jiaojiao Zhao, Enmao Diao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2510.07651v4 Announce Type: replace-cross Abstract: Large language models (LLMs) with extended context windows enable powerful applications but impose significant memory overhead, as caching all key-value (KV) states scales linearly with sequence length and batch size. Existing cache eviction ...

📖 Read original article


610. SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation ​

Author: Xiaobei Zhao, Xingqi Lyu, Xin Chen, Xiang Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2510.14357v2 Announce Type: replace-cross Abstract: Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily relying on manual operations or fixed railways for movement. The A2A benchmark and the AgriVLN method pion...

📖 Read original article


611. Self-Attention to Operator Learning-based 3D-IC Thermal Simulation ​

Author: Zhen Huang, Hong Wang, Wenkai Yang, Muxi Tang, Depeng Xie, Ting-Jung Lin, Yu Zhang, Wei W. Xing, Lei He
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR

arXiv:2510.15968v2 Announce Type: replace-cross Abstract: Thermal management in 3D ICs is increasingly challenging due to higher power densities. Traditional PDE-solving-based methods, while accurate, are too slow for iterative design. Machine learning approaches like FNO provide faster alternatives...

📖 Read original article


612. NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning ​

Author: Zhi Zhang, Yixian Shen, Congfeng Cao, Ekaterina Shutova
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2510.18940v2 Announce Type: replace-cross Abstract: Existing parameter-efficient fine-tuning (PEFT) methods primarily fall into two categories: addition-based and selective in-situ adaptation. The former, such as LoRA, introduce additional modules to adapt the model to downstream tasks, offeri...

📖 Read original article


613. Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation ​

Author: Dhrupad Bhardwaj, Julia Kempe, Tim G. J. Rudner
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, stat.ME, stat.ML

arXiv:2510.21891v2 Announce Type: replace-cross Abstract: To deploy large language models (LLMs) in high-stakes application domains that require substantively accurate responses to open-ended prompts, we need reliable, computationally inexpensive methods that assess the trustworthiness of long-form ...

📖 Read original article


614. QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture ​

Author: Shvetank Prakash, Andrew Cheng, Mark Mazumder, Arya Tschand, Varun Gohil, Jeffrey Ma, Jason Yik, Zishen Wan, Jessica Quaye, Elisavet Lydia Alvanaki, Avinash Kumar, Chandrashis Mazumdar, Tuhin Khare, Alexander Ingare, Ikechukwu Uchendu, Radhika Ghosal, Abhishek Tyagi, Chenyu Wang, Andrea Mattia Garavagno, Sarah Gu, Alice Guo, Grace Hur, Luca P. Carloni, Tushar Krishna, Ankita Nayak, Amir Yazdanbakhsh, Vijay Janapa Reddi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.LG, cs.SE

arXiv:2510.22087v3 Announce Type: replace-cross Abstract: The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent from current large language model (LLM) evaluations. To this end, we present QuArch (pronounced 'quark')...

📖 Read original article


615. SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery ​

Author: Qiwei Ma, Xukun Lu, Wang Liu, Puhong Duan, Xudong Kang, Shutao Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2510.22665v4 Announce Type: replace-cross Abstract: Synthetic Aperture Radar (SAR) is a critical imaging modality due to its all-weather operational capability. Although recent advances in self-supervised learning and masked image modeling (MIM) have enabled SAR foundation models, these approa...

📖 Read original article


616. Reasoning about Intent for Ambiguous Requests ​

Author: Irina Saparina, Mirella Lapata
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2511.10453v4 Announce Type: replace-cross Abstract: Large language models often respond to ambiguous requests by implicitly committing to one interpretation, frustrating users and creating safety risks when that interpretation is wrong. We propose generating a single structured response that e...

📖 Read original article


617. iLTM: Integrated Large Tabular Model ​

Author: David Bonet, Mar\c{c}al Comajoan Cara, Alvaro Calafell, Daniel Mas Montserrat, Alexander G. Ioannidis
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2511.15941v2 Announce Type: replace-cross Abstract: Tabular data underpins decisions across science, industry, and public services. Despite rapid progress, advances in deep learning have not fully carried over to the tabular domain, where gradient-boosted decision trees (GBDTs) remain a defaul...

📖 Read original article


618. VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging ​

Author: Ming Zhong, Yuanlei Wang, Liuzhou Zhang, Ruichuan An, Renrui Zhang, Hao Liang, Ming Lu, Ying Shen, Wentao Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2511.18121v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visual information. Unlike humans who naturally bridge details and high-level concepts, models tend to t...

📖 Read original article


619. Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM ​

Author: Adarsh Kumarappan, Ayushi Mehrotra
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2511.18721v4 Announce Type: replace-cross Abstract: The SmoothLLM defense provides a certification guarantee against jailbreaking attacks, but it relies on a strict "k-unstable" assumption that rarely holds in practice. This strong assumption can limit the trustworthiness of the provided safet...

📖 Read original article


620. Automating Deception: Scalable Multi-Turn LLM Jailbreaks ​

Author: Adarsh Kumarappan, Ananya Mujoo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves the way for a more significant one, to bypass safety alignments, pose a persistent threat to Large La...

📖 Read original article


621. Length-MAX Tokenizer for Language Models ​

Author: Dong Dong, Weijie Su
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2511.20849v2 Announce Type: replace-cross Abstract: We introduce a new tokenizer for language models that minimizes the average tokens per character, thereby reducing the number of tokens needed to represent text during training and to generate text during inference. Our method, which we refer...

📖 Read original article


622. Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics ​

Author: Yuan Gao, Jin Song, Yiyun Fei, Gongzhe Li, Ruigao Yang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2512.05098v2 Announce Type: replace-cross Abstract: In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily target portraits and artistic images, lacking a systematic evaluation of interior scenes. We introduce Sp...

📖 Read original article


623. Multilingual Agent-Based World Modeling for Social Science ​

Author: Xuan Zhang, Wenxuan Zhang, Anxu Wang, See-Kiong Ng, Yang Deng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.MA, cs.SI

arXiv:2512.07195v2 Announce Type: replace-cross Abstract: Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly monolingual without cross-lingual interaction, an essential property of real societies. We introduce MA...

📖 Read original article


624. The Theory of Strategic Evolution: Games with Endogenous Players and the Seven Laws of Strategic Replicators ​

Author: Kevin Vallier
Published: 8/11/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.CY, cs.MA, econ.TH

arXiv:2512.07901v4 Announce Type: replace-cross Abstract: Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. Rational players do not control their replication, and replicators do not choose strategically. Contemporary AI systems expos...

📖 Read original article


625. Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models ​

Author: Bj"orn Deiseroth, Max Henning H"oth, Kristian Kersting, Letitia Parcalabescu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2512.11614v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) relies on retrieved context to guide large language models (LLM), yet treats the retrieval as a heuristic rather than verifiable evidence -- leading to unsupported answers, hallucinations, and reliance on ...

📖 Read original article


626. Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality ​

Author: Sirui Chen, Jingji Chen, Siqi Zhu, Ziheng Jiang, Yanghua Peng, Xuehai Qian
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DC, cs.AI

arXiv:2512.20968v2 Announce Type: replace-cross Abstract: Distributed attention is essential for scaling large language models (LLMs) to long contexts, yet existing methods either have limited parallelism or incur high communication costs. Ulysses uses efficient all-to-all communication but cannot s...

📖 Read original article


627. TGIF: Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs ​

Author: Chenchen Lin, Sanbao Su, Rachel Luo, Yuxiao Chen, Yan Wang, Marco Pavone, Fei Miao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2601.03100v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) typically rely on a single late-layer feature from a frozen vision encoder, leaving the encoder's rich hierarchy of visual cues under-utilized. MLLMs still suffer from visually ungrounded hallucination...

📖 Read original article


628. IndexTTS 2.5 Technical Report ​

Author: Yunpei Li, Xun Zhou, Jinchao Wang, Lu Wang, Yong Wu, Siyi Zhou, Yiquan Zhou, Yining Wang, Yaogen Yang, Zhetao Hu, Shiyao Duan, Jiacheng Xu, Jingchen Shu, Bin Xia
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2601.03888v5 Announce Type: replace-cross Abstract: In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-based Text-to-Semantic (T2S) module and a non-autoregressive Semantic-to-Mel (S2M) module, which togethe...

📖 Read original article


629. ReMIND: Orchestrating Modular Large Language Models for Controllable Serendipity A REM-Inspired System Design for Emergent Creative Ideation ​

Author: Makoto Sato
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.07121v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used not only for problem solving but also for creative ideation; however, generating ideas that are both novel and coherent remains challenging. While high-temperature sampling can promote origin...

📖 Read original article


630. Layerwise goal-oriented adaptivity for neural ODEs: an optimal control perspective ​

Author: Michael Hinterm"uller, Michael Hinze, Denis Korolev
Published: 8/11/2026, 4:00:00 AM
Categories: math.OC, cs.AI

arXiv:2601.07397v2 Announce Type: replace-cross Abstract: In this work, we propose a novel layerwise adaptive construction method for neural network architectures. Our approach is based on a goal--oriented dual-weighted residual technique for the optimal control of neural differential equations. Thi...

📖 Read original article


631. Expert-Guided Multimodal Fusion for Unified Emotion and Sentiment Analysis ​

Author: Jiaqi Qiao, Xinran Li, Yifan Lyu, Xiujuan Xu, Liu Yu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2601.07565v2 Announce Type: replace-cross Abstract: Multimodal emotion understanding requires the integration of heterogeneous data sources, including text, audio, and visual modalities, while simultaneously addressing discrete emotion recognition and continuous sentiment analysis. We propose ...

📖 Read original article


632. LAUDE: LLM-Assisted Unit Test Generation and Debugging of Hardware DEsigns ​

Author: Deeksha Nandal, Riccardo Revalor, Soham Dan, Debjit Pal
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2601.08856v3 Announce Type: replace-cross Abstract: Unit tests are critical in the hardware design lifecycle to ensure that component design modules are functionally correct and conform to the specification before they are integrated at the system level. Thus developing unit tests targeting va...

📖 Read original article


633. RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation ​

Author: Yue Chang, Rufeng Chen, Zhaofan Zhang, Yi Chen, Yifan Tian, Sihong Xie
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO

arXiv:2601.10168v3 Announce Type: replace-cross Abstract: Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet current 3DSG construction methods suffer from semantic inconsistencies caused by noisy cross-image a...

📖 Read original article


634. Communication-efficient distributed hazard difference estimation for heterogeneous multi-site survival data ​

Author: Ziwen Wang, Siqi Li, Marcus Eng Hock Ong, Nan Liu
Published: 8/11/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG

arXiv:2601.14609v2 Announce Type: replace-cross Abstract: Multi-site collaboration can power survival models that no single hospital could fit alone, but privacy rules and protected computing environments block patient-level data sharing and the persistent server connections required by iterative fe...

📖 Read original article


635. Hybrid Mamba-Attention Neural Architecture for Channel Estimation ​

Author: Dianxin Luan, Chengsi Liang, Jie Huang, Zheng Lin, Kaitao Meng, John Thompson, Cheng-Xiang Wang, Ozgur Akan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.SP

arXiv:2601.17108v3 Announce Type: replace-cross Abstract: This paper proposes a hybrid Mamba-attention neural architecture to achieve improved channel estimation for orthogonal frequency-division multiplexing (OFDM) waveforms, particularly for configurations with a large number of subcarriers. By in...

📖 Read original article


636. SNR-Edit: Structure-Aware Noise Rectification for Inversion-Free Flow-Based Editing ​

Author: Lifan Jiang, Boxi Wu, Yuhang Pei, Tianrun Wu, Yongyuan Chen, Yan Zhao, Shiyu Yu, Deng Cai
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2601.19180v2 Announce Type: replace-cross Abstract: Inversion-free image editing using flow-based generative models challenges the prevailing inversion-based pipelines. However, existing approaches rely on fixed Gaussian noise to construct the source trajectory, leading to biased trajectory dy...

📖 Read original article


637. Temporal Sepsis Modeling: a Relational and Explainable-by-Design Framework ​

Author: Vincent Lemaire, N'edra Meloulli, Pierre Jaquet
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.21747v4 Announce Type: replace-cross Abstract: Sepsis remains one of the most complex and heterogeneous syndromes in intensive care. While deep learning models achieve competitive performance in early sepsis prediction, their decision processes often remain difficult to interpret clinical...

📖 Read original article


638. Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic ​

Author: Xingyu Zhao, Darsh Sharma, Rheeya Uppaal, Yiqiao Zhong
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.22510v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanistic studies reveal the discrepancy between LLMs and humans in skill compositions, the learning dynam...

📖 Read original article


639. Universal One-third Time Scaling in Learning Peaked Distributions ​

Author: Yizhou Liu, Ziming Liu, Cengiz Pehlevan, Jeff Gore
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2602.03685v3 Announce Type: replace-cross Abstract: Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origin remains debatable. Through systematic analysis of toy models and empirical evaluation of LLMs, we sho...

📖 Read original article


640. On the Infinite Width and Depth Limits of Predictive Coding Networks ​

Author: Francesco Innocenti, El Mehdi Achour, Rafal Bogacz
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE

arXiv:2602.07697v3 Announce Type: replace-cross Abstract: Predictive coding (PC) is a biologically plausible alternative to standard backpropagation (BP) that minimises an energy function with respect to network activities before updating weights. Recent work has improved the training stability of d...

📖 Read original article


641. SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online Transfer ​

Author: Nathan Samuel de Lara, Florian Shkurti
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2602.17632v3 Announce Type: replace-cross Abstract: Modern offline Reinforcement Learning (RL) methods find performant actor-critics, however, fine-tuning these actor-critics online with value-based RL algorithms typically causes immediate drops in performance. We provide evidence consistent w...

📖 Read original article


642. From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves ​

Author: Haritz Puerto, Haonan Li, Xudong Han, Timothy Baldwin, Iryna Gurevych
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2602.24210v3 Announce Type: replace-cross Abstract: Large reasoning models (LRMs) produce reasoning traces (RTs) that often contain sensitive information. These leaky thoughts are difficult to control and frequently violate explicit privacy directives. Because RTs can be exposed through prompt...

📖 Read original article


643. Attn-QAT: 4-Bit Attention With Quantization-Aware Training ​

Author: Peiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang, Will Lin, Wei Zhou, Hao Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.00040v3 Announce Type: replace-cross Abstract: Achieving reliable 4-bit attention is a prerequisite for end-to-end FP4 computation on emerging FP4-capable GPUs, yet attention remains the main obstacle due to FP4's tiny dynamic range and attention's heavy-tailed activations. This paper pre...

📖 Read original article


644. Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks ​

Author: Delip Rao, Chris Callison-Burch
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2603.00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduced to exact programmatic checks. Yet the underlying judges remain vulnerable to position bias, stoch...

📖 Read original article


645. LLMs Remember First, Forget Last: Dual-Process Interference in Large Language Models ​

Author: Sourav Chattaraj, Kanak Raj
Published: 8/11/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL

arXiv:2603.00270v3 Announce Type: replace-cross Abstract: Large language models can process millions of tokens, yet how they handle conflicting information within context remains poorly understood. From patient health logs tracking evolving vital signs to legal documents with superseding clauses, re...

📖 Read original article


646. PolypSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering ​

Author: Trong-Thang Pham, Loc Nguyen, Anh Nguyen, Hien V. Nguyen, Ngan Le
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2603.07066v2 Announce Type: replace-cross Abstract: Generative diffusion models are increasingly used for medical imaging data augmentation, but text prompting cannot produce causal training data. Re-prompting rerolls the entire generation trajectory, altering anatomy, texture, and background....

📖 Read original article


647. Adversarial Latent-State Training for Robust Policies in Partially Observable Domains ​

Author: Angad Singh Ahuja
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2603.07313v4 Announce Type: replace-cross Abstract: Robustness under latent distribution shift remains challenging in partially observable reinforcement learning. We formalize a focused setting where an adversary selects a hidden initial latent distribution before the episode, termed an advers...

📖 Read original article


648. Efficient Cross-View Localization in 6G Space-Air-Ground Integrated Network ​

Author: Min Hao, Yanbing Xu, Maoqiang Wu, Jinglin Huang, Chen Shang, Jiacheng Wang, Ruichen Zhang, Jiawen Kang, Dusit Niyato, Zhu Han, Wei Ni
Published: 8/11/2026, 4:00:00 AM
Categories: cs.NI, cs.AI

arXiv:2603.11398v2 Announce Type: replace-cross Abstract: Recently, visual localization has become an important supplement to improve localization reliability, and cross-view approaches can greatly enhance coverage and adaptability. Meanwhile, future 6G will enable a globally covered mobile communic...

📖 Read original article


649. Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages ​

Author: Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2603.12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre...

📖 Read original article


650. Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification ​

Author: Zenan Li, Ziran Yang, Deyuan He, Haoyu Zhao, Andrew Zhao, Shange Tang, Kaiyu Yang, Aarti Gupta, Zhendong Su, Chi Jin
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2603.19329v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can generate plausible code but offer limited guarantees of correctness. Formally verifying that implementations satisfy specifications requires constructing machine-checkable proofs, a task that remains beyond cu...

📖 Read original article


651. SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs ​

Author: Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence, Rohan Bhakta, Izzy F. Thomas, Mingyun Cao, Chung-Hao Tsai, Zihao Zhou, Yidong Zhao, Hao Liu, Alessandro Marinoni, Alexey Arefiev, Rose Yu
Published: 8/11/2026, 4:00:00 AM
Categories: physics.comp-ph, cs.AI, cs.DC, cs.LG

arXiv:2603.20253v3 Announce Type: replace-cross Abstract: Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To ad...

📖 Read original article


652. ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention ​

Author: Xinyan Wang, Xiaogeng Liu, Ming Pei, Chaowei Xiao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2603.22016v3 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant verification, repeated attempts, or unnecessary exploration that wastes computation and can even overturn ...

📖 Read original article


653. SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection ​

Author: Kexian Tang, Jiani Wang, Shaowen Wang, Kaifeng Lyu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2603.22213v2 Announce Type: replace-cross Abstract: While large language models (LLMs) are pretrained on massive amounts of data, their knowledge coverage remains incomplete in specialized, data-scarce domains, motivating extensive efforts to study synthetic data generation for knowledge injec...

📖 Read original article


654. A Sobering Look at Tabular Data Generation via Probabilistic Circuits ​

Author: Davide Scassola, Dylan Ponsford, Adri'an Javaloy, Sebastiano Saccani, Luca Bortolussi, Henry Gouk, Antonio Vergari
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.23016v2 Announce Type: replace-cross Abstract: Tabular data is more challenging to generate than text and images, due to its heterogeneous features and much lower sample sizes. On this task, diffusion-based models are the current state-of-the-art (SotA) model class, achieving almost perfe...

📖 Read original article


655. Explaining, Verifying, and Aligning Semantic Hierarchies in Vision-Language Model Embeddings ​

Author: Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein, Edgar Heinert, Annika M"utze, Marvin Keller, Sparsh Tiwari, Georgii Mikriukov, Diedrich Wolter, Jae Hee Lee, Matthias Rottmann
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.26798v2 Announce Type: replace-cross Abstract: Vision-language model (VLM) encoders such as CLIP enable strong retrieval and zero-shot classification in a shared image-text embedding space, yet the semantic organization of this space is rarely inspected. We present a post-hoc framework to...

📖 Read original article


656. Critic-Free Deep Reinforcement Learning for Maritime Coverage Path Planning on Irregular Hexagonal Grids ​

Author: Carlos S. Sep'ulveda, Gonzalo A. Ruz
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE, cs.RO

arXiv:2603.28385v2 Announce Type: replace-cross Abstract: Maritime surveillance missions, such as search and rescue and environmental monitoring, rely on the efficient allocation of sensing assets over vast and geometrically complex areas. Traditional Coverage Path Planning (CPP) approaches depend o...

📖 Read original article


657. To Memorize or to Retrieve: Scaling the Interaction Between Pretraining and Retrieval ​

Author: Karan Singh, Michael Yu, Varun Gangal, Zhuofu Tao, Sachin Kumar, Emmy Liu, Steven Y. Feng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2604.00715v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) improves language model (LM) performance by providing relevant context at test time for knowledge-intensive situations. In this work, we systematically study the trade-off between pretraining and retrieval...

📖 Read original article


658. Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines ​

Author: Jingjie Ning, Xueqi Li, Chengyu Yu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL

arXiv:2604.01029v2 Announce Type: replace-cross Abstract: Multi-LLM revision pipelines, in which a second model reviews and improves a draft produced by a first, are widely assumed to derive their gains from genuine error correction. We question this assumption with a controlled decomposition experi...

📖 Read original article


659. Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding ​

Author: Tao Jin, Phuong Minh Nguyen, Naoya Inoue
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.02047v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates large language model inference by drafting multiple candidate tokens and verifying them in a single forward pass. Candidates are organized as a tree: deeper trees accept more tokens per step, but adding depth ...

📖 Read original article


660. CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge ​

Author: Mete Ismayilzada, Renqing Cuomao, Daniil Yurshevich, Anna Sotnikova, Lonneke van der Plas, Antoine Bosselut
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.03374v2 Announce Type: replace-cross Abstract: Creative problem-solving requires combining multiple cognitive abilities, including logical reasoning, lateral thinking, analogy-making, and commonsense knowledge, to discover insights that connect seemingly unrelated pieces of information. H...

📖 Read original article


661. Large Language Models Align with the Human Brain during Creative Thinking ​

Author: Mete Ismayilzada, Simone A. Luchini, Abdulkadir Gokce, Badr AlKhamissi, Antoine Bosselut, Antonio Laverghetta Jr., Lonneke van der Plas, Roger E. Beaty
Published: 8/11/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.CL

arXiv:2604.03480v2 Announce Type: replace-cross Abstract: Creative thinking is a fundamental aspect of human cognition, and divergent thinking-the capacity to generate novel and varied ideas-is widely regarded as its core generative engine. Large language models (LLMs) have recently demonstrated imp...

📖 Read original article


662. Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders ​

Author: Atahan Dokme, Sriram Vishwanath
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.03919v2 Announce Type: replace-cross Abstract: We present the first systematic study of Sparse Autoencoders (SAEs) on video representations. Standard SAEs decompose video into interpretable, monosemantic features but destroy temporal coherence: hard TopK selection produces unstable featur...

📖 Read original article


663. Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning in Agents ​

Author: Neharika Jali, Anupam Nayak, Gauri Joshi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.05164v3 Announce Type: replace-cross Abstract: As LLM reasoning performance plateaus, improving inference-time compute efficiency is crucial to mitigate overthinking and long thinking traces even for simple queries. Prior approaches including length regularization, adaptive routing, and d...

📖 Read original article


664. SALLIE: Generation-Free Hidden-State Detection of Jailbreaks and Prompt Injections Across Text and Vision ​

Author: Guy Azov, Ofer Rivlin, Guy Shtar
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2604.06247v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) are vulnerable to jailbreaks and prompt injections delivered through text or images. Existing defenses often narrow threat coverage or add inference cost through input transformat...

📖 Read original article


665. Continual Visual Anomaly Detection on the Edge: Benchmark and Efficient Solutions ​

Author: Manuel Barusco, Francesco Borsatti, David Petrovic, Davide Dalle Pezze, Gian Antonio Susto
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.06435v2 Announce Type: replace-cross Abstract: Visual Anomaly Detection (VAD) is a critical task for many applications including industrial inspection and healthcare. While VAD has been extensively studied, two key challenges remain largely unaddressed in conjunction: edge deployment, whe...

📖 Read original article


666. Multi-objective Evolutionary Merging Enables Efficient Reasoning Models ​

Author: Mario Iacobelli, Adrian Robert Minut, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Iacopo Masi, Emanuele Rodol`a
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.06465v2 Announce Type: replace-cross Abstract: Reasoning models achieve strong performance on complex problems by leveraging long chains of thought, but this deliberate reasoning incurs substantial inference-time cost. The Long-to-Short (L2S) reasoning problem seeks to preserve accuracy w...

📖 Read original article


667. TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories ​

Author: Yen-Shan Chen, Sian-Yao Huang, Cheng-Lin Yang, Yun-Nung Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG, cs.SE

arXiv:2604.07223v2 Announce Type: replace-cross Abstract: As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final outputs to intermediate execution traces. While safety guardrails are well-benchmarked for natural languag...

📖 Read original article


668. In-context superposition: human-like working memory interference in large language models ​

Author: Hua-Dong Xiong, Li Ji-An, Jiaqi Huang, Robert C. Wilson, Kwonjoon Lee, Xue-Xin Wei
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.09670v2 Announce Type: replace-cross Abstract: Intelligent systems must maintain and manipulate task-relevant information online to adapt to dynamic environments and changing goals. This capacity, known as working memory, is fundamental to human reasoning and intelligence. Despite their r...

📖 Read original article


669. Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification ​

Author: Patrick Lutz, Themistoklis Haris, Arjun Chandra, Aditya Gangrade, Venkatesh Saligrama
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.11613v4 Announce Type: replace-cross Abstract: Transformers can perform in-context classification from a few labeled examples, yet the inference-time algorithm remains opaque. We study multi-class linear classification in the hard no-margin regime and make the computation identifiable by ...

📖 Read original article


670. Fairness is Not Flat: Geometric Phase Transitions Against Shortcut Learning ​

Author: Nicolas Rodriguez-Alvarez (Instituto de Educacion Secundaria Parquesol, Valladolid, Spain)
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.11704v2 Announce Type: replace-cross Abstract: Deep Neural Networks are highly susceptible to shortcut learning, frequently memorizing low-dimensional spurious correlations instead of underlying causal mechanisms. This phenomenon not only degrades out-of-distribution robustness but also i...

📖 Read original article


671. The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models ​

Author: Akshay Paruchuri, Ishan Chatterjee, Henry Fuchs, Ehsan Adeli, Piotr Didyk
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV

arXiv:2604.14363v2 Announce Type: replace-cross Abstract: Multimodal language models systematically underperform on visual perception tasks, yet the structure underlying this failure remains poorly understood. We propose centroid replacement, mapping tokens to their nearest K-means centroid and remo...

📖 Read original article


672. Controllable Video Object Insertion via Multi-View Priors ​

Author: Qi Xia, Peishan Cong, Yichen Yao, Ziyi Wang, Yaoqin Ye, Yuexin Ma
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2604.14556v2 Announce Type: replace-cross Abstract: Video object insertion places a user-specified object in an existing dynamic scene. Existing methods typically condition generation on text or a single reference image. Consequently, object appearance is underconstrained under viewpoint chang...

📖 Read original article


673. Switching Theory for Q-Learning ​

Author: Donghwan Lee
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SY, eess.SY

arXiv:2604.19569v5 Announce Type: replace-cross Abstract: Q-learning is a fundamental algorithmic primitive in reinforcement learning. This paper develops a new framework for analyzing constant step-size tabular Q-learning from a switching linear system (SLS) viewpoint. In particular, we derive a st...

📖 Read original article


674. Hybrid Policy Distillation for LLMs ​

Author: Wenhong Zhu, Ruobing Xie, Rui Wang, Pengfei Liu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2604.20244v2 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is a powerful paradigm for compressing large language models (LLMs), whose effectiveness depends on intertwined choices of divergence direction, optimization strategy, and data regime. We break down the design of e...

📖 Read original article


675. Model Predictive Control of Hybrid Dynamical Systems ​

Author: Ricardo G. Sanfelice, Berk Altin
Published: 8/11/2026, 4:00:00 AM
Categories: math.OC, cs.AI, cs.RO, cs.SY, eess.SY, math.DS

arXiv:2604.21989v2 Announce Type: replace-cross Abstract: The problem of controlling hybrid dynamical systems using model predictive control (MPC) is formulated and sufficient conditions for asymptotic stability of a set are provided. Hybrid dynamical systems are modeled in terms of hybrid equations...

📖 Read original article


676. From Local to Cluster: A Unified Framework for Causal Discovery with Latent Variables ​

Author: Zongyu Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.22416v3 Announce Type: replace-cross Abstract: Latent variables pose a fundamental obstacle to both causal discovery and inference. Local approaches exploiting direct neighborhood relations provide little beyond immediate dependencies. Cluster-level methods, though capable of broader reas...

📖 Read original article


677. UGAF-ITS: A Standards Harmonization Framework and Validation Tool for Multi-Framework AI Governance in Distributed Intelligent Transportation Systems ​

Author: Talal Ashraf Butt, Muhammad Iqbal, Razi Iqbal
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2604.22789v2 Announce Type: replace-cross Abstract: Organizations deploying AI-enabled Intelligent Transportation Systems face fragmented governance: ISO/IEC~42001 demands a certifiable management system, the EU AI Act imposes binding high-risk obligations from August~2026, and the NIST AI Ris...

📖 Read original article


678. Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards ​

Author: Taha Hammadia, Lucas Rea, Ahmad Mohammad Saber, Amr Youssef, Deepa Kundur
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2604.23341v3 Announce Type: replace-cross Abstract: The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-making but exposes new vulnerabilities to prompt-based adversarial attacks. This paper evaluates the risk ...

📖 Read original article


679. Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective ​

Author: Jiaming Yang, Chenwei Tang, Liangli Zhen, Jiancheng Lv
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, math.IT

arXiv:2604.25975v2 Announce Type: replace-cross Abstract: Key-Value (KV) caching is essential for large language model inference, yet its memory overhead poses a critical bottleneck for long-context generation. Existing eviction policies predominantly rely on empirical heuristics, lacking a rigorous...

📖 Read original article


680. Culturally Situated AI Safety for Youth: Saudi Arabian Perspectives of Youth, Parents and Teachers ​

Author: Aljawharah Alzahrani, Tanusree Sharma
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.CY, cs.ET

arXiv:2604.26494v2 Announce Type: replace-cross Abstract: Generative AI tools are widely used by youth and have introduced new privacy and safety challenges. While prior research has explored youths safety in GenAI within a Western context, it often overlooks the cultural, religious, and social dime...

📖 Read original article


681. Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation ​

Author: Shouren Wang, Wang Yang, Chuang Ma, Debargha Ganguly, Vikash Singh, Chaoda Song, Xinpeng Li, Xianxuan Long, Vipin Chaudhary, Xiaotian Han
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2604.27201v3 Announce Type: replace-cross Abstract: Hybrid-thinking language models expose explicit /think and /no_think modes, but current designs do not separate them cleanly. Even in /no_think mode, models often emit long and self-reflective responses, causing reasoning leakage. Existing wo...

📖 Read original article


682. The Safety-Aware Denoiser for Text Diffusion Models ​

Author: Amman Yusuf, Zhejun Jiang, Mijung Park
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.08116v3 Announce Type: replace-cross Abstract: Recent work on text diffusion models offers a promising alternative to autoregressive generation, but controlling their safety remains underexplored. Existing safety approaches are geared toward autoregressive models and typically rely on pos...

📖 Read original article


683. In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores ​

Author: Zeyu Tang, Sang T. Truong, Deonna Owens, Shreyas Sharma, Yibo Jacky Zhang, Brando Miranda, Sanmi Koyejo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY

arXiv:2605.12530v2 Announce Type: replace-cross Abstract: LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally unreliable: surface-level prompt construction choices, although ...

📖 Read original article


684. Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy ​

Author: Adarsh Kumarappan, Ananya Mujoo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.12991v3 Announce Type: replace-cross Abstract: LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widely attributed to RLHF-induced sycophancy. We test this attribution across four model families...

📖 Read original article


685. HEART: Exploiting Head Heterogeneity in Sparse Attention for Video Diffusion ​

Author: Xuzhe Zheng, Yuexiao Ma, Jing Xu, Xiawu Zheng, Rongrong Ji, Fei Chao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.14513v2 Announce Type: replace-cross Abstract: Sparse attention accelerates video diffusion by allowing each attention head to focus on only a small subset of interactions. Existing methods already construct head-specific sparse patterns conditioned on the input. However, we find that the...

📖 Read original article


686. DeltaPrompts: Escaping the Zero-Delta Trap in Multimodal Distillation ​

Author: Jaehun Jung, Hyunwoo Kim, Brandon Cui, Ximing Lu, David Acuna, Prithviraj Ammanabrolu, Yejin Choi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2605.15532v3 Announce Type: replace-cross Abstract: Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process are typically chosen via simple heuristics or aggregated from off-the-shelf datasets. We reveal a critica...

📖 Read original article


687. Post-Deployment Accountability in AI Governance: A Cross-Regulatory Empirical Analysis of AI Incidents ​

Author: Ummara Mumtaz, Rabi Noor, Summaya Mumtaz
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2605.16281v3 Announce Type: replace-cross Abstract: Post-deployment accountability has become central to AI governance, yet little empirical evidence shows whether monitoring, incident reporting, and impact assessment obligations are visible when AI systems fail. This study analyzes real-world...

📖 Read original article


688. Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap ​

Author: Aysa Xuemo Fan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2605.16283v3 Announce Type: replace-cross Abstract: Large-scale AI deployment data and controlled learning experiments characterize different consequences of the same technology. Deployment telemetry shows that AI use is concentrated in skilled work and frequently supports immediate task perfo...

📖 Read original article


689. Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control ​

Author: Rohith Uppala
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2605.18414v2 Announce Type: replace-cross Abstract: Large language models increasingly operate as autonomous agents that select and invoke tools from large registries. We identify a critical gap: when unauthorized tools are visible in an agent's context, models select them in adversarial scena...

📖 Read original article


690. Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance ​

Author: Jing Chen, Shixiang Pan, Yujie Fan, Haocheng Ye, Haitao Xu, Wenqiang Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.18793v2 Announce Type: replace-cross Abstract: Accurate spatiotemporal pattern analysis is critical in fields such as urban traffic, meteorology, and public health monitoring. However, existing methods face performance bottlenecks, typically yielding only incremental gains and often exhib...

📖 Read original article


691. How to Build Marcus's Algebraic Mind: Algebro-Deterministic Substrate over Galois Fields ​

Author: Hiroyuki Chuma, Kanji Otsuk, Yoichi Sato
Published: 8/11/2026, 4:00:00 AM
Categories: cs.NE, cs.AI

arXiv:2605.21379v3 Announce Type: replace-cross Abstract: In The Algebraic Mind (2001), Marcus held that any adequate cognitive architecture needs operations over variables, recursively structured representations, and an individual/kind distinction, and that multilayer perceptrons support none of th...

📖 Read original article


692. Don't Retrain, Just Reuse: Recovering Dual-Target Molecules from Single-Target Diffusion Models ​

Author: Qingyuan Zeng, Pengxiang Cai, Zixin Guan, Ziyang Chen, Anglin Liu, Xinyao Lai, Jintai Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.25681v2 Announce Type: replace-cross Abstract: Designing a single molecule that modulates two targets is a promising strategy for polypharmacology, but it remains substantially harder than standard single-target generation because one candidate must satisfy two binding requirements while ...

📖 Read original article


693. Beyond Questions: Evaluating LLM's Knowledge Expression ​

Author: Luca Giordano, Simon Razniewski
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2605.26937v2 Announce Type: replace-cross Abstract: Parametric knowledge in large language models (LLMs) is a cornerstone of their success, yet remains poorly understood. Existing knowledge benchmarks typically rely on predefined questions (e.g., "What is the birth date of M.L. King?"), evalua...

📖 Read original article


694. Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation ​

Author: Zhiyuan Yang, Jiahao Cheng, Vincent Quoc-Huy Trinh, Mahdi S. Hosseini
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2605.30716v2 Announce Type: replace-cross Abstract: Generating clinically useful pathology reports for pathology cases from whole-slide images (WSIs) is challenging due to gigapixel resolution, long visual-token sequences, and the complexity of case-level reasoning, where a single case may con...

📖 Read original article


695. SimSD: Simple Speculative Decoding in Diffusion Language Models ​

Author: Junxia Cui, Haotian Ye, Runchu Tian, Hongcan Guo, Jinya Jiang, Haoru Li, Chaojie Ren, Yiming Huang, Kaijie Zhu, Zhongkai Yu, Kun Zhou, Jingbo Shang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.02544v2 Announce Type: replace-cross Abstract: Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster inference through parallel or blockwise decoding. However, their masked language modeling formulation remain...

📖 Read original article


696. dots.tts Technical Report ​

Author: Shi Lian, Changtao Li, Bohan Li, Hankun Wang, Da Zheng, Junfeng Tian, Yufeng Ma, Colin Zhang, Kai Yu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI, eess.AS

arXiv:2606.07080v2 Announce Type: replace-cross Abstract: We present dots$.$tts, a 2B-parameter continuous autoregressive text-to-speech (TTS) foundation model that models speech in a continuous latent space. Compared with existing continuous autoregressive models, our key innovations are threefold....

📖 Read original article


697. Enhancing AI Interpretability with Localised Architectures ​

Author: Ian Seet, Jonas Bozenhard, Simon Ostermann
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.07998v3 Announce Type: replace-cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs), raise concerns over the interpretability, safety and sustainability of these large and opaque AI models. The power of such architectures is derived not only f...

📖 Read original article


698. Contemporary AI lacks the imagination to diverge or negate in science ​

Author: Honglin Bao, Siyang Wu, Xiao Liu, Sida Li, Shiyun Cao, James A. Evans
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2606.08251v3 Announce Type: replace-cross Abstract: Bold claims that AI will accelerate scientific discovery have raced ahead of evidence from working scientists, yet large-scale, scientist-in-the-loop evidence is scarce. Here we mount the largest evaluation to date, inviting authors of 121,64...

📖 Read original article


699. Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation ​

Author: Bruce Changlong Xu, Adarsh Kumarappan, Mu Zhou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.ET

arXiv:2606.09864v2 Announce Type: replace-cross Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact. In this study, we explore ...

📖 Read original article


700. Anomaly Detection and Root Cause Analysis for Microservice Systems ​

Author: Luan Pham
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SE, cs.AI

arXiv:2606.09942v3 Announce Type: replace-cross Abstract: Microservice systems are widely used to build cloud applications, yet their complexity makes failures inevitable, degrading user experience and causing economic loss. Automated anomaly detection and root cause analysis (RCA) are now active re...

📖 Read original article


701. Towards a Bridge Layer Between Bibliographic and Formalized Mathematical Knowledge ​

Author: A. Mayeux
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.LO

arXiv:2606.11430v3 Announce Type: replace-cross Abstract: Mathematical knowledge is split between bibliographic databases (e.g., MathSciNet, zbMATH Open) and formal proof libraries (e.g., Lean's mathlib), preventing unified access to published results and their formalizations. We propose a relationa...

📖 Read original article


702. Two-Layer Linear Auto-Regressive Models Estimate Latent States ​

Author: Yahya Sattar, Sunmook Choi, Leo Maynard-Zhang, Yassir Jedra, Maryam Fazel, Sarah Dean
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SY, eess.SY, math.OC, stat.ML

arXiv:2606.12691v2 Announce Type: replace-cross Abstract: Auto-regressive models have emerged as powerful tools for sequential data, from language to video. Understanding how and why these models learn latent representations remains an open theoretical question. In this work, we demonstrate that whe...

📖 Read original article


703. SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model ​

Author: Xiaoxin Lu, Ranran Haoran Zhang, Rui Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2606.14574v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmarks evaluate whether LLM-generated plans execute successfully, they overlook a critical type of failure:...

📖 Read original article


704. Learning aligned EEG representations with subject-specific encoders ​

Author: Bruna J. Lopes, Gabriel Schwartz, Sylvain Chevallier, Raphael Y. de Camargo, Bruno Aristimunha
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.16462v2 Announce Type: replace-cross Abstract: Cross-subject EEG decoding promises more training data, but it also exposes neural networks to strong inter-subject distribution shifts. We study whether task supervision and architecture alone can learn subject-aligned representations. We re...

📖 Read original article


705. OmniV2X: A Generative Foundation Planner for Efficient End-to-End Cooperative Driving ​

Author: Juntong Peng, Juanwu Lu, Yupeng Zhou, Can Cui, Yaobin Chen, Ziran Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2606.21165v2 Announce Type: replace-cross Abstract: We present OmniV2X, a generative foundation model for vehicle-to-everything (V2X) cooperative driving. The model directly interprets independent context sequences comprising multi-modal and multi-agent observations. The new design mitigates t...

📖 Read original article


706. Unsupervised Disentanglement Without Compromises : How Functional Orthogonality Enforces Identifiability ​

Author: Mathieu Cyrille Simon, Pascal Frossard, Christophe De Vleeschouwer
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.21385v2 Announce Type: replace-cross Abstract: This paper explores unsupervised disentangled representation learning from a functional perspective. We define latent concepts as factors that influence observations through locally orthogonal directions, formalized as an orthogonality constr...

📖 Read original article


707. Decodable but Not Faithful: Coupling Natural-Language Rationales to Programmatic Verifiers ​

Author: Vatsal Ananthula, Adarsh Kumarappan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2606.21678v2 Announce Type: replace-cross Abstract: Language models can generate plausible rationales for their predictions, but these explanations may not faithfully represent the model's internal reasoning. We propose verifier-coupled reasoning, a framework that inserts inline claims into re...

📖 Read original article


708. Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior ​

Author: Vyom Agarwal, Mokshda Gangrade, Siddharth Pal, Jerry Wu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2606.22790v2 Announce Type: replace-cross Abstract: In this paper, we investigate the tradeoffs between compute allocation and model performance for two speech processing tasks: Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). We propose a unified framework that analyze...

📖 Read original article


709. The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection ​

Author: Nicolas M. M"uller, Pascal Debus
Published: 8/11/2026, 4:00:00 AM
Categories: cs.SD, cs.AI

arXiv:2606.23335v2 Announce Type: replace-cross Abstract: Provenance watermarking is increasingly treated as a safeguard for synthetic speech, whether built directly into speech-generation models such as Chatterbox, provided through dedicated techniques such as AudioSeal, or deployed by commercial p...

📖 Read original article


710. ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory ​

Author: Habibullah Akbar
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.25156v3 Announce Type: replace-cross Abstract: Native length extrapolation remain a weakly solvable problem in language modeling due to trade-off balancing between exact retrieval fidelity, long-document likelihood, and inference efficiency. We present ATMA as a disciplined diagnostic for...

📖 Read original article


711. EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis ​

Author: Huaqiu Li, Jiahao Wang, Sijia Cai, Hualian Sheng, Bing Deng, Jieping Ye, Wenhan Luo
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2606.25465v2 Announce Type: replace-cross Abstract: While image stylization has been studied extensively, video stylization remains a critical and largely unsolved challenge in the field of intelligent content creation. Existing methods, usually utilizing a reference image as the style prior, ...

📖 Read original article


712. Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction ​

Author: Chenguang Wang, Ming Li, Xinyue Zeng, Zhuochun Li, Hong Jiao, Tianyi Zhou, Dawei Zhou
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.LG

arXiv:2606.28186v3 Announce Type: replace-cross Abstract: Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on costly human calibration or item-level textual representations,...

📖 Read original article


713. DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training ​

Author: Haisen Luo, Yiwei Liu, Haoning Wang, Dan Liu, Junxi Yin, Haotian Wang, Lei Zhang, Xiaoyu Tian, Shuaiting Chen, Yuansheng Song, Baoyan Guo, Xiongfei Yan, Bolan Yang, Chengwei Liu, Ming Cui, Jiong Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.30345v4 Announce Type: replace-cross Abstract: Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in complex reasoning tasks. Existing self-distillation and reinforcement learning methods lack explicit mechanis...

📖 Read original article


714. Can LLMs Rank? A Tale of Triads and Triage ​

Author: Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler, Sanmay Das
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI

arXiv:2606.30412v2 Announce Type: replace-cross Abstract: From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered as judges of consequential decisions that require ranking people for scarce resources. Ranking large g...

📖 Read original article


715. Learning Cardiac Motion Priors for Implicit Neural Representations ​

Author: Andrew Bell, George Webber, Steffen E Petersen, Andrew P King, Muhummad Sohaib Nazir, Alistair Young
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.00955v3 Announce Type: replace-cross Abstract: Implicit neural representations (INRs) are well suited to cardiac motion estimation, providing continuous, compact representations of motion fields. However, fitting an INR to each image sequence is time-consuming and sensitive to the optimis...

📖 Read original article


716. Safeguarding LLM Agents from Misalignment through Provenance Analysis ​

Author: Yining She, Yiliang Liang, Eunsuk Kang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.01236v2 Announce Type: replace-cross Abstract: As LLM agents gain increasing access to powerful tools, ensuring that their actions align with the user's intent becomes critical. When an agent's proposed action deviates from that intent---a phenomenon called misalignment---it may cause har...

📖 Read original article


717. NeuroBridge: Bridging Multi-Task MRI Knowledge for Neurodegenerative Disease Diagnosis ​

Author: Mengyu Li, Guoyao Shen, Chad W. Farris, Xin Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2607.01401v2 Announce Type: replace-cross Abstract: Accurate MRI-based identification of Alzheimer's disease (AD), mild cognitive impairment (MCI), and related dementias remains challenging because disease-related structural changes are often subtle and heterogeneous. We developed NeuroBridge,...

📖 Read original article


718. Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention ​

Author: Siyu Ding, Mingchuan Ma, Jiabo Tong, Xingrun Xing, Ziming Wang, Guoqi Li
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.04422v2 Announce Type: replace-cross Abstract: Recent NVFP4 pretraining work has primarily optimized Transformer linear projections, leaving persistent optimizer states, optimizer computation, and low-precision attention forward--backward paths less explored. We present \textbf{Full-Stack...

📖 Read original article


719. UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents ​

Author: Niu Lian, Tongbo Chen, Zhehao Yu, Chengzhen Duan, Fazhan Liu, Hui Liu, Pei Fu, Jian Luan, Heng Qu, Shu-Tao Xia, Jinpeng Wang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CV, cs.LG, cs.MM

arXiv:2607.04425v2 Announce Type: replace-cross Abstract: Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-platform interaction. However, unified multi-platform GUI learning remains challenging: high-quality cro...

📖 Read original article


720. RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ​

Author: Qiang Liu, Taian Guo, Ruizhi Qiao, Xing Sun
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.04713v2 Announce Type: replace-cross Abstract: Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks. However, in long-horizon, multi-turn tasks characterized by sparse outcome rewards, directly training with ou...

📖 Read original article


721. Safe Bayesian Optimization with Counterfactual Policies ​

Author: Katherine Avery, Bruno Castro da Silva, David Jensen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.05620v2 Announce Type: replace-cross Abstract: In many decision-making settings, new interventions are acceptable only if they do not reduce outcomes below some established threshold. For example, in clinical medicine, new treatments are often acceptable only if they do not worsen outcome...

📖 Read original article


722. Digital Fragmentation and Generative AI Use Across 103 Million Application Events ​

Author: Sumer S. Vaid, Ashley V. Whillans
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.ET, stat.AP

arXiv:2607.06681v2 Announce Type: replace-cross Abstract: Knowledge workers switch between applications thousands of times per day, spending nearly a tenth of the work year transitioning between digital applications in a process called digital fragmentation. Whether this fragmentation reflects who a...

📖 Read original article


723. EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins ​

Author: Joshua Pickard, Wei Qi, Na Li, Ann Woolley, Lisa Cosimi, Roy Kishony, Deborah Hung
Published: 8/11/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, cs.SY, eess.SY, math.OC

arXiv:2607.08793v4 Announce Type: replace-cross Abstract: Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches learn fixed strategies for sepsis treatment, limiting adaptability to changing clinical objectives during...

📖 Read original article


724. ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception ​

Author: Weichen Zhang, Shiquan Yu, Yinan Zhu, Peizhi Tang, Shilong Ji, Zhiyuan Deng, Tianyi Lyu, Haoyang Wang, Xin Zeng, Chen Gao, Yong Li, Xinlei Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.RO, cs.AI

arXiv:2607.10180v2 Announce Type: replace-cross Abstract: We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. The benchmark decomposes active perception into three hierarchical tasks: Aerial Embodied Question An...

📖 Read original article


725. Instruction Set and Language for Hypergraphs ​

Author: Mario Pascual-Gonzalez, Ezequiel Lopez-Rubio
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.PL

arXiv:2607.10194v2 Announce Type: replace-cross Abstract: We present IsalHG, a method for representing the structure of any finite, connected hypergraph of bounded hyperedge arity as a string over a compact instruction alphabet $\Sigma_{\mathrm{HG}}$. The encoding is executed by a small virtual mach...

📖 Read original article


726. A Physics-Inspired Classical Digital Twin of Cortical Dynamics: A Band-Stratified Metriplectic Port-Hamiltonian Neural Network Learned from Brain-Computer-Interface EEG ​

Author: Dibakar Sigdel
Published: 8/11/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI

arXiv:2607.10439v4 Announce Type: replace-cross Abstract: We present a physics-inspired classical digital twin of brain-computer- interface (BCI) data: a graph neural network constrained to a band-stratified, metriplectic port-Hamiltonian form, with parameters learned from scalp EEG recorded during ...

📖 Read original article


727. Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update ​

Author: Daocheng Fu, Rong Wu, Yu Yang, Jianbiao Mei, Licheng Wen, Pinlong Cai, Xuemeng Yang, Yong Liu, Botian Shi, Yu Qiao
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.11505v2 Announce Type: replace-cross Abstract: Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward behaviors from policy exploration. While on-policy distillation alleviates this by consolidating independ...

📖 Read original article


728. Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation ​

Author: Jasmine Brazilek, Maheep Chaudhary, Zoe Lu, Miles Tidmarsh
Published: 8/11/2026, 4:00:00 AM
Categories: cs.MA, cs.AI, cs.CR

arXiv:2607.15434v5 Announce Type: replace-cross Abstract: Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. ...

📖 Read original article


729. LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4 ​

Author: Mobina Kashaniyan, Amirhossein Ghassemi, Nasser Mozayani
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.15509v2 Announce Type: replace-cross Abstract: We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition. Each large language model independ...

📖 Read original article


730. An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism ​

Author: Mobina Kashaniyan, Mehrdad Ashtiani, Amirhossein Ghassemi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.PF

arXiv:2607.15511v2 Announce Type: replace-cross Abstract: Serverless computing provides automatic resource management and pay-per-use execution, but effective autoscaling remains challenging because of dynamic workloads, cold-start latency, and dependencies among functions. We present a dependency-a...

📖 Read original article


731. Understanding Reasoning from Pretraining to Post-Training ​

Author: Jingyan Shen, Ang Li, Salman Rahman, Yifan Sun, Micah Goldblum, Matus Telgarsky, Pavel Izmailov
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.16097v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions ...

📖 Read original article


732. OpenMHC: Accelerating the Science of Wearable Foundation Models ​

Author: Narayan Schuetz, Yuze Bai, Lianggang Pan, Edgar Eggert, Favour Nerrise, Juan Delgado-SanMartin, Max Rosenblattl, Milana Gurbanova, Mohammad Asadi, Anders Johnson, Paul Schmiedmayer, Dennis Wang, Allan Lawrie, Daniel Seung Kim, Xin Liu, Akshay Paruchuri, Ehsan Adeli, Euan Ashley, Kelly W. Zhang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.16235v3 Announce Type: replace-cross Abstract: Mobile and wearable devices offer an unprecedented opportunity for continuous, passive health monitoring and active health coaching. However, the largest wearable datasets are not publicly available for research, and leading wearable foundati...

📖 Read original article


733. LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models ​

Author: Zeyu Xu, Xingzhong Hou, Pengkai Guo, Siling Lin, Xiao Xu, Menghua Zhai, Haoyu Chen, Yunke Zhang, Fei Huang
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.16305v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Experts (MoE) models to improve performance limits deployment in resource-constrained environments due to the...

📖 Read original article


734. Cost Accounting for Reactive Computational Graphs: Exhaustive Sweeps, Sequential Mutation, and the Backward-Locality Gap ​

Author: Abdallah Khemais (ISITCOM, University of Sousse)
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.18323v2 Announce Type: replace-cross Abstract: Exhaustive site-by-site interventions on a neural network's computational graph -- activation-patching sweeps, circuit-discovery searches, systematic ablation studies -- mutate the graph at every candidate site, and their cost is dominated by...

📖 Read original article


735. Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models ​

Author: Gautam Rajendrakumar Gare, Jia Shi, Zhiqiu Lin, Deepak Pathak, John Galeotti, Deva Ramanan
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV

arXiv:2607.18695v2 Announce Type: replace-cross Abstract: A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP with the resulting descriptors. We show that these descriptors carry little visual evidence of their own:...

📖 Read original article


736. Riemannian Deep Learning: Modules, Networks, and Geometries ​

Author: Ziheng Chen
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.DG

arXiv:2607.19305v3 Announce Type: replace-cross Abstract: Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operat...

📖 Read original article


737. ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems ​

Author: Elias Hossain, Md Mehedi Hasan Nipu, Fatema Tuj Johora Faria, Tasfia Nuzhat Ornee, Maleeha Sheikh
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.MA

arXiv:2607.19430v2 Announce Type: replace-cross Abstract: Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel through which an adversary can smuggle instructions. Existing defenses guard only the input bou...

📖 Read original article


738. Multilevel Graph Wavelet Compressed Sensing with Scale-Aware Neural Recovery ​

Author: Amirhossein Nouranizadeh, Sarang Rajendra Patil, Alan John Varghese, Varsha Narayanan, Amit Chakraborty, Mengjia Xu
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.20857v2 Announce Type: replace-cross Abstract: Scientific machine learning methods such as neural operators and physics-informed neural networks have advanced engineering applications and inverse problems, but their training typically requires large volumes of simulated data. This makes d...

📖 Read original article


739. Unified Static-Dynamic Pruning for Efficient LLM Inference ​

Author: Jinhyeok Kim, Yejoon Lee, Jaeyoung Do
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.AR, cs.LG

arXiv:2607.21985v2 Announce Type: replace-cross Abstract: The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dominate inference cost. Weight pruning offer...

📖 Read original article


740. How Context Attribution Handles What the Model Already Knows ​

Author: Quoc-Huy Trinh, Lin Zhu, Sebastian Szyller
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.23804v2 Announce Type: replace-cross Abstract: Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works show the initial success in attributing the con- tributive score of the contexts. However, we observe th...

📖 Read original article


741. Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text ​

Author: Aman Kumar, Lasitha Vidyaratne, Dipanjan D Ghosh, Arnab Chakrabarti, Ahmed K Farahat
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CL, cs.AI

arXiv:2607.26368v2 Announce Type: replace-cross Abstract: Financial disclosures may contain numerical, temporal, referential, factual, and policy inconsistencies that require different evidence and reasoning to diagnose. We study \emph{fine-grained inconsistency classification}: given a passage know...

📖 Read original article


742. Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks ​

Author: Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Yi Feng, Xiao Luo, Zijian Xiao, Haowen Xu, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2607.26574v2 Announce Type: replace-cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or...

📖 Read original article


743. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System ​

Author: Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Junyan Ye, Jinjing Zhao, Zehui Chen, Lin Chen, Renye Yan, Feng Zhao, Pheng-Ann Heng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2607.27380v2 Announce Type: replace-cross Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Exist...

📖 Read original article


744. Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration ​

Author: Rana Abou Khamis, Hala Assal, Ashraf Matrawy
Published: 8/11/2026, 4:00:00 AM
Categories: cs.HC, cs.AI

arXiv:2607.28650v2 Announce Type: replace-cross Abstract: While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration of GenAI into system administration may involve deeper shifts in professional practice that are not ye...

📖 Read original article


745. Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity ​

Author: Ramya Varunsegar
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.00143v2 Announce Type: replace-cross Abstract: Automated attack chain generation is critical for modern cybersecurity, yet manual construction fails to scale as adversary behaviors expand. While classical AI planning using the Planning Domain Definition Language (PDDL) offers a formal met...

📖 Read original article


746. Beyond Static Anchors: Bounded Prototype Conditioning for Language-Free Medical Anomaly Detection ​

Author: Yibo Wan, Jinyu Cai, See-kiong Ng
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.00442v2 Announce Type: replace-cross Abstract: Medical anomaly detection identifies abnormal images and localizes lesions under scarce supervision while generalizing across organs and modalities. Existing CLIP-based methods reduce annotation requirements through vision--language alignment...

📖 Read original article


747. Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks ​

Author: Haoyu Zhang, Xiangchen Guan, Shibo Zheng, Mohammad Zandsalimy, Shanu Sushmita
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CR, cs.AI

arXiv:2608.01043v2 Announce Type: replace-cross Abstract: We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR). T...

📖 Read original article


748. TransNRank: Towards Accurate Neoantigen Ranking with Transformer ​

Author: Zhiyin An, Yuenan Hou, Shumeng Duan, Yiming Zhou, Yuanting Zheng, Leming Shi
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CE, cs.AI

arXiv:2608.01924v2 Announce Type: replace-cross Abstract: Personalized neoantigen prediction is challenging due to the scarcity of positive samples, the noise of the experimental data, the severe class imbalance trait and the complex of immunogenicity features. Prior arts, such as linear regression ...

📖 Read original article


749. FAST-GS: Frequency Aware Space-time Gaussian Splatting for Photorealistic Dynamic Novel View Synthesis ​

Author: Zhengyang Zhang, Ziyu Lu, PengCheng Li, Hongbo Duan, Yi Liu, Pengting Luo, Peiyu Zhuang, Xinghui Li, Shaohua Ma
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.01958v2 Announce Type: replace-cross Abstract: 4D Gaussian Splatting (4DGS) excels in dynamic 3D reconstruction and real-time novel view synthesis via efficient 4D Gaussian representations and parallelizable rendering. However, existing 4DGS approaches rely on a single polynomial to model...

📖 Read original article


750. A Trust-region Framework for Moment Estimation ​

Author: Oluwasegun A. Somefun
Published: 8/11/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SY, eess.SP, eess.SY

arXiv:2608.04026v2 Announce Type: replace-cross Abstract: In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment estimation mechanisms, such as \textsc{Adam}, in stochastic gradient optimization. Specifically, the magnitude of the update step associated ...

📖 Read original article


751. Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics ​

Author: Jessica Y. Bo, Paula Akemi Aoyagui, Shalaleh Rismani, Dipto Das, Syed Ishtiaque Ahmed, Ashton Anderson
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC

arXiv:2608.05656v2 Announce Type: replace-cross Abstract: Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks favor technical methods, such as model benchmarks and LLM simulations, often sidelining empiri...

📖 Read original article


752. Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model ​

Author: Saad Ahmed, Md Khalid Syfullah
Published: 8/11/2026, 4:00:00 AM
Categories: cs.CV, cs.AI

arXiv:2608.06252v2 Announce Type: replace-cross Abstract: Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on personal devices could widen access to education and services. Existing systems use controlled-setting dataset...

📖 Read original article


753. SCALE: Scientific Concept Aggregation via LLMs and Embeddings for Fine-Grained Taxonomy Extension ​

Author: Daniele Raimondi, Feichi Lu, Oliver Grun, Mariia Eremina, Andrea Perlato
Published: 8/11/2026, 4:00:00 AM
Categories: cs.DL, cs.AI

arXiv:2608.07254v2 Announce Type: replace-cross Abstract: The increasing specialization of scientific research challenges existing classification systems, which provide effective representations of broad disciplines and research topics but often fail to capture the fine-grained conceptual structure ...

📖 Read original article