arXiv cs.AI - 2026-08-14 ​
306 items collected.
1. Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes ​
Author: Alexander Liss, Nicholas Desmond, Santiago Gil Gallego
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11207v1 Announce Type: new Abstract: When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversati...
2. Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration ​
Author: Patrik P. S"uli, Gy"orgy Eigner, Roland Holl'os
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, cs.SE
arXiv:2608.11210v1 Announce Type: new Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from...
3. A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph ​
Author: Aalok Thakkar
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.SC, math.CO
arXiv:2608.11211v1 Announce Type: new Abstract: Conway's 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists. We report a systematic, fully reproducible attack by an autonomous AI research agent, scored under the track's partial-credit metric. Our ...
4. Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts ​
Author: Parvel Gu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.11212v1 Announce Type: new Abstract: Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quantization read by a protected BF16 gate -- pushes tokens across decision boundaries and flips which experts fire. Th...
5. Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop ​
Author: Igor Itkin
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.stat-mech, cs.CL, cs.LG, cs.MA, physics.soc-ph
arXiv:2608.11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any singl...
6. AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research ​
Author: Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11216v1 Announce Type: new Abstract: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous r...
7. MaSRead: Content-Addressed Reading of Replicated Latent Stores ​
Author: Carlos Baquero, Lu'is Brito, Jo~ao Resende
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2608.11218v1 Announce Type: new Abstract: Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text. Merged by a conflict-free replicated data type, these fragments form a store that converges under any delivery order or duplication. ...
8. From Monolithic to Modular: Segment-level Automatic Prompt Optimization ​
Author: Nikita Kulin, Viktor Zhuravlev, Artur Khairullin, Sergey Muravyov, Ilya Makarov, Daniil Sukhorukov, Ekaterina Averkova
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.11219v1 Announce Type: new Abstract: Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then app...
9. LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs ​
Author: Timur Zakarin, Sergei Voitov, Sergei Shumilin, Evgeny Burnaev
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.11220v1 Announce Type: new Abstract: Nowadays, the creation of a process flow diagram (PFD) and its subsequent transformation into a piping and instrumentation diagram (P&ID) is predominantly performed manually. Applying artificial intelligence in the task could potentially lead not only ...
10. A Conceptual Framework for Refining Influence Knowledge from Simulation Evidence in Cyber-Physical Systems ​
Author: Barbara da Silva Oliveira (UniCA, Laboratoire I3S - COMRED, KAIROS), Julien Deantoni (UniCA, Laboratoire I3S - COMRED, KAIROS), Nicolas Ferry (Laboratoire I3S - COMRED, KAIROS, UniCA)
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11221v1 Announce Type: new Abstract: Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artefacts tailored to their specific domains of expertise. The behaviour of these systems emerges from the interaction between those artefacts and their operation...
11. Harnessing agent memory to build lifelong AI partners for materials scientists ​
Author: Siyu Liu, Bo Hu, Beilin Ye, He Cao, David J. Srolovitz, Tongqi Wen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci, cs.CE, cs.CL, cs.MA
arXiv:2608.11224v1 Announce Type: new Abstract: Materials research advances through accumulated experience - scripts that work, protocols that are trusted, warnings attached to failed calculations or experiments, and judgement that links a new question to an old result. This experience is essential ...
12. Identity from the Outside: A Conceptual Framework and Research Program for AI Personality Clones ​
Author: Luc E. Brunet
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11225v1 Announce Type: new Abstract: AI "personality clones" force a re-examination of personal identity in operational terms. Setting aside the hard problem of consciousness, we approach identity through the indiscernibility of manifestations, as assessed by an observer over a duration. ...
13. Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet ​
Author: Eliseo Curcio
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.SY, eess.SY
arXiv:2608.11226v1 Announce Type: new Abstract: Reinforcement-learning post-training dominates modern language-model development, yet its power behavior on GPU hardware has not been characterized, and datacenters manage GPU power with workload-blind mechanisms, static caps and reactive throttling, t...
14. Forecasting Side Effects of Activation Steering ​
Author: Chong Yong Ong, Alson Wei Jie Sim, Peixin Zhang, Jun Sun
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.11227v1 Announce Type: new Abstract: Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, steering often produces unintended side effects on other behaviors, making ...
15. Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version) ​
Author: Jack Mirenzi, Henny Admoni
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.RO
arXiv:2608.11229v1 Announce Type: new Abstract: Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly. Preference-based reward learning typically casts ...
16. The Edge-based Contiguous p-median Problem with Connections to Logistics Districting ​
Author: Zeyad Kassem, Adolfo R. Escobedo
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.DM
arXiv:2608.11230v1 Announce Type: new Abstract: This paper introduces the edge-based contiguous p-median (ECpM) problem to partition the roads in a network into a given number of compact and contiguous territories. Two binary programming models are introduced, both of which incorporate a network dis...
17. LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs ​
Author: Yirui Liu, Ruoling Qi, Longwen Wang, Xuaner Wu, Jian Chen, Yuxin Jin, Jiawei Shao, Xuelong Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11231v1 Announce Type: new Abstract: LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV cache underlies its core operations: matching reusable token chunks, concatenati...
18. InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk ​
Author: Yuan Gao (Wanxiang), Zeren Yang (Wanxiang), Junnan Li (Wanxiang), Shawn (Wanxiang), Zhong, Ahmed Dajani, Mai Zheng, Andrea Arpaci-Dusseau, Remzi Arpaci-Dusseau
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.OS
arXiv:2608.11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well s...
19. CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference ​
Author: Yifan Wu, Yufeng Zhang, Kenli Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11235v1 Announce Type: new Abstract: Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon. Many predictions stabilize early, but blockwise decoding continues until all positions are resolved, causing repeated dense for...
20. Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction ​
Author: Jiaquan Zhang, Shuxu Chen, Haifan Meng, Yi Lu, Zhihan Lyu, Fan Mo, Wei Dong, Yang Yang, Chaoning Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11237v1 Announce Type: new Abstract: Neural operators have shown strong potential for learning solution operators of partial differential equations (PDEs). However, long-horizon autoregressive prediction remains challenging: local errors accumulate as spectral inconsistency, phase misalig...
21. Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability ​
Author: Jeonghwan Choi, Taewon Yun, Minjeong Ban, Gyeonghun Sun, Jae-Gil Lee, Hwanjun Song
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11238v1 Announce Type: new Abstract: Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of use...
22. VQ-bench: A Composable Vector Quantization Framework ​
Author: Ashwin Padaki, Amir Ingber, Edo Liberty
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.DB
arXiv:2608.11240v1 Announce Type: new Abstract: Vector quantization is an old problem but has recently become central to AI infrastructure. It is therefore experiencing a surge of renewed engineering and research activity. This paper provides a unified framework for developing and benchmarking new q...
23. RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle ​
Author: Dongyang Ao, Kaixiang Fang, Shijie Xu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.IR
arXiv:2608.11241v1 Announce Type: new Abstract: Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency trilemma: general autonomy (interpreting operator intent, generating glue code zero-shot), industrial determinism (s...
24. The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification ​
Author: Yoshinori Watanabe
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.11243v1 Announce Type: new Abstract: We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if q is the data distribution and (p(...
25. BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model ​
Author: Jia-Rui Lin, Junxi Guo, Keyin Chen, Peng Pan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.11244v1 Announce Type: new Abstract: Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retrieval and manual cross-clause interpretation, which cannot reliably support multi-clause reasoning, m...
26. Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach ​
Author: Chaofan Zhai, Yicheng Song, Ravi Bapna, Junyao Ye
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG
arXiv:2608.11245v1 Announce Type: new Abstract: Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from low engagement and poor long term learning effectiveness. To address these challenges, we introduce AI Tutor, a r...
27. Towards the Harness of Embodied Agents ​
Author: Qi Wang, Tianyi Wang, Chengyang Li, Shikun Ban, Yurun Chen, Yizhong Ge, Jason Qin, Chengtai Li, Wentao Zhu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.RO
arXiv:2608.11246v1 Announce Type: new Abstract: The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. We ask whether the same paradigm extends to embodied agents in the physical world. We pr...
28. Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier ​
Author: Zafar Hussain, Kristoffer Nielbo
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.11247v1 Announce Type: new Abstract: Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs. Each agent sees what the others assert before ...
29. EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents ​
Author: Yuxi Qian, Yuxiang Ren
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2608.11248v1 Announce Type: new Abstract: Long-term memory is essential for language agents operating across extended interactions and evolving tasks. Existing memory-augmented agents mainly focus on storing and retrieving past experience, but the quality of stored memories may degrade over ti...
30. AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search ​
Author: Weicheng Ye, Youran Sun, Xingyu Ren, Shunyao Yu, Chugang Yi, Haizhao Yang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA, q-fin.CP, q-fin.PM
arXiv:2608.11250v1 Announce Type: new Abstract: Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, verify its own evidence, and preserve how each candidate was produced. We present AgonAlpha, an architecture that se...
31. Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning ​
Author: Suyash Mishra
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.ET, cs.GT, cs.MA
arXiv:2608.11252v1 Announce Type: new Abstract: Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local verification: checking at each step that the entity is representable in the chosen tool, that parameters are comp...
32. Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures ​
Author: Bongseok Kim, Suman Chakraborty, Gary Huang, Mehek Mathur, Guang Lin, Li Qiao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, physics.comp-ph
arXiv:2608.11255v1 Announce Type: new Abstract: Accurate prediction of vapor--liquid equilibrium (VLE) for hydrocarbon-nitrogen mixtures remains challenging for cubic equations of state, particularly across broad ranges of composition and hydrocarbon chain length. While deep learning models can prov...
33. Adaptive Hybrid Particle Swarm Optimization with Gradient Descent ​
Author: Aryan Gurudeo
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.NE
arXiv:2608.11258v1 Announce Type: new Abstract: Gradient injection helps Particle Swarm Optimization (PSO) only when the swarm has identified a basin with smooth local structure, not universally. We propose Adaptive Hybrid PSO (AHPSO), which uses a sigmoid function on swarm diversity to automaticall...
34. Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning ​
Author: Shibo Gao, Peipei Yang, Xu-Yao Zhang, Linlin Huang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2608.11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals. Existing approaches exhibit a "when-what" dissociation: traditional DNN-based methods localize when anomalies occur but lack semantic understanding, ...
35. Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations ​
Author: Vasundra Srinivasan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.11323v1 Announce Type: new Abstract: Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, $\tau^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total varian...
36. Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence ​
Author: Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, Handuo Zhang, Feng Chen, Fuchao Yang, Xiang Wang, Jiacheng Lin, Siting Li, Zixuan Liu, Chi Han, Zhenhailong Wang, Kunlun Zhu, Lawrence Zhao, Yueqi Guo, Kailong Wen, Feng Xing, Yiling Guo, Lidong Bing, David Tan, Bo An, Heng Ji, Sheng Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11341v1 Announce Type: new Abstract: Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a...
37. Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval ​
Author: Archan Dutta, Vyanktesh Kanungo
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.IR, cs.LG
arXiv:2608.11343v1 Announce Type: new Abstract: Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. The March 2026 r...
38. Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces ​
Author: Mengyu Chen, Feiyu Lu, Chun-Fu Chen, Lucas Vinh Tran, Jay Katukuri
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11354v1 Announce Type: new Abstract: Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UI...
39. From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate ​
Author: Pardis Taghavi, Santosh Bhavani
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.11381v1 Announce Type: new Abstract: We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-aligned specialists;...
40. When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs ​
Author: Utkarsh Bahuguna
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2608.11403v1 Announce Type: new Abstract: Self-consistency (SC) via majority vote is a widely used way to spend inference-time compute: sample N chains of thought, return the plurality answer. On the full GPQA Diamond benchmark (198 graduate-level science questions), majority voting reduces pe...
41. Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology ​
Author: Del Coburn, Scott Sanner, Dan Silver
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.11420v1 Announce Type: new Abstract: Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users. When OpenAI (2026) reports that more than 5% of ChatGPT messages globally are healthcare-related, the transpare...
42. Benchmarking LLM Judges for Mobile Agent Evaluation ​
Author: Ziqiang Wan, Li Gu, Zhixiang Chi, Zhi Liu, Seyed Mehdi Ayyoubzadeh, Yuanhao Yu, Yang Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV
arXiv:2608.11434v1 Announce Type: new Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for systematically evalua...
43. A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization ​
Author: Kelvin P. Idanwekhai, Enes Kelestemur, Benjamin Strickland, Matthew Hart, Steini Davidsson, Angelos Angelopoulos, Ron Alterovitz, Marcello DeLuca, Alexander Tropsha
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, q-bio.QM
arXiv:2608.11483v1 Announce Type: new Abstract: Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic constraints. We present SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration...
44. From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation ​
Author: Alireza S. Ziabari, Kat Ellis, Colleen Chan, Ding Tong
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.11493v1 Announce Type: new Abstract: Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw...
45. Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models ​
Author: Mingyu Zong, Sampad Mohanty, Bhaskar Krishnamachari
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11583v1 Announce Type: new Abstract: Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safet...
46. EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval ​
Author: Huiqi Miao, Xinbao Sun, Bo Wang, Fanyu Meng, Lijun Mei, Na Wu, Di Jin, Chao Deng, Junlan Feng
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11584v1 Announce Type: new Abstract: Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retriev...
47. CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications ​
Author: Linqiang Guo (Peter), Li Gu (Peter), Zihuan Jiang (Peter), Zhixiang Chi (Peter), Siobhan Reid (Peter), Ziqiang Wang (Peter), Yuanhao Yu (Peter), Wei Liu (Peter), Yang Wang (Peter), Tse-Hsun (Peter), Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11588v1 Announce Type: new Abstract: Mobile GUI agents remain brittle when deployed to applications absent from source training. We study novel-app generalization under a limited target interaction budget and without target demonstrations. We introduce CoAdapt-GUI, a test-time adaptation ...
48. Learning from Online User Feedback for Shopping Agents ​
Author: Haobo Zhang, Kelong Mao, Sulong Xu, Simiu Gu, Zhicheng Dou
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11604v1 Announce Type: new Abstract: Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide valuable supervision for improving these agents. However, existing approaches prim...
49. Foresight Without Seeing: Latent Futures for World Action Models ​
Author: Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction. Existing WAMs differ in how predictive dynamics are exposed to the action pathway. Ex...
50. MBA: Multimodal Benchmark and Agents for Real-World Business Ideation ​
Author: Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2608.11616v2 Announce Type: new Abstract: Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus intro...
51. Making AI-Generated Feedback Matter: From Provision to Student Enactment ​
Author: Omar Alsaiari, Nilufar Baghaei, Jason M. Lodge, Dragan Ga\v{s}evi'c, Naomi Winstone, Hassan Khosravi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11625v1 Announce Type: new Abstract: Feedback processes strongly influence student learning, yet their educational value depends on addressing two distinct challenges: providing high-quality, timely, and individualised feedback at scale, and supporting students to interpret, evaluate, and...
52. CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement ​
Author: Kuangzhao Yang, Ziliang Zhao, Zhicheng Dou
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11631v1 Announce Type: new Abstract: In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In such cases, directly producing an answer often leads to overgeneralized, erroneous, or low-informat...
53. XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication ​
Author: Wooseong Yang, Wei-Chieh Huang, Weizhi Zhang, Yu Wang, Philip S. Yu, Junhyun Lee
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11676v1 Announce Type: new Abstract: Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discardin...
54. AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection ​
Author: Touseef Hasan, Mounika Ghanta, Souvika Sarkar, Ujjwal Guin
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.IR, cs.MA
arXiv:2608.11679v1 Announce Type: new Abstract: Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity and volume of raw sen...
55. FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents ​
Author: Yuhao Zhang, O. Ozan Koyluoglu, Thejas Venkatesh, Richard Diehl Martinez, Vishank Bhatia, Arash Alidoust, Ashwin Paranjape
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.11683v1 Announce Type: new Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have large...
56. HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting ​
Author: Xikai Sun, Cangtian Zhou, Kebin Liu, Ke Ma, Xu Wang, Zaishu Chen, Haotian Wang, Li Liu, Yunhao Liu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11692v1 Announce Type: new Abstract: Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spatially disjoint camera views. We formulate this setting as Joint Multi-Scene Understanding (JMSU). With open-world...
57. Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning ​
Author: Lang Cao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11705v1 Announce Type: new Abstract: Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones. However, we show that the same request can elicit substantially differe...
58. Proportional Analogies on Probability Distributions via Bayesian Updating ​
Author: Pierre-Alexandre Murena
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11724v1 Announce Type: new Abstract: Analogies are quaternary relations of the form "A is to B as C is to D". Among the various formalizations of analogical reasoning, proportional analogies provide an important axiomatic framework by characterizing valid analogies through a set of postul...
59. Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents ​
Author: Zining Huang, Haoran Que, Hong Zeng, Ge Zhang, Zuo Wang, Jin Chen, Haodong Wang, Zhongfei Hou, Changxin Pu, Shen Yan, Wenhao Huang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11727v1 Announce Type: new Abstract: When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate rules in the user turn, while coding-agent benchmarks emphasize final task success...
60. HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry ​
Author: Haoran Pei, Zhao Su, Zetao Lin, Haoran Li, Jun Shen, Qi Zhu, Lan Guo, Qingguo Zhou, Binbin Yong
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11768v1 Announce Type: new Abstract: The adaptive neuro-fuzzy inference system (ANFIS) is an interpretable reasoning framework capable of generating explicit IF-THEN fuzzy rules, making it suitable for tasks requiring transparent reasoning. However, existing ANFIS models generally constru...
61. The Sleeping Agent: What Gist-Based Context Compression Loses and Why ​
Author: Nicholas E. Kyrkewood
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.11775v1 Announce Type: new Abstract: Gist-based context compression---summarising older conversation history into compact representations---is a common approach in long-horizon language model agents, yet its effect on different types of memory retrieval is poorly understood. We use Salien...
62. Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents ​
Author: Gen Dong, Yanjie Gao, Liqun Li, Tianyin Xu, Yu Hua, Fan Yang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11888v1 Announce Type: new Abstract: Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results of agent skills: som...
63. Policy-as-logic for robust reasoning over rules ​
Author: Rahul Nair, Bastian Lipka, Elizabeth Daly
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.SC
arXiv:2608.11905v1 Announce Type: new Abstract: In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries must respect written policies or rules. We present a hybrid symbolic approach that expresses policies in formal ...
64. OEIS Open: How many conjectures can language models turn into theorems? ​
Author: Tom Adamczewski
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11941v2 Announce Type: new Abstract: We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conjectures had previously been attempted only with a bespoke agent, our open-source evaluation code runs...
65. ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models ​
Author: Zhou Liu, Chaoyang Han, Zewei Pan, Zeli Su, Wentao Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11949v1 Announce Type: new Abstract: Roles provide an interpretable interface for organizing language-model agents, yet most multi-agent systems treat them as hand-written prompt labels disconnected from learned behavior and parameter updates. We argue that a useful role should instead be...
66. Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection ​
Author: Chaoran Chen, Vy Nguyen, Ziji Zhang, Abhinav Gullapalli, Ziyi Wang, Yuxuan Lu, Dakuo Wang, Jing Huang, Zhou Yu, Jin Lai
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.11977v1 Announce Type: new Abstract: Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail transiently, persistently, or silently. Robust recovery therefore requires more than repeated retries: an agent may ...
67. Claim-Level Reliability Assessment for Efficient Test-Time Reasoning ​
Author: Sen Xu, Wei Wang, Shixi Liu, Jixin Min, Yingwei Dai, Zhibin Yin, Yirong Chen, Junlin Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.11994v1 Announce Type: new Abstract: We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallocates test-time compute from additional solution sampling to targeted ve...
68. CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations ​
Author: Xingyu Yan, Tingting Dai, Antonio De Domenico, Mohamed Sana, Nicola Piovesan, Changchang Li, Bowen Liu, Kun Jiang, Mengjie Zhang, Dingcheng Shan, Jing-Cheng Pang, Chenwei Wu, Sijie Wu, Lianying Chao, Haoran Cai, Jiantao Ye, Xubin Li, Simon Mark Lucas, Xin Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12002v1 Announce Type: new Abstract: Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints. However...
69. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence ​
Author: Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.HC, cs.LG, cs.MA
arXiv:2608.12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic explorati...
70. Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges ​
Author: Xi Chen, Jie Mu, Mo Xuan, Qun Shao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12097v1 Announce Type: new Abstract: Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to judge but leave criterion composition implicit, even when natural-language rules state it. We introduce Graph-Structured Rubrics (GSR), which compil...
71. GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings ​
Author: Shivali Dalmia, Sumukha Thoppanahalli, Mohammadreza Sediqin, Abhishek Mukherji
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12133v1 Announce Type: new Abstract: Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables, and embedded images. Existing LLM and VLM systems face hallucinated content, table structure degradation, and lack governed workflows extending b...
72. Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation ​
Author: Rodrigo Guedes de Souza, Alison R. Panisson
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2608.12150v1 Announce Type: new Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels (64--4,096), ...
73. How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models ​
Author: Aleksandra Kalisz, Jack Simons, Krisztina Sinkovics, Noam Ghenassia, Shikha Surana, Henry Moss, Paul Duckworth
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2608.12192v1 Announce Type: new Abstract: Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint. Existing guidance methods, ...
74. An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS ​
Author: Yuzhong Shen, Masha Sosonkina, Peng Xu, Mark S. Gordon
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12249v1 Announce Type: new Abstract: Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and across much of computational science the work simply goes undone. We propose an agentic workflow that takes this wor...
75. VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies ​
Author: Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12282v1 Announce Type: new Abstract: Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}...
76. Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models ​
Author: Saman Marandi, Yu-Shu Hu, Mohammad Modarres
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.12304v1 Announce Type: new Abstract: Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional objectives to underlying structural elements. However, DML construction typically relies on expert interpretation of technical documenta...
77. Evaluating LLM Generated Detection Rules in Cybersecurity ​
Author: Anna Bertiger, Bobby Filar, Aryan Luthra, Stefano Meschiari, Aiden Mitchell, Sam Scholten, Vivek Sharath
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2509.16749v1 Announce Type: cross Abstract: LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners. Here, we present an open-source evaluation framework and benchmark metrics for eva...
78. Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs ​
Author: Ruoxi Zhao, Maziar Raissi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11232v1 Announce Type: cross Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a framework with two compl...
79. Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets ​
Author: Mark Shapiro
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.11233v1 Announce Type: cross Abstract: A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing. Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent Block, and a Coda, ...
80. TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation ​
Author: Jiahui Zhang, Ziwei Zhang, Yipeng Wang, Yibo Liu, Haozhou Pang, Yikai Hu, Hongyan Ren, Lan Zhou, Qi Gan, Kai Sheng
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11236v1 Announce Type: cross Abstract: Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist evaluation f...
81. Reinforcement Learning based DBMS Buffer Pool Auto-Tuning for Optimal Memory Utilization ​
Author: Yifan Wang, Patrick Royer, Rapha"el F'eraud, David Delande
Published: 8/14/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2608.11239v1 Announce Type: cross Abstract: Administering Database Management Systems (DBMS) instances requires Database Administrators (DBA) to balance performance in terms of Service Level Agreement (SLA) against resource usage, often prompting RAM over-allocation that wastes memory. We intr...
82. Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction ​
Author: Zhiqi Wang, Yichi Zhang, Dongwon Lee, Yuchen Yang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11242v1 Announce Type: cross Abstract: When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks. We identify a class of user-issued instructions, Session Constraints (SCs), such as "do not delete any emails until I confirm," that are meant to ...
83. Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression ​
Author: Angelo Nardone, Paolo Ferragina
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IT, cs.LG, math.IT
arXiv:2608.11249v1 Announce Type: cross Abstract: We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language...
84. Variable Selection in the Context of AI Fairness ​
Author: Ivan Luciano Danesi, Chiara Frigerio, Fabio Maccaferri, Giorgio Alessandro Motta, Pietro Zecca
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG
arXiv:2608.11251v1 Announce Type: cross Abstract: Fairness in AI systems has become more important with recent regulatory demands, such as the EU AI Act. Traditional approaches often do not take into account philosophical ethics and social awareness. Variable selection processes, in particular, can ...
85. Methodologies for Improving the Quality of AI Tutoring in K-12 Education ​
Author: Tushar Udeshi, Anna Khazenzon, Kabir Khan, Nick Breen, RJ Corwin, Chris DiGiano, Kodi Weatherholtz, Marek Zaluski
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.11259v1 Announce Type: cross Abstract: Many AI tutors leverage large language models (LLMs) today. Given that LLMs are opaque black boxes, robust evaluation and live experimentation to measure the impact of every change are essential. We pioneered AI-powered tutoring for K-12 with the lau...
86. Agent Safety Should Be a Runtime Contract ​
Author: Albus W. Ng, Yi Han, Jusheng Zhang, Wenhao Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.11274v1 Announce Type: cross Abstract: The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally insufficient for autonomous agents that execute code, mutate files, send messages, and modif...
87. Every pooling rule has its world: matching probability combination rules to situations and stakes ​
Author: Tanel Tammet, Priit J"arv, Dirk Draheim
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ME, cs.AI
arXiv:2608.11275v1 Announce Type: cross Abstract: Systems often need to combine two numerical assessments of the same yes/no question. The appropriate formula depends on what the numbers represent and on how the sources are related. Averaging is correct when one of several alternative interpretation...
88. Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification ​
Author: Rofiqul Islam, Lilatul Ferdouse
Published: 8/14/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, cs.LG
arXiv:2608.11280v1 Announce Type: cross Abstract: Skin cancer diagnosis from dermoscopic images remains challenging due to high intra-class variability, inter-class similarity, class imbalance, and the limited interpretability of deep learning models. This paper proposes an uncertainty-aware and exp...
89. Federated Learning for Distributed CNC Tool Wear Prediction ​
Author: Afsana Khan, Morris Stallmann, Marcin Pietrasik, Charis Kouzinopoulos, Anna Wilbik
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11281v1 Announce Type: cross Abstract: Tool wear prediction is an important task in CNC machining, where accurate monitoring of tool condition supports product quality and process reliability. Machine learning methods have shown potential for this task, but their use in industrial environ...
90. Physics-Informed Implicit Neural Representations for Improved Myocardial Perfusion MRI Quantification ​
Author: Christos Tsepas, Chang Yan, Maximilian Fuetterer, Sebastian Kozerke, Cian M Scannell
Published: 8/14/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.LG
arXiv:2608.11282v1 Announce Type: cross Abstract: Quantifying myocardial perfusion from cardiac magnetic resonance (CMR) can be achieved by fitting tracer-kinetic models to the dynamic contrast-enhanced MR data. However, fitting the observed data with multi-compartment exchange models, which describ...
91. Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models ​
Author: Guobin Zhao, Xiao-Yan Li
Published: 8/14/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI
arXiv:2608.11283v1 Announce Type: cross Abstract: Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity. Existing validation approaches...
92. SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation ​
Author: Dongsu Song, DaeYun GO, Boseung Seo, Jay Hoon Jung
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CR, cs.LG
arXiv:2608.11285v1 Announce Type: cross Abstract: Despite the practical relevance of sparse decision-based black-box threats, they have received limited attention in semantic segmentation. To bridge this gap, we adapt the most representative decision-based black-box sparse attacks from the classific...
93. CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification ​
Author: Gawon Lim
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.11287v1 Announce Type: cross Abstract: Long-tailed classification poses a reliability challenge because models trained on imbalanced data are unevenly reliable across frequent and underrepresented classes. While existing methods address imbalance through re-balancing, adjustment, represen...
94. Backdoor Decontamination Dynamics in LLM Agents ​
Author: Gabriel Huang, Abhay Puri, L'eo Boisvert, Alexandre Drouin, Perouz Taslakian, Spandana Gella, Christopher Pal
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.11295v1 Announce Type: cross Abstract: Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing. Assuming defenders do not know the existing trigger, they cannot unlearn it directly. O...
95. Clinical Feasibility of Low-Magnification Fluorescence Imaging for Breast Cancer Margin Detection Using Texture Analysis and Deep Learning ​
Author: Pouya Afshin, Tianling Niu, Tongtong Lu, David Helminiak, Julie Jorns, Mollie Patton, Tina Yen, Donghye Ye, Bing Yu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.11317v1 Announce Type: cross Abstract: High-resolution images of unprocessed surgical breast tissue can be obtained using microscopy with ultraviolet surface excitation (MUSE). This technique is considered a promising method for checking surgical margins during breast cancer surgery. In t...
96. Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction ​
Author: Yi Liu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11318v1 Announce Type: cross Abstract: Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent. We develop a decision-resource view of terminal symmetry: process evidence supplies directionality, terminal corresp...
97. Socioduality: A Relational Process Framework for Human-AI Interaction ​
Author: Mehmed Zahid \c{C}"ogenli
Published: 8/14/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.11322v1 Announce Type: cross Abstract: Human-AI research often evaluates individual capabilities, combined performance, or final outputs, but these approaches do not preserve how one party's response becomes part of the conditions under which the other party's next contribution is formed....
98. Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings ​
Author: Tran Le Vu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE, math.OC
arXiv:2608.11324v1 Announce Type: cross Abstract: This paper proposes a contextual quality-diversity evolutionary reinforcement-learning controller, CQD-ERL, for the supervisory control of a tropical, water-cooled chiller plant and its associated air side. Rather than converging to a single scalaris...
99. Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation ​
Author: Md Maklachur Rahman, Tracy Hammond
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, eess.IV
arXiv:2608.11335v1 Announce Type: cross Abstract: Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries. We propose Dual-Domain Cross-Modal Decoding (DD-CMD) for clinical text-...
100. Self-evolving network verifiers ​
Author: Ioannis Protogeros, Tibor Schneider, Laurent Vanbever
Published: 8/14/2026, 4:00:00 AM
Categories: cs.NI, cs.AI
arXiv:2608.11340v1 Announce Type: cross Abstract: Symbolic network verifiers can reason about correctness across vast spaces of routing inputs and failures, but only for the protocols and features an expert has encoded by hand. Creating and maintaining a faithful model of the control plane is both d...
101. Governing Agentic AI in FinTech ​
Author: Henry Han
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, q-fin.RM
arXiv:2608.11344v2 Announce Type: cross Abstract: Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding gov...
102. Dynamics Models for Offline Hyperparameter Selection in Real-World RL ​
Author: Jordan Coblin, Han Wang, Martha White, Adam White
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11349v1 Announce Type: cross Abstract: A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unavailable and online experimentation is costly. Prior work has proposed calibration models trained on offline dat...
103. Gaze Target Estimation Anywhere with Concepts ​
Author: Xu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou, Tianyu Xu, Adarsh Kowdle, Inki Kim, James M. Rehg
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.11367v1 Announce Type: cross Abstract: Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify...
104. AI Guardrail Survival under Single-Cycle Agentic Self-Summarization ​
Author: Ted Kwartler, Alan Aqrawi, Arian Abbasi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.11392v2 Announce Type: cross Abstract: Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dropping a standing safety constraint during compaction drives behavioral violations across many models (Governanc...
105. TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs ​
Author: Valentin Rodionov, Shamil Assylbekov
Published: 8/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.11415v1 Announce Type: cross Abstract: Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists. Such deployment assumes the model can distinguish reliable scientific literature from unreliable literature, a capability that...
106. Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards ​
Author: Sim'on Pati~no Idarraga, Erick Silva, Rehana Yasmin, Ali Shoker
Published: 8/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.SY, eess.SY
arXiv:2608.11451v1 Announce Type: cross Abstract: Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never miss. The reason is structural: they learn statistical patterns rather than the physical conditions that guara...
107. TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation ​
Author: Haoqi Hu, Tongji Luo, Li Zhang, Boning Zhou
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.11452v1 Announce Type: cross Abstract: Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image renders the meaning of a poem. The task is many-sided: a good illustration must be visually sound, faithful to the ...
108. PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition ​
Author: Vasant G. Honavar, Satish Kumar Keshri, Neil Ashtekar, Zehao Liu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11465v1 Announce Type: cross Abstract: PAC-Bayes theory provides generalization guarantees by controlling the Kullback--Leibler (KL) divergence between posterior and prior distributions over a chosen hypothesis representation. However, predictive risk depends only on the predictive behavi...
109. The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark ​
Author: Jeremy Spence, Nicholas Assaderaghi, Jinhao Zhu, Nikil Ravi, Raluca Ada Popa, Guannan Wei, Yangruibo Ding, Zhuo Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.SE
arXiv:2608.11469v1 Announce Type: cross Abstract: AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequential to cybersecurity, including malware, firmware, and proprietary applications, is available only a...
110. HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging ​
Author: Hyo Seo Kim, Ren Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11499v1 Announce Type: cross Abstract: Task vectors enable model merging without joint retraining. In practice, the subset of task vectors to be merged may vary, but many existing methods use scalar tuning for a particular subset, requiring repeated tuning across subsets and restricting t...
111. Strengthening Full Justified Representation: Efficient Verification and Computation ​
Author: Nicholas Teh
Published: 8/14/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, econ.TH
arXiv:2608.11500v1 Announce Type: cross Abstract: Full justified representation (FJR) is among the strongest known satisfiable proportionality axioms for approval-based committee elections. Recent work has shown that an FJR committee can be found in polynomial time, but verifying whether a given com...
112. Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task ​
Author: Xiaoyang Hu, Mike Angstadt, Shane Storks, Zan Huang, Aman Taxali, Alex Weigard, Richard L. Lewis, Chandra Sripada
Published: 8/14/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI
arXiv:2608.11510v1 Announce Type: cross Abstract: Congruency effects, observed in conflict tasks such as Stroop and flanker tasks, have been investigated for nearly a century in psychology and neuroscience, but their mechanistic basis is not fully understood. We introduce a verbal-only LLM conflict ...
113. Let it Cook: Learning to Wait in Sequential Decision Making ​
Author: Christopher Watson, Arjun Krishna, Dinesh Jayaraman, Rajeev Alur
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11511v1 Announce Type: cross Abstract: In sequential decision making, an agent typically observes its environment and acts at every timestep. However, such active participation may not always be necessary; tasks such as brewing coffee include periods that are served equally well by lettin...
114. Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation ​
Author: Alex Deaconu, Anubhav Gupta, Manaal Basha, Nicholas Haydu, Gema Rodr'iguez-P'erez
Published: 8/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL
arXiv:2608.11513v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly integrated into software engineering workflows, helping developers write, debug, test, and maintain code. While prompt wording and structure are known to influence model performance, the impact of psychol...
115. Keep the Future, Drop the Rollout: RIFT for World Action Models ​
Author: Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.11521v2 Announce Type: cross Abstract: World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across fou...
116. Hierarchical Federated Transfer Learning in Digital Twin-Based Vehicular Networks ​
Author: Qasim Zia, Saide Zhu, Haoxin Wang, Zafar Iqbal, Yingshu Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11532v1 Announce Type: cross Abstract: In recent research on the Digital Twin-based Vehicular Ad hoc Network(DT-VANET), Federated Learning (FL) has shown its ability to provide data privacy. However, Federated learning struggles to adequately train a global model when confronted with data...
117. Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment ​
Author: Weize Cai, Yongqi Dong, Zhida Shao, Zixin Fu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV
arXiv:2608.11537v1 Announce Type: cross Abstract: Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rende...
118. A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era ​
Author: Dalton Ross Smith, Wilburn Whittington, Alejandro Martinez, Aidan Duncan, Gang Li
Published: 8/14/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.CY, cs.SY
arXiv:2608.11540v1 Announce Type: cross Abstract: The convergence of artificial intelligence (AI), Industrial Internet of Things, cyber-physical systems, and advanced robotics is reshaping manufacturing faster than engineering curricula can adapt, widening the gap between the competencies required o...
119. Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents ​
Author: Dylan Bouchard, Mohit Singh Chauhan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.11552v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where th...
120. From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection ​
Author: Zepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen, Fei Wang, Daiguo Zhou
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, eess.IV
arXiv:2608.11562v1 Announce Type: cross Abstract: Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplo...
121. Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs ​
Author: Vu Duc Anh, Nhat M. Hoang, Do Xuan Long, Cong-Duy Nguyen, Ponhvoan Srey, Luu Anh Tuan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11573v1 Announce Type: cross Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learning based, two-stage...
122. RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation ​
Author: Yueyuan Li, Zexi Chen, Weijie Xi, Mingyang Jiang, Songan Zhang, Hanyang Zhuang, Ming Yang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.11580v1 Announce Type: cross Abstract: Autonomous driving simulation requires diverse and scalable lane-level HD maps to support long-horizon evaluation across complex road networks. Existing approaches either rely on handcrafted or reconstructed real-world maps, which limits scalability,...
123. A Hybrid Framework of Vision Transformer and Gated Recurrent Unit for Detection of Mosquito Diseases ​
Author: Danial Sharifrazi, Saadat Behzadi, Nouman Javed, Roohallah Alizadehsani, Prasad N. Paradkar, Asim Bhatti
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.11582v1 Announce Type: cross Abstract: Identifying dengue virus-infected mosquitoes from control mosquitoes is a major challenge in analyzing mosquito locomotion behavior due to the small size and complexity of the video background. Conventional AI methods are often unable to extract accu...
124. Dion3: Full-Stack Orthogonal Updates ​
Author: Noah Amsel, Jack Zhang, Kwangjun Ahn, Ali Naeimi, Austin Feng, Berlin Chen, Tri Dao, John Langford
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11612v1 Announce Type: cross Abstract: The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step. When weights are sharded, communication overhead compounds this computational cost, eroding the benefits of Muon in many settings. We pr...
125. FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting ​
Author: Rentao Gu, Yihang Ding, Junjie Li, Yi Ding, Weijing Sang, Xiaoli Huo, Xin Qin, Yuefeng Ji
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NI, eess.SP
arXiv:2608.11623v1 Announce Type: cross Abstract: Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods rely heavily on textual prompts for modality alignment-introducing nontrivial computational overhead and failing...
126. Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs ​
Author: Nimet Beyza Bozdag, Emre Can Acikgoz, Gokhan Tur, Dilek Hakkani-T"ur
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11624v1 Announce Type: cross Abstract: Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each o...
127. Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones ​
Author: Oshan A. B. Yalegama, Wageesha N. Manamperi
Published: 8/14/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, eess.SP
arXiv:2608.11627v1 Announce Type: cross Abstract: The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied to speech enhancement in noisy environments. Estimating the Re...
128. Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents ​
Author: Jun He, Deying Yu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.MA, cs.AI
arXiv:2608.11632v1 Announce Type: cross Abstract: Persistent AI agents accumulate versioned state across long horizons, but storage retention alone does not identify authoritative state. Without an explicit control plane, unmediated updates by models, tools, and background workers risk stale overwri...
129. Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting ​
Author: Xikai Sun, Kebin Liu, Haotian Wang, Li Liu, Xu Wang, Yunhao Liu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.11655v1 Announce Type: cross Abstract: Motion-centric video reasoning is fundamental to interactive applications such as robotic manipulation and autonomous navigation. However, multimodal large language models (MLLMs) typically process videos through sparse uniform sampling to control vi...
130. Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models ​
Author: Yoshihiko Kayama
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, nlin.CG
arXiv:2608.11657v1 Announce Type: cross Abstract: We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization problem into a continuous dynamical system within the macroscopic logit space. By establishing a non-linear home...
131. Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning ​
Author: Zijian Zhao, Sen Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA
arXiv:2608.11658v1 Announce Type: cross Abstract: Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically after deployment, and retraining a policy for each new objective is prohibitively expensive. For a single agent, th...
132. Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing ​
Author: Tianci Liu, Zihan Dong, Tianchun Li, Yi-Chung Chen, Qiming Cao, Xingchen Wang, Shiyang Wang, Zichen Miao, Linjun Zhang, Haoyu Wang, Jing Gao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.11660v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which upda...
133. Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads ​
Author: Zijian Zhao, Sen Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11661v1 Announce Type: cross Abstract: A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings. This architecture has been developed independently in operator learning, bipartite matching, contrastive vision...
134. Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL ​
Author: Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.11669v1 Announce Type: cross Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer. The rubric, however, is a fixed proxy for quality, never a complete descr...
135. GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs ​
Author: Kai Yang, Jingwei Xu, Wanyu Wang, Kai-Yuan Guo, Zhenbo Yu, Yi Wang, Yu Qiao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11674v1 Announce Type: cross Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradation, and response-length inflation. Although prior work has characteri...
136. Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation ​
Author: Duy Tran Thanh, Yeejin Lee, Byeongkeun Kang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM
arXiv:2608.11681v1 Announce Type: cross Abstract: This work addresses the challenge of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to recognize both predefined and unseen object categories without exhaustive human annotations. Existing methods of...
137. APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference ​
Author: Alish Kanani, Layan Badawi, Umit Y. Ogras
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.LG
arXiv:2608.11688v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency. However, MoE inference at the edge is fundamentall...
138. The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance ​
Author: Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11694v1 Announce Type: cross Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while keeping its meaning a...
139. REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation ​
Author: Yang Sun, Lichao Ma, Houyuan Qin, Yuxin Liu, Hanyang Lu, Yao Zhu, Pinlong Cai, Guohang Yan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11698v2 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond direct imitation, b...
140. Consolidator: Learning Persistent Routed Memory Across Context Boundaries ​
Author: Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11701v1 Announce Type: cross Abstract: Copying short-term memory (STM) into a slower store can preserve state across a context boundary, but persistence alone does not ensure that the retained state influences subsequent memory access. We test this distinction in a Phasor Memory Network (...
141. Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing ​
Author: Ziqiang Li, Yun Liu, Gouhei Tanaka
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11704v1 Announce Type: cross Abstract: Dynamic Time Warping (DTW)-based Nearest-Neighbor (NN) classifiers are effective for time-series classification but are vulnerable to mislabeled training samples and require numerous DTW computations during inference. We propose DTW-based Granular Ba...
142. High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions ​
Author: Hongyan Wang, Jiayu Huang, Haotian Zheng, Xin Gao, Chi Ding, Ying Liu, Xia Wang, Qing Xu, Keqiang Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11713v1 Announce Type: cross Abstract: Multi-objective Bayesian optimization (MOBO) is effective in identifying the Pareto fronts for expensive black-box problems. However, most current MOBO approaches are limited to low-dimensional decision space due to its exponential sampling complexit...
143. When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use ​
Author: Siddharth Chauhan, Thomas Butler, Abhishek Singhania, Pankaj Porwal, Honey Gupta
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11715v1 Announce Type: cross Abstract: The reliability of Large Language Models (LLMs) for API calling degrades in multilingual settings. A common failure occurs when a model selects the correct tool but generates argument values in an inconsistent language, which we term Argument Languag...
144. Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation ​
Author: Yuanmin Huang, Chen Chen, Geng Hong, Xiaoyu You, Hui Xue, Zhenxing Qian, Mi Zhang, Min Yang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.11732v1 Announce Type: cross Abstract: Proprietary text-to-image diffusion models are increasingly distributed as hosted services and downloadable checkpoints, making their intellectual property (IP) protection an increasingly critical concern when model leakage, copying, or unauthorized ...
145. A 12-CNOT Double Qubit Excitation Gate ​
Author: Irfansha Shaik
Published: 8/14/2026, 4:00:00 AM
Categories: quant-ph, cs.AI
arXiv:2608.11733v1 Announce Type: cross Abstract: Effective implementation of high-level quantum gates is essential for practical quantum computing. To the best of our knowledge, we present the first reported 12-CNOT decomposition of the double qubit excitation operator, improving upon state-of-the-...
146. Locating and Controlling Implicit Personalization in Large Language Models ​
Author: Yueru Yan, Siqi Wu, Thai Le
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.11735v1 Announce Type: cross Abstract: Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and th...
147. Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System ​
Author: Haoyu Zhang, Shuoxun Zhang, Peng Ye, Lin Zhang, Jiakang Yuan, Shenghong Yi, Yuening Wang, Tao Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.11738v1 Announce Type: cross Abstract: Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses distinct challenges arising from extreme scale variation, arbitrary camera orientations, and high object density....
148. G0.5: One Autoregressive Stream for Robot Reasoning and Action ​
Author: Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao
Published: 8/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.11739v1 Announce Type: cross Abstract: The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregre...
149. JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis ​
Author: Ran Li, Huiguo He, Jiahuan Cao, Junle Liu, Hiuyi Cheng, Lianwen Jin
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.11741v1 Announce Type: cross Abstract: The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieva...
150. MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning ​
Author: Shiji Zhou, Kunlin Lyu, Lei Zhang, Ruodong Wang, Yifan Sun
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2608.11749v1 Announce Type: cross Abstract: Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradient manipulation. However, most existing methods flatten model parameters into vectors and perform gradient manip...
151. Instruction Alignment for Binary Code Representation Learning ​
Author: Huaijin Wang, Shuai Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CR
arXiv:2608.11766v1 Announce Type: cross Abstract: Binary code representation learning is a fundamental problem in software security and reverse engineering. Existing methods mainly learn function-level embeddings that capture coarse-grained semantic relationships between binary functions, but they l...
152. GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation ​
Author: Ofir Ben Shoham, Shrutendra Harsola, Vignesh Subrahmaniam, Shravan Mohan, Yakov Gazman, Oded Vainas
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.11787v1 Announce Type: cross Abstract: Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision is difficult: histor...
153. TELLME: Test-Enhanced Learning for Language Model Enrichment ​
Author: Minjun Kim, Inho Won, Hyeonseok Lim, MinKyu Kim, Junghun Yuk, Wooyoung Go, Jongyoul Park, Jungyeul Park, KyungTae Lim
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11788v1 Announce Type: cross Abstract: Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such as the difficulty of acquiring large-scale domain-specific datasets an...
154. Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion ​
Author: Adrian Rauchfleisch, Andreas Jungherr
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2608.11794v1 Announce Type: cross Abstract: The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of content provenance and AI involvement. But the effects of such disclosures remain uncertain. We test two disclosu...
155. Towards Model-based Run-time Cybersecurity: On Control-Flow Anomaly Detection, Attack Identification, and Hardware Monitoring ​
Author: Martin Sachenbacher, Martin Leucker, Alexander Weiss, Aliyu Tanko Ali
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.11802v1 Announce Type: cross Abstract: Methods to increase the resilience of systems to cyber-attacks become increasingly important. Control-flow monitoring provides a principled basis to ensure integrity and detect possible anomalies at run-time. Once anomalies have been detected, so-cal...
156. How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment ​
Author: Guang Yang, Fengchen Liu, Alex Wang, Homa Hosseinmardi, Amir Ghasemian
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL
arXiv:2608.11816v1 Announce Type: cross Abstract: State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced benchmark of 200 core e...
157. User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling ​
Author: Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magn'usson, Praveen Kumar Donta
Published: 8/14/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG, cs.NI, cs.PF
arXiv:2608.11840v1 Announce Type: cross Abstract: Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. We propose a collaborative distributed inference system combining dedicated infrastructure with resou...
158. LookBack: Where and How to Score LVLM Responses via Visual Reference Usage ​
Author: Beomsik Cho, Jinhyeong Kim, Dongseok Lee, Jaehyung Kim
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2608.11847v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text-level hallucinations; they also hallucinate...
159. Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical Spectra ​
Author: Waleed Waseer, Muhammad Shahid Jabbar, Muhammad Sohail Ibrahim, Shujaat Khan
Published: 8/14/2026, 4:00:00 AM
Categories: physics.optics, cs.AI, cs.CV
arXiv:2608.11860v1 Announce Type: cross Abstract: Data-driven inverse design enables efficient generation of nanophotonic structures with prescribed optical responses, but spectrum-to-geometry mapping remains challenging due to non-uniqueness and fine geometric features. This work presents a two-sta...
160. CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation ​
Author: Xue Yang, Rigui Zhou, ShiZheng Jia, Dax Enshan Koh, Siong Thye Goh, Young-Wook Cho, YaoChong Li, Xuezhi Ma, Hongyu Chen, Xin Wang
Published: 8/14/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.CV
arXiv:2608.11884v1 Announce Type: cross Abstract: Quantum generative adversarial networks (QGANs) have attracted increasing attention for image generation using parameterized quantum circuits. Existing amplitude-based approaches face two key limitations: pixel locations are typically encoded by comp...
161. DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation ​
Author: Anik Pramanik, Murat Kantarcioglu, Vincent Oria, Shantanu Sharma
Published: 8/14/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CL, cs.IR
arXiv:2608.11889v1 Announce Type: cross Abstract: Prompting-based (\textit{i}.\textit{e}., non-fine-tuning) Text-to-SQL methods, where underlying large language model parameters are not changed for the task, face three problems: (\textit{i})~relying on coarse-grained schema information that may not ...
162. Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework ​
Author: Avinash Agarwal, Vridhi Jain
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC
arXiv:2608.11891v1 Announce Type: cross Abstract: Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. Assessing the progress of such national ecosystems is complicated by inconsistent benchmark reporting, p...
163. Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models ​
Author: Hao Zhang, Jiaxin Qi, Zhijiang Tang, Jianqiang Huang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.11907v1 Announce Type: cross Abstract: As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols ...
164. Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians ​
Author: Timothy Heightman, Elena Orlova, Philip Mantrov, Aleksei Ustimenko
Published: 8/14/2026, 4:00:00 AM
Categories: quant-ph, cond-mat.dis-nn, cond-mat.str-el, cs.AI
arXiv:2608.11911v2 Announce Type: cross Abstract: A central promise of useful quantum advantage is the ability to compute ground states of Hamiltonian systems beyond the reach of classical simulation methods. Here we demonstrate that this problem can be effectively amortized across an arbitrary and ...
165. Accuracy and Order Sensitivity Diverge Under Label-Free Strategies ​
Author: Karl Hanna, Chen Feng
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.11947v1 Announce Type: cross Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test whether preventing a mod...
166. TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement ​
Author: Karim Aly, Alexei Sharpanskykh, Jacco Hoekstra
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11951v1 Announce Type: cross Abstract: Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disruptions with substantial operational, economic, and safety costs. Such events are rare in historical records, leaving insufficient trai...
167. Causal inference for group-contaminated structured outcomes: observable quotients, lossless reduction and exact randomization inference ​
Author: Usef Faghihi, Amir Saki
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ME, cs.AI
arXiv:2608.11954v1 Announce Type: cross Abstract: Structured potential outcomes such as microscopy images may be recorded after an unknown, unit-specific transformation. If that transformation can depend on treatment, covariates or the intrinsic outcome, raw-coordinate analyses may mix biological ef...
168. LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation ​
Author: Zhixin Zhang, Xinke Jiang, Zhibang Yang, Weixuan Xu, Guohong Qiu, Xu Chu, Junfeng Zhao, Yasha Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.11967v1 Announce Type: cross Abstract: Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is reflection: assessing trajectory progress, identifying missing evidence...
169. HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs ​
Author: Kangning Zhang, Haotian Fang, Xukun Luo, Hao Yin, Yang Gao, Peng Yan, Weiwen Liu, Weinan Zhang, Yong Yu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2608.11980v1 Announce Type: cross Abstract: Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence. This paradigm enables a unified generation interface for item IDs, h...
170. Remote Sensing and Machine Learning-Based Analysis of Land Use and Vegetation Change in Dhaka District, Bangladesh ​
Author: Muhammad Masud Tarek, Md. Alamgir Hossain, Md. Samiul Islam, Muntasir Hasan Kanchan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2608.12001v1 Announce Type: cross Abstract: Rapid urbanization in Dhaka District, Bangladesh has triggered substantial alterations in land use and environmental conditions, necessitating systematic monitoring for informed urban planning and ecological sustainability. This study employs remote ...
171. RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks ​
Author: Jinjun Huang, Zhongzhen Wen, Tongtong Xu, Meng Yan, Xin Xia, Zhongxin Liu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.12004v1 Announce Type: cross Abstract: In modern AI frameworks, GPU kernels are key to overall system performance. Combining usability, portability, and near-handwritten CUDA performance, Triton is widely adopted for implementing GPU kernels. Recent advances show the potential of large la...
172. Dual-Model Sentiment Analysis of Consumer Reviews in the Retail Coffee Sector Using Machine Learning and Deep Learning Approaches ​
Author: Muntasir Hasan Kanchan, Md. Alamgir Hossain, Md. Samiul Islam, Muhammad Masud Tarek
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2608.12007v1 Announce Type: cross Abstract: Consumer reviews play an important role in shaping brand perception and business strategies, particularly in service-driven industries such as retail coffee. This study presents a comparative sentiment analysis framework for Starbucks customer review...
173. From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices ​
Author: Tuhinangshu Gangopadhyay, Rasmus Adler, Peter Liggesmeyer, Jan Reich
Published: 8/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2608.12025v1 Announce Type: cross Abstract: Medical devices are becoming more software-intensive, connected, and AI-enabled. Their development requires risk-management evidence aligned with ISO 14971 and, for software, IEC 62304. This evidence must be kept consistent across requirements, desig...
174. Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision ​
Author: Shaojie Zhang, Ke Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12027v1 Announce Type: cross Abstract: Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption. Existing deep constrained c...
175. LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration ​
Author: Enhuai Liu, Yunke Wang, Yutong Wang, Changming Sun, Chang Xu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12032v1 Announce Type: cross Abstract: Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic cost that dominates as resolution and duration grow. Sparse attention reduces this cost without retraining, but ...
176. How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging ​
Author: Yiheng Xiong, Luisa Gall'ee, Daniel Santak Wolf, Heiko Hillenhagen, Michael G"otz
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12035v1 Announce Type: cross Abstract: Deploying unsupervised domain adaptation (UDA) in clinical practice requires choosing which algorithm to use and which of its trained models to ship. However, the deployment (target) domain is unlabeled, so models cannot be evaluated directly on it, ...
177. Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations ​
Author: Lior Baruch, Moshe Butman, Kfir Bar, Doron Friedman
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12062v1 Announce Type: cross Abstract: Developing dialogue systems capable of engaging in multi-turn, goal-oriented conversations remains a significant challenge, especially in specialized domains with limited data. This research proposes a novel framework called Preference Tree Optimizat...
178. Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL ​
Author: Martin Schuck, Maks Sorokin, Simone Manni, Duy Ta, Angela P. Schoellig, Marco Hutter, Simon Le Cleac'H, Jan Br"udigam
Published: 8/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI
arXiv:2608.12063v1 Announce Type: cross Abstract: Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping. To bypass this limitation, we lev...
179. Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models ​
Author: Shukrullo Nazirjonov, Sai Prasanna, Anna Manasyan, Georg Martius
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.12078v1 Announce Type: cross Abstract: Learning world models from offline trajectories enables agents to accomplish different tasks through planning. Object-centric (OC) representations, which decompose a scene into a set of slots that bind to its objects, have been proposed as an inducti...
180. Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion ​
Author: David Bechtoldt, Sidney Bender
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12083v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) achieve strong predictive performance on graph-structured data across domains such as chemistry, biology, and network analysis, yet they provide no intrinsic explanation of their predictions. This limits their adoption in...
181. Confidence Calibration of Deep Learning Systems ​
Author: Coby Penso
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2608.12100v1 Announce Type: cross Abstract: In high-stakes applications, reliable confidence estimates are as important as the predictions themselves. Confidence calibration ensures that predicted probabilities reflect the likelihood of correctness, making it essential for safe deployment of d...
182. No One to Blame: A Framework of Constitutive AI Unaccountability ​
Author: Long Hoang Nguyen, Eva Sp"athe, Sebastian Lins, Ali Sunyaev
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.12104v1 Announce Type: cross Abstract: The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominantly frames AI accountability gaps as barriers that can be overcome through better standards, transparency, and i...
183. QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving ​
Author: Yilin Liu, Rui Meng, Wangze Ni, Jianxin Yan, Heng Cao, Libin Zheng, Peng Cheng, Jinfei Liu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12121v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations. Position-Independent Caching (PIC) mitigates it by reusing precomputed Key-Value (KV) across positions, but its efficienc...
184. Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control ​
Author: Josef Liyanjun Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.OS
arXiv:2608.12123v1 Announce Type: cross Abstract: LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what cha...
185. Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation ​
Author: Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer
Published: 8/14/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.CL, cs.MA
arXiv:2608.12125v1 Announce Type: cross Abstract: As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation p...
186. Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids: From Robust Offline Optimization to Full-Bandit Learning ​
Author: Vaneet Aggarwal
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CC, math.OC
arXiv:2608.12134v1 Announce Type: cross Abstract: We study nonnegative submodular maximization subject to a general matroid when the offline algorithm is given an arbitrary controlled value oracle. Our main result is an adversarial resilience theorem for the Spiteful Greedy Swap Poisson Process (SGS...
187. A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench ​
Author: Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.HC, cs.IR, cs.LG
arXiv:2608.12138v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income s...
188. Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers ​
Author: 'I~nigo de Troya, Maurus Enbergs, Neelke Doorn, Roel Dobbe
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.SY, eess.SY
arXiv:2608.12166v1 Announce Type: cross Abstract: Algorithm registers have been championed as a means of providing transparency on the use of algorithms in public services. Yet potential publics differ in their expectations of what should be made transparent and how, as well as in their interest in ...
189. HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation ​
Author: Ruochen Li, Shuang Chen, Wenke E, Farshad Arvin, Amir Atapour-Abarghouei
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12187v1 Announce Type: cross Abstract: Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdependencies inh...
190. Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation ​
Author: Md Yassir Mottalib, Md Yousuf, Eklachur Rahman Bhuiyan, S M Ahsan Habib, Sonjoy Kumar Dey, Md. Salahuddin Gazi, Molay Kumar Roy, Asaduzzaman Anik
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.12190v1 Announce Type: cross Abstract: With the increasing complexity of cyber assaults in cloud environments, adaptable security solutions are needed that can support real-time detection and autonomous response. In this paper, we propose a reinforcement learning-based dynamic cyber defen...
191. HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks ​
Author: Zhao Su, Yuxin Xia, Haoran Li, Jun Shen, Qi Zhu, Qingguo Zhou, Binbin Yong
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12194v1 Announce Type: cross Abstract: Kolmogorov-Arnold Networks (KANs) enhance nonlinear function approximation by replacing scalar weights with learnable univariate functions. However, assigning an independent function to every connection results in substantial parameter redundancy, li...
192. M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation ​
Author: Jing Zhu, Ye Wang, Fumin Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12196v1 Announce Type: cross Abstract: Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-driven approaches often fail to exploit the rich mathematical structure inherent in medical images. We investigate whether explicit mathematical ...
193. NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation ​
Author: Jiarui Ma, Jianghan Wang, Yuheng Ma, Ziyi Zhuang, Xiaoguang Liu
Published: 8/14/2026, 4:00:00 AM
Categories: eess.SY, cs.AI, cs.SY
arXiv:2608.12197v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their reliability on simulator-facing SPICE netlist recognition and manipulation remains poorly understood and is rarely separated from high-level design reasoning. A...
194. Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment ​
Author: Jean-Pierre Busch, Guido Linden, Jan Bergmann, Lutz Eckstein
Published: 8/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2608.12198v1 Announce Type: cross Abstract: Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of automated vehicles, especially in complex environments. However, their complex nature and lack of tran...
195. Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge ​
Author: Arda Uzunoglu, Benjamin Van Durme, Daniel Khashabi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.12218v2 Announce Type: cross Abstract: Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer contexts will only help the model ...
196. SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward ​
Author: Zile Zhou, Huining Yuan, Weichen Zhang, Xinlei Chen, Xiao-ping Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12220v1 Announce Type: cross Abstract: Existing Vision-Language Models (VLMs) exhibits a critical bottleneck in robust spatial reasoning. Recent reinforcement learning (RL) methods aim to close this gap with verifiable outcomes, yet they suffer from poor credit assignment across intermedi...
197. Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification ​
Author: Kazi Nabiul Alam, Pooneh Bagheri Zadeh, Akbar Sheikh-Akbari
Published: 8/14/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV, eess.SP
arXiv:2608.12227v1 Announce Type: cross Abstract: Hyperspectral imaging (HSI) offers nondestructive assessment of fish freshness by detecting biochemical alterations across spectral bands. However, conventional deep learning approaches do not fully address the particular characteristics of HSI data,...
198. Few-Shot Ordinal Learning for Day-Wise Freshness Estimation with Hyperspectral Fish Images ​
Author: Kazi Nabiul Alam, Pooneh Bagheri Zadeh, Akbar Sheikh-Akbari
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, eess.IV, eess.SP
arXiv:2608.12230v1 Announce Type: cross Abstract: Non-destructive food quality assessment has increasingly benefited from hyperspectral imaging (HSI), which captures spectral signatures linked to biochemical changes during storage. Estimating day-wise freshness, however, remains challenging owing to...
199. How Organizations Use AI: Evidence from ChatGPT ​
Author: Aaron Chatterji, David Holtz, Neel Rakholia, Prasanna Tambe, Gawesha Weeratunga
Published: 8/14/2026, 4:00:00 AM
Categories: econ.GN, cs.AI, cs.HC, q-fin.EC
arXiv:2608.12236v1 Announce Type: cross Abstract: We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company financial data through March 2026. These linked data enable a privacy-preserving anal...
200. HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression ​
Author: Yuefeng Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM
arXiv:2608.12239v1 Announce Type: cross Abstract: Use this plain-text version for the arXiv abstract field: Learned image compression (LIC) models achieve strong rate-distortion performance but are hindered by high computational complexity and encoding-decoding mismatches across heterogeneous hardwa...
201. VICBench: A Multi-Language Benchmark for Code Vulnerability Detection ​
Author: Jin Lu, Xuening Han, Yang Zhong, Lin Tan, Kevin Luo, Andrew Gacek, Neha Rungta
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.SE
arXiv:2608.12246v1 Announce Type: cross Abstract: Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the full range of vulnerabl...
202. One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL ​
Author: Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2608.12253v1 Announce Type: cross Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: becau...
203. Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams ​
Author: Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu, Yongke Yao, Jinhao Du, Wei He, Kai Zou, Zechao Li, Jingdong Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12262v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scie...
204. Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents ​
Author: Junliang Liu, Ruoyu Li, Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Jingheng Xu, Laizhong Cui
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2608.12273v1 Announce Type: cross Abstract: LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publishers: a static skil...
205. A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery ​
Author: Rafi Ibn Sultan, Chengyin Li, Yiannos Demetriou, Ahmed I. Ghanem, Joshua P. Kim, Justine Cunningham, Hassan Bagher-Ebadian, Dongxiao Zhu, Kundan S. Thind
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12274v1 Announce Type: cross Abstract: Background: Accurate segmentation of the Left Anterior Descending (LAD) artery in 3D free-breathing, non-contrast CT is critical for cardiac dose sparing in thoracic radiotherapy. The LAD is extremely small, has poor soft-tissue contrast, and varies ...
206. Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages ​
Author: Avijit Roy, Proma Roy
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2608.12278v1 Announce Type: cross Abstract: Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization sche...
207. Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence ​
Author: Aman Tyagi, Hemanth Boinpally, Jonathan Chen, Douglas Gebert, Steven Hickson
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.MM
arXiv:2608.12290v1 Announce Type: cross Abstract: Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity cau...
208. Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations ​
Author: AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdel
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12299v1 Announce Type: cross Abstract: Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolut...
209. Redistribution-based Cost Inference Improves Sparse Safe Offline RL ​
Author: Ebenezer Gelo (University of the Witwatersrand), Geraud Nangue Tasse (University of the Witwatersrand), Steven James (University of the Witwatersrand), Benjamin Rosman (University of the Witwatersrand)
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.12306v1 Announce Type: cross Abstract: Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. We frame this as a tem...
210. AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses ​
Author: Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2608.12307v1 Announce Type: cross Abstract: Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. In this paper, we ask whether such t...
211. DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation ​
Author: Yan Deng, Fei Xu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2608.12308v1 Announce Type: cross Abstract: Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability. Although recent VLA models offer a promis...
212. Causal Agent based on Large Language Model ​
Author: Kairong Han, Kun Kuang, Ziyu Zhao, Junjian Ye, Fei Wu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL
arXiv:2408.06849v3 Announce Type: replace Abstract: The large language model (LLM) has achieved significant success across various domains. However, the inherent complexity of causal problems and causal theory poses challenges in accurately describing them in natural language, making it difficult fo...
213. On Benchmarking Human-Like Intelligence in Machines ​
Author: Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2502.20502v2 Announce Type: replace Abstract: Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-generated data, are increasingly posited as approximate models of human cognition. However, we argue that many c...
214. OpenAg: Democratizing Agricultural Intelligence ​
Author: Srikanth Thudumu, Jason Fisher
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2506.04571v3 Announce Type: replace Abstract: Agriculture is undergoing a major transformation driven by artificial intelligence (AI), machine learning, and knowledge representation technologies. However, current agricultural intelligence systems often lack contextual understanding, explainabi...
215. Deep Fictitious Play-Based Potential Differential Games for Learning Human-Like Interaction at Unsignalized Intersections ​
Author: Kehua Chen, Ryan Feng Lin, Shucheng Zhang, Yinhai Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2506.12283v2 Announce Type: replace Abstract: Modeling vehicle interactions at unsignalized intersections is a challenging task due to the complexity of the underlying game-theoretic processes. Although prior studies have attempted to capture interactive driving behaviors, most approaches reli...
216. DREAMS: Density Functional Theory Based Research Engine for Agentic Materials Simulation ​
Author: Ziqi Wang, Hongshuo Huang, Hancheng Zhao, Changwen Xu, Shang Zhu, Jan Janssen, Venkatasubramanian Viswanathan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cond-mat.mtrl-sci
arXiv:2507.14267v2 Announce Type: replace Abstract: Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical outputs are difficult to trust: agents lose context, game verification checks, and can produce large volumes of plausible yet invalid results. We i...
217. On the Definition of Intelligence ​
Author: Kei-Sing Ng
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2507.22423v3 Announce Type: replace Abstract: To engineer AGI, we should first capture the essence of intelligence in a species-agnostic form that can be evaluated, while being sufficiently general to encompass diverse paradigms of intelligent behavior, including reinforcement learning, genera...
218. SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives ​
Author: Vincent Siu, Nicholas Crispino, David Park, Nathan W. Henry, Zhun Wang, Yang Liu, Dawn Song, Chenguang Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2509.13450v3 Announce Type: replace Abstract: We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work highlights the general capabilities of representation steering, we focus on safety perspe...
219. Behavior and Representation in Open-Weight Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection ​
Author: Francesca Da Ros, Luca Di Gaspero, Kevin Roitero
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2512.13374v2 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) open new perspectives for automation in optimization, yet little is known about whether their internal representations capture problem structure or algorithmic behavior. We investigate whether represe...
220. Credo: Declarative Control of LLM Pipelines via Beliefs and Policies ​
Author: Duo Lu, Andrew Crotty, U\u{g}ur \c{C}etintemel
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.DB
arXiv:2604.14401v2 Announce Type: replace Abstract: Agentic AI systems are becoming commonplace in domains that require long-lived, stateful decision-making in continuously evolving conditions. As such, correctness depends not only on the output of individual model calls, but also on how to best ada...
221. Towards Human Motion World Models via Executable Behaviour Representations ​
Author: Rimvydas Rubavicius, Manisha Dubey, N. Siddharth, Subramanian Ramamoorthy
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2604.18064v2 Announce Type: replace Abstract: Human motion world models should capture motion's intentionality by being executable: adaptable to different actions and capable of assessing motion quality. To achieve this, we introduce a domain-specific language ExAct that represents human motio...
222. Tools as Continuous Flow for Evolving Agentic Reasoning ​
Author: Tairan Huang, Siyu Shang, Qiang Chen, Xiu Su, Yi Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2605.07339v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing methods rely on a step-wise paradigm that lacks a global perspective, which causes error accumulation over long hor...
223. Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching ​
Author: Diego Gosmar, Deborah A. Dahl
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.MA
arXiv:2605.29055v2 Announce Type: replace Abstract: This paper describes an approach to hallucination detection and mitigation using a HOPE-inspired Nested Learning architecture with Continuum Memory Systems (CMS) and semantic similarity caching, tested on a hybrid benchmark of 310 prompts (217 epis...
224. Moxia: A Trust-First Neuro-Symbolic Execution Architecture for Self-Explaining Mathematical Reasoning ​
Author: Alessio Bruno
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2606.00671v3 Announce Type: replace Abstract: We present Moxia (formerly AXIOM), a trust-first neuro-symbolic architecture for self-explaining mathematical reasoning over natural-language input. Its language model is strictly a canonicalizer: it rewrites informal problem text into a narrow sch...
225. RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit ​
Author: Amirhossein Ghaffari, Ali Goodarzi, Huong Nguyen, Simo Hosio, Lauri Lov'en, Ekaterina Gilman
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.SI
arXiv:2606.06027v2 Announce Type: replace Abstract: Community-conditioned language model adaptation needs choices about data collection, community definition, and evaluation that are currently made independently in each study, making it hard to compare assumptions or reuse artifacts. We present Redd...
226. Teaching agentic AI to learn expert reasoning for rare disease diagnosis ​
Author: Minh-Ha Nguyen, Erica Gray, Bryce A. Schuler, Kevin W. Byram, Chih-Ting Yang, Fan Ma, Hua Xu, Wu-Chen Su, Chao Yan, Wei-Qi Wei, Adam Wright, Lisa Bastarache, Josh Peterson, Lingyao Li, Siyuan Ma, Undiagnosed Diseases Network, Rizwan Hamid, Thomas A. Cassini, Cathy Shyr
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2606.16149v3 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) rank the correct disease first in only 35.4% of benchmark cases. Here we show that this expert reasoning can be ...
227. Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising ​
Author: Tianci Liu, Zihan Dong, Linjun Zhang, Haoyu Wang, Jing Gao, Emre Kiciman, Ranveer Chandra, Wei-Ting Chen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.00407v2 Announce Type: replace Abstract: Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design. Solely relying on prespecified templates or user verbose instructions, they fail to capture la...
228. The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy ​
Author: Chunzheng Zhu, Lei Tian, Bohan Tan, Ziqi Zhou, Yuxuan Sun, Yijun Wang, Chengchao Lv, Yilin Wen, Yijun He, Jinghao Lin, Yihang Chen, Chee Wei Tan, Qianshan Wei, Lei Zhao, Bin Pu, Kenli Li, Yuan Xue, Jianxin Lin
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.11175v2 Announce Type: replace Abstract: The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical imaging AI, moving it from task-specific predictors toward autonomous agents that perceive, reason, pl...
229. SportD: How do VLMs physically strategize? ​
Author: Jasin Cekinmez, Addison J. Wu, Haotian Xia, Kyumin Andrew Shim, Anay Putty, Jinglin Xiao, Zhuohan Liu, Leo Liu, Weining Shen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CV
arXiv:2607.14616v3 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset an...
230. The Human-AI Substitution Principle: When will you be replaced by AI in your organization? ​
Author: Bonny Banerjee, Shreya Singh
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, econ.GN, q-fin.EC
arXiv:2607.20781v2 Announce Type: replace Abstract: Artificial Intelligence (AI) is rapidly transforming organizations, raising a fundamental organizational and economic question: when will a human employee be replaced by AI? We present an analytical model for studying Human--AI Task Allocation (HAT...
231. A foundation model of numerical intelligence with cross-disciplinary generalization ​
Author: Chenghan Wu, Zongmin Yu, Liu Yang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2607.28432v3 Announce Type: replace Abstract: Intelligence is commonly understood as the ability to acquire and apply knowledge, adapt to unfamiliar situations and solve new problems. Large language models exhibit this capacity by inferring task-relevant knowledge from textual context and appl...
232. Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Baseline, and a Dual-Head Model ​
Author: Gospel Bassey, Samuel Bassey, Vincent Fakiyesi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.05685v2 Announce Type: replace Abstract: Most public benchmarks for machine-condition monitoring come from test rigs, where faults are induced on purpose and every event is known. Real production fields rarely offer that. They give you sensor histories with no fault log attached, which is...
233. Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains ​
Author: Hiroki Naito
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CY
arXiv:2608.07474v2 Announce Type: replace Abstract: Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains once AI output velocity V exceeds human cognitive capacity C_max. The operative constraint, however, is V x L, where L is per-item cognitive load...
234. Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability ​
Author: Alvin Spivey, Yu Huang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.LG
arXiv:2608.10300v2 Announce Type: replace Abstract: Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human reviewers may each expose rich internal states, while operational exchange requires a narrow share...
235. HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment ​
Author: Xiaokang Qu, Yiting Lin
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI, cs.DL
arXiv:2608.10584v2 Announce Type: replace Abstract: Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery. Existing scholar assessment methods predominantly rely on bibliometric indicators and reputation proxies, while recent...
236. ComBodied Agents: a New Paradigm of Human-Centric Agentic AI ​
Author: Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Yao, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.AI
arXiv:2608.10915v2 Announce Type: replace Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor wh...
237. Deep Activity Model: A Generative Approach for Human Mobility Pattern Synthesis ​
Author: Xishun Liao, Qinhua Jiang, Brian Yueshuai He, Yifan Liu, Chenchen Kuai, Jiaqi Ma
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2405.17468v3 Announce Type: replace-cross Abstract: Human mobility plays a crucial role in transportation, urban planning, and public health, but current approaches face important limitations. Existing deep learning models tend to overlook the semantic interdependencies among activities and ho...
238. ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation ​
Author: Xiaoman Zhang, Hong-Yu Zhou, Xiaoli Yang, Oishi Banerjee, Juli'an N. Acosta, Mohammed Baharoon, Josh Miller, Ouwen Huang, Pranav Rajpurkar
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL
arXiv:2411.15122v2 Announce Type: replace-cross Abstract: AI-driven models have demonstrated significant potential in automating radiology report generation for chest X-rays. However, there is no standardized benchmark for objectively evaluating their performance. To address this, we present ReXrank...
239. Explainability in Practice: A Survey of Explainable NLP Across Various Domains ​
Author: Hadi Mohammadi, Robert A. Bagheri, Anastasia Giachanou, Daniel L. Oberski
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2502.00837v3 Announce Type: replace-cross Abstract: Natural Language Processing (NLP) is now embedded in critical sectors including healthcare, finance, and customer relationship management, where models such as GPT-4o, Gemini, and BERT increasingly inform decisions. The black-box nature of th...
240. Proportional Committee Elections with Positive and Negative Votes ​
Author: Sonja Kraiczy, Georgios Papasotiropoulos, Grzegorz Pierczy'nski, Piotr Skowron
Published: 8/14/2026, 4:00:00 AM
Categories: cs.GT, cs.AI
arXiv:2503.01985v2 Announce Type: replace-cross Abstract: In the classic committee election setting each voter approves a subset of candidates and the goal is to select $k$ winners based on these preferences. A central focus of recent research in the area has been to achieve proportional representat...
241. Program Semantic Inequivalence Game with Large Language Models ​
Author: Antonio Valerio Miceli-Barone, Vaishak Belle, Ali Payani
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.PL
arXiv:2505.03818v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-trivial reasoning about program semantics. Finding training examples to teach LLMs to solve these tasks ...
242. COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification ​
Author: Mariano Rivera, Angello Hoyos
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2505.18315v3 Announce Type: replace-cross Abstract: We introduce \textbf{CoLoRA} (Convolutional Low-Rank Adaptation), a parameter-efficient fine-tuning method for convolutional neural networks (CNNs). CoLoRA extends LoRA to convolutional layers by decomposing kernel updates into lightweight de...
243. P2MFDS: A Privacy-Preserving Multimodal Fall Detection System for Elderly People in Bathroom Environments ​
Author: Haitian Wang, Yiren Wang, Xinyu Wang, Yumeng Miao, Yuliang Zhang, Yu Zhang, Atif Mansoor
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2506.17332v2 Announce Type: replace-cross Abstract: By 2050, people aged 65 and over are projected to make up 16% of the global population. As aging is closely associated with increased fall risk, particularly in wet and confined environments such as bathrooms where over 80% of falls occur. Al...
244. Small Data Explainer -- The impact of small data methods in everyday life ​
Author: Maren Hackenberg, Sophia G. Connor, Fabian Kabus, June Brawner, Ella Markham, Mahi Hardalupas, Areeq Chowdhury, Rolf Backofen, Anna K"ottgen, Angelika Rohde, Nadine Binder, Harald Binder, the Collaborative Research Center 1597 Small Data
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2507.11773v2 Announce Type: replace-cross Abstract: The emergence of breakthrough artificial intelligence (AI) techniques has led to a renewed focus on how small data settings, i.e., settings with limited information, can benefit from such developments. This includes societal issues such as ho...
245. Commonsense on Demand: Generating and Selectively Integrating Commonsense Knowledge for Natural Language Inference ​
Author: Chathuri Jayaweera, Brianna Yanqui, Bonnie J. Dorr
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2507.15100v3 Announce Type: replace-cross Abstract: Natural Language Inference (NLI) determines whether a premise entails, contradicts, or is neutral with respect to a hypothesis. The task is often framed as emulating human inference, in which commonsense knowledge plays a major role. This stu...
246. Quantization-Aware Neuromorphic Architecture for Skin Lesion Classification on Resource-Constrained Devices ​
Author: Haitian Wang, Xia Cheng, Xinyu Wang, Fiona Wei, Zichen Geng
Published: 8/14/2026, 4:00:00 AM
Categories: eess.IV, cs.AI, cs.CV
arXiv:2507.15958v5 Announce Type: replace-cross Abstract: On-device skin lesion analysis is constrained by the compute and energy cost of conventional CNN inference and by the need for lightweight calibration under clinical data shift. Neuromorphic processors provide event-driven sparse computation,...
247. Empowering Children to Create AI-Enabled Augmented Reality Experiences ​
Author: Lei Zhang, Shuyao Zhou, Amna Liaqat, Tinney Mak, Brian Berengard, Emily Qian, Andr'es Monroy-Hern'andez
Published: 8/14/2026, 4:00:00 AM
Categories: cs.HC, cs.AI, cs.GR, cs.PL
arXiv:2508.08467v2 Announce Type: replace-cross Abstract: Despite their potential to enhance children's learning experiences, AI-enabled AR technologies are predominantly used in ways that position children as consumers rather than creators. We introduce Capybara, an AR-based and AI-powered visual p...
248. Ethics Practices in AI Development: An Empirical Study Across Roles and Regions ​
Author: Wilder Baldwin, Sepideh Ghanavati, Manuel Woersdoerfer
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC, cs.SE
arXiv:2508.09219v3 Announce Type: replace-cross Abstract: Recent advances in AI applications have raised growing concerns about the need for ethical guidelines and regulations to mitigate the risks posed by these technologies. In this paper, we present a mixed-methods survey study - combining statis...
249. CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D ​
Author: Mohamad Amin Mirzaei, Pantea Amoie, Ali Ekhterachian, Matin Mirzababaei, Babak Khalaj
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2509.24528v4 Announce Type: replace-cross Abstract: Object retrieval from a scene has become a new trend of research due to its numerous applications. Recent approaches achieve zero-shot, open-vocabulary 3D semantic mapping by assigning embedding vectors to 2D class-agnostic masks generated vi...
250. Adaptive Online Learning with LSTM Networks for Energy Price Prediction ​
Author: Salih Salihoglu, Ibrahim Ahmed, Afshin Asadi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2510.16898v2 Announce Type: replace-cross Abstract: Accurate prediction of electricity prices is crucial for stakeholders in the energy market, particularly for grid operators, energy producers, and consumers. This study focuses on developing a predictive model leveraging Long Short-Term Memor...
251. LiDAR-based 3D Change Detection at City Scale ​
Author: Hezam Albaqami, Haitian Wang, Xinyu Wang, Muhammad Ibrahim, Zainy M. Malakan, Abdullah M. Algamdi, Mohammed H. Alghamdi, Ajmal Mian
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.RO
arXiv:2510.21112v3 Announce Type: replace-cross Abstract: High-definition 3D city maps enable city planning and change detection, which is essential for municipal compliance, map maintenance, and asset monitoring, including both built structures and urban greenery. Conventional Digital Surface Model...
252. MicroAUNet: Boundary-Enhanced Multi-scale Fusion with Knowledge Distillation for Colonoscopy Polyp Image Segmentation ​
Author: Ziyi Wang, Yuanmei Zhang, Baoying Ye, Yimei Jiang, Leilei Gu, Suncheng Xiang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2511.01143v2 Announce Type: replace-cross Abstract: Early and accurate segmentation of colorectal polyps is critical for reducing colorectal cancer mortality, which has been extensively explored by academia and industry. However, current deep learning-based polyp segmentation models either com...
253. BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents ​
Author: Kaiyuan Zhang, Mark Tenenholtz, Kyle Polley, Jerry Ma, Denis Yarats, Ninghui Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2511.20597v2 Announce Type: replace-cross Abstract: The integration of artificial intelligence (AI) agents into web browsers introduces security challenges that go beyond traditional web application threat models. Prior work has identified prompt injection as a new attack vector for web agents...
254. A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation ​
Author: Xiaocan Li, Shiliang Wu, Zheng Shen
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC
arXiv:2512.06547v4 Announce Type: replace-cross Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL setting. Decoupled loss used in decoupled PPO improves coupled-loss style of algorithms' (e.g., standard ...
255. Probably Approximately Correct Maximum A Posteriori Inference ​
Author: Matthew Shorvon, Frederik Mallmann-Trenn, David S. Watson
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.16083v2 Announce Type: replace-cross Abstract: Computing the conditional mode of a distribution, better known as the maximum a posteriori (MAP) assignment, is a fundamental task in probabilistic inference. However, MAP is generally intractable, and remains hard even under many common stru...
256. Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models ​
Author: Moritz Miller, Florent Draye, Bernhard Sch"olkopf
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2602.04718v5 Announce Type: replace-cross Abstract: A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not ...
257. LLM-Powered Automatic Translation and Urgency in Crisis Scenarios ​
Author: Belu Ticona, Antonis Anastasopoulos
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2602.13452v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly proposed for crisis preparedness and response, particularly for multilingual communication. However, their suitability for high-stakes crisis contexts remains insufficiently evaluated. This work e...
258. How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers? ​
Author: Pritam Sil, Durgaprasad Karnam, Vinay Reddy Venumuddala, Pushpak Bhattacharyya
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.CL
arXiv:2603.00056v2 Announce Type: replace-cross Abstract: STEM Mental models can play a critical role in assessing students' conceptual understanding of a topic. They not only offer insights into what students know but also into how effectively they can apply, relate to, and integrate concepts acros...
259. Post-Training with Policy Gradients: Optimality and the Base Model Barrier ​
Author: Alireza Mousavi-Hosseini, Murat A. Erdogdu
Published: 8/14/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2603.06957v2 Announce Type: replace-cross Abstract: We study post-training linear autoregressive models with outcome and process rewards. Given a context $\boldsymbol{x}$, the model must predict the response $\boldsymbol{y} \in Y^N$, a sequence of length $N$ that satisfies a $\gamma$ margin co...
260. Representation Finetuning for Continual Learning ​
Author: Haihua Luo, Xuming Ran, Tommi K"arkk"ainen, Huiyan Xue, Zhonghua Chen, Qi Xu, Fengyu Cong
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2603.11201v3 Announce Type: replace-cross Abstract: The world is inherently dynamic, and continual learning aims to enable models to adapt to ever-evolving data streams. While pre-trained models have shown powerful performance in continual learning, they still require finetuning to adapt effec...
261. A Simple Efficiency Incremental Learning Framework via Vision-Language Model with Nonlinear Multi-Adapters ​
Author: Haihua Luo, Xuming Ran, Jiangrong Shen, Timo H"am"al"ainen, Zhonghua Chen, Qi Xu, Fengyu Cong
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI
arXiv:2603.11211v3 Announce Type: replace-cross Abstract: Incremental Learning (IL) aims to learn new tasks while preserving previously acquired knowledge. Integrating the zero-shot learning capabilities of pre-trained vision-language models into IL methods has marked a significant advancement. Howe...
262. Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation ​
Author: Petter T"ornberg
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2603.13891v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderation and hiring. Across 19 LLMs and two experiments totaling more than 4 million annotation judgments, w...
263. VLM2Rec: Resolving Modality Collapse in Vision-Language Model Embedders for Multimodal Sequential Recommendation ​
Author: Junyoung Kim, Woojoo Kim, Wonbin Kweon, Jaehyung Lim, Dongha Kim, Hwanjo Yu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI
arXiv:2603.17450v2 Announce Type: replace-cross Abstract: Sequential Recommendation (SR) in multimodal settings typically relies on small frozen pretrained encoders, which limits semantic capacity and prevents Collaborative Filtering (CF) signals from being fully integrated into item representations...
264. REVERE: Reflective Evolving Research Engineer ​
Author: Balaji Dinesh Gangireddi, Aniketh Garikaparthi, Manasi Patwardhan, Arman Cohan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI
arXiv:2603.20667v2 Announce Type: replace-cross Abstract: Existing prompt-optimization techniques rely on local signals, causing poor generalization across tasks. In addition, they also rely on weak update mechanisms, such as full-prompt rewrites or unstructured merges, which cause knowledge loss an...
265. Designing Agentic AI-Based Screening for Portfolio Investment ​
Author: Mehmet Caner, Agostino Capponi, Nathan Sun, Jonathan Y. Tan
Published: 8/14/2026, 4:00:00 AM
Categories: q-fin.PM, cs.AI, cs.MA, q-fin.ST
arXiv:2603.23300v2 Announce Type: replace-cross Abstract: We introduce a new agentic artificial intelligence (AI) platform for portfolio management. Our architecture consists of three layers. First, two large language model (LLM) agents are assigned specialized tasks: one agent screens for firms wit...
266. Evaluation and Hardening of LLM System Instructions Against Extraction via Encoding Attacks ​
Author: Anubhab Sahu, Diptisha Samanta, Reza Soosahabi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2604.01039v4 Announce Type: replace-cross Abstract: System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensitive operational context in agentic AI applications. These instructions may contain sensitive informatio...
267. Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic Prompting ​
Author: Roland M"uhlenbernd
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.02512v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly exhibit human-like patterns of pragmatic and social reasoning. This paper addresses two related questions: do LLMs approximate human social meaning not only qualitatively but also quantitatively, and ...
268. Uncertainty as a Planning Signal: Multi-Turn Decision Making for Goal-Oriented Conversation ​
Author: Xinyi Ling, Ye Liu, Reza Averly, Xia Ning
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.03924v2 Announce Type: replace-cross Abstract: Goal-oriented conversational systems require making sequential decisions under uncertainty about the user's intent, where the algorithm must balance information acquisition and target commitment over multiple turns. Existing approaches addres...
269. TEMPER: Testing Emotional Perturbation in Quantitative Reasoning ​
Author: Atahan Dokme, Benjamin Reichman, Larry Heck
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.07801v2 Announce Type: replace-cross Abstract: Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, real-world queries are often wrapped in frustration, urgency or enthusiasm. Does emotional framing alone ...
270. DORA Explorer: Improving the Exploration Ability of LLMs Without Training ​
Author: Priya Gurjar, Md Farhan Ishmam, Kenneth Marino
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2604.17244v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents for sequential decision-making struggle to produce diverse outputs. This leads to insufficient exploration, suboptimal solutions, and repeated actions. Actions are generated at the sequence level, but existin...
271. Making Gaussian Kolmogorov-Arnold Networks Reliable and Accurate ​
Author: Amir Noorizadegan, Sifan Wang, Leevan Ling
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CE, cs.AI, math.AP
arXiv:2604.21174v3 Announce Type: replace-cross Abstract: Kolmogorov-Arnold Networks (KANs) replace fixed activations with learnable univariate edge functions whose behavior depends strongly on the chosen basis. Gaussian radial basis functions provide a simple and efficient alternative to splines, b...
272. Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents ​
Author: Benjamin Probst, Andreas Happe, J"urgen Cito
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2604.27143v2 Announce Type: replace-cross Abstract: Cloud-based Large Language Models (LLMs) can perform autonomous penetration-testing sub-tasks such as Linux privilege escalation, but raise security, privacy, and sovereignty concerns. Locally hosted open-weight models avoid these issues, yet...
273. Analytic Bridge Diffusions for Controlled Path Generation ​
Author: Michael Chertkov
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.stat-mech, cs.AI, cs.SY, eess.SY, math.OC
arXiv:2605.02961v2 Announce Type: replace-cross Abstract: Most modern bridge-diffusion methods achieve finite-time transport by specifying an interpolation, Schrodinger-bridge, or stochastic-control objective and then learning the associated score or drift field with a neural network. In contrast, w...
274. CAR: Query-Guided Confidence-Aware Reranking for Retrieval-Augmented Generation ​
Author: Zhipeng Song, Yizhi Zhou, Xiangyu Kong, Jiulong Jiao, Xuezhou Ye, Chunqi Gao, Xueqing Shi, Yu Wang, Yuhang Zhou, Heng Qi
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2605.04495v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) relies on evidence ranking to determine what information is exposed to the generator, yet existing retrieval and reranking methods primarily estimate query--document relevance. Relevance, however, is not e...
275. Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement ​
Author: Nicholas S. Kersting, Vittorio Castelli, Chieh Ting Yeh, Xinzhu Wang, Saad Taame, Khaoula Allak
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY
arXiv:2605.05103v3 Announce Type: replace-cross Abstract: We introduce the \textbf{Concept Field} of a text corpus: a local drift field with pointwise uncertainty, estimated in sentence-embedding space from the deltas between consecutive sentences. Given a candidate sentence transition, we score its...
276. Pretraining large language models with MXFP4 on Native FP4 Hardware ​
Author: Musa Cim, Sarthak Arora, Poovaiah Palangappa, Miro Hodak, Ravi Dwivedula, Meena Arunachalam, Mahmut Taylan Kandemir
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.09825v4 Announce Type: replace-cross Abstract: Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable? We address this question through a controlled study of MXFP4 quantization in transformer trainin...
277. Cavity-Enhanced Collective Quantum Processing with Polarization-Encoded Qubits ​
Author: Kamil Wereszczy'nski, J'ozef Cyran, Adam Brzezowski, Dawid Za{\l}u.zny, Robert Potoniec, Kasper Wi'sniowski, Agnieszka Michalczuk
Published: 8/14/2026, 4:00:00 AM
Categories: quant-ph, cs.AI
arXiv:2605.10473v2 Announce Type: replace-cross Abstract: We introduce a cavity-enhanced optical architecture for collective quantum processing in which logical qubits are encoded in the polarization subspace of recirculating intracavity modes. The physical carrier and computational degree of freedo...
278. TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning ​
Author: Matthew M. Hong, Jesse Zhang, Anusha Nagabandi, Abhishek Gupta
Published: 8/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2605.12236v2 Announce Type: replace-cross Abstract: Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with behavioral cloning (BC), which produces narrow action distributions that lack the coverage necessary for do...
279. memorywire: A Vendor-Neutral Wire Format for Agent Memory Operations ​
Author: Thamilvendhan Munirathinam
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.DC
arXiv:2606.01138v4 Announce Type: replace-cross Abstract: Agent-memory frameworks -- mem0, Letta/MemGPT, Cognee, Zep/Graphiti, MemoryOS, MemTensor -- each ship their own SDK, storage layout, and operational vocabulary. There is no shared wire format: every integration is bespoke, every migration reb...
280. Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association ​
Author: Matvei Shelukhan, Timur Mamedov, Aleksandr Chukhrov, Karina Kvanchiani
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2606.02022v2 Announce Type: replace-cross Abstract: Multi-view object association is an important computer vision problem that underlies many multi-camera perception tasks. While this task is naturally formulated as a constrained one-to-one matching problem, recent works heavily rely on pairwi...
281. Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning ​
Author: Xuekang Wang, Zhuoyuan Hao, Shuo Hou, Hao Peng, Juanzi Li, Xiaozhi Wang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2606.04923v2 Announce Type: replace-cross Abstract: Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe ...
282. ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset ​
Author: Luciano Duarte, Olga Ovcharenko, Sebastian Schelter
Published: 8/14/2026, 4:00:00 AM
Categories: cs.DB, cs.AI
arXiv:2606.09648v2 Announce Type: replace-cross Abstract: Multi-modal data management has emerged as a central research topic in the database community, spanning data integration, semantic query processing, and data quality assessment. Despite this growing interest, the community lacks large-scale, ...
283. FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning ​
Author: Steven Oh, Jason Jingzhou Liu, Tony Tao, Philip Han, Kenneth Shaw, Satoshi Funabashi, Ruslan Salakhutdinov, Deepak Pathak
Published: 8/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG, cs.SY, eess.SY
arXiv:2606.12406v2 Announce Type: replace-cross Abstract: Contact-rich manipulation requires force sensitivity, but many robot arms lack dedicated force sensors due to their high cost. We present Neural External Torque Estimation (NEXT), a data-driven method that estimates external joint torques wit...
284. ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory ​
Author: Habibullah Akbar
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.25156v4 Announce Type: replace-cross Abstract: Length extrapolation in language models involves competing objectives: retrieval fidelity, long-document likelihood, short-context quality, and inference cost. We present ATMA, a 378M-parameter hybrid recipe that combines Polar Attention with...
285. Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop ​
Author: Chenmu Zhang, Boris I. Yakobson
Published: 8/14/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.AI, cs.LG
arXiv:2606.29717v2 Announce Type: replace-cross Abstract: Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. A decade of work has produced standard public benchmarks and many published machine-learning models for the task (D...
286. From World Models to World Action Models: A Concise Tutorial for Robotics ​
Author: Xiaoxiong Zhang, Xiong Zeng, Wei Zhang
Published: 8/14/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.SY, eess.SY
arXiv:2607.00836v5 Announce Type: replace-cross Abstract: Rather than providing an exhaustive survey, this paper presents a concise tutorial on world models and world action models for robotics. After reading the tutorial, readers should have a clear understanding of what constitutes a "world", how ...
287. Prompt-Driven Exploration ​
Author: Sunshine Jiang, John Marangola, David Zhang, Raghuram Kowdeed, Ruiyang Luo, Nitish Dashora, Richard Li, Pulkit Agrawal, Zhang-Wei Hong
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.08837v2 Announce Type: replace-cross Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject stochasticity in the action space, but such jitter only yields rollouts close to the original. Escap...
288. LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes ​
Author: Michael Solodko, Steven Gong, Guangwei Yu, Satya Krishna Gorti, Jesse C. Cresswell, Victor Zhong
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2607.12310v3 Announce Type: replace-cross Abstract: While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Answering questions over enterprise and scientific data lakes requires systems to navigate heterogeneous, ...
289. Reducing Per-Sample Interference in Stochastic Optimization ​
Author: Apostolos Avranas
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2607.16261v2 Announce Type: replace-cross Abstract: Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments. While effective, this standard practice can produce parameter updates that actively increase the loss of...
290. Cryptographically verifiable authorization for autonomous AI agents: A falsifiable hypothesis and proof-of-concept ​
Author: M. Llamb'i-Morillas, D. Fern'andez-Fern'andez
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CR, cs.AI
arXiv:2607.21325v2 Announce Type: replace-cross Abstract: Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limited human oversight. Existing authentication and authorization mechanisms establish identity and delegate authority, but do not inher...
291. Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking ​
Author: Peijia Guo, Wenxuan Xie, ZiGuang Li, Ming Li
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.03859v2 Announce Type: replace-cross Abstract: Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexplored and largely unresolved challenge. Prior work on LLM-generated-text detection targets AI involveme...
292. Continual Learning in Transition ​
Author: Zhiyan Hou, Dan Zhang, Tao Feng, Liyuan Wang, Wei Li, Xiangzhao Hao, Hongyan An, Junfeng Fang, Haokai Ma, Zhaohui Xu, Xinyu Tang, Haiyun Guo, Jinqiao Wang, Tat-Seng Chua
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06216v2 Announce Type: replace-cross Abstract: Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architectural designs, and weight adaptation. However, emerging paradig...
293. ED-CSP: Crystal Structure Prediction from Electron Diffraction ​
Author: Germain Poloudenny, Ya"el Fr'egier, Arnaud Demorti`ere
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2608.06448v2 Announce Type: replace-cross Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. Existing ED-based learning methods mainly predict crystallographic labels, reconstruct struc...
294. Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations? ​
Author: Osvaldo Quinjica, Eric Bennett, Xinchen Yang, Andrew Schonebaum, Marine Carpuat
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.08283v2 Announce Type: replace-cross Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation met...
295. Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol ​
Author: Christoph Trattner
Published: 8/14/2026, 4:00:00 AM
Categories: cs.HC, cs.AI
arXiv:2608.08882v2 Announce Type: replace-cross Abstract: AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after using such a tool, what can the user still do on their own? I call this epistemic transfer. It refers t...
296. Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction ​
Author: Jingxian Xu, Yuhao Huang, Rusi Chen, Yanfeng Zhou, Dong Ni
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.09182v2 Announce Type: replace-cross Abstract: Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Existing localization methods have advanced, among which multi-stage refinement is a superior solution. Alth...
297. Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning ​
Author: Yuting Liu, Wei Wu, Jianzhe Zhao, Guibing Guo
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.09507v2 Announce Type: replace-cross Abstract: Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full preference...
298. From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models ​
Author: Si'an Xie, Jiaxun Liu, Biao Yang, Wei Yuan, Fan Yang, Tingting Gao, Ming Wu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10444v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively unexamined capab...
299. Persistent Recursive Worlds Enable Autonomous Software Evolution ​
Author: Beichen Huang, Zhenyu Liang, Bowen Zheng, Ran Cheng
Published: 8/14/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.MA, cs.NE
arXiv:2608.10450v2 Announce Type: replace-cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity through persistent sessions, memories, managers or shared context. We introduce EvoX G...
300. Inferential Capability Does Not Determine Legal Scope ​
Author: Nicola Fabiano
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.10601v2 Announce Type: replace-cross Abstract: Two instruments of EU digital law place inference at their centre and mean different things by it. Article 3(1) of the AI Act uses the capability to infer constitutively: it is the central feature separating the regulated category from conven...
301. ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering ​
Author: Akrin Zheng, Alexander Wu, Alaia Liu
Published: 8/14/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.CL
arXiv:2608.10679v2 Announce Type: replace-cross Abstract: Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are work by-products in which required organizational relations remain implicit across heterogeneou...
302. Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization ​
Author: Tal Oved, Roi Pony, Oshri Naparstek, Udi barzelay
Published: 8/14/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.NE
arXiv:2608.10694v2 Announce Type: replace-cross Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs an answering LLM over a validation set, so the evaluator's price tier dictates total search cost. We re...
303. Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization ​
Author: Swarnim Maheshwari, Syed Imam Ali, Vineeth N. Balasubramanian
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2608.10798v2 Announce Type: replace-cross Abstract: Most image colorization systems operate in $Lab$ space by predicting chroma ($ab$) while preserving an input-derived luminance channel ($L$). While effective on standard benchmarks, this fixed-luminance design restricts brightness changes and...
304. Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse ​
Author: Zhenyan Zheng, Yunyao Zhang, Junxi Sheng, Junqing Yu, Zikai Song
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10810v2 Announce Type: replace-cross Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions. Existing emotion benchmarks mainly annotat...
305. Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation ​
Author: Chris Han, Pengzhi Gao, Pei Fu, Jian Luan
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CL, cs.AI
arXiv:2608.10812v2 Announce Type: replace-cross Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that av...
306. Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis ​
Author: Benjamin Faveri (CEIMIA, Carleton University), Brie Bhasin (University of Ottawa)
Published: 8/14/2026, 4:00:00 AM
Categories: cs.CY, cs.AI
arXiv:2608.11006v2 Announce Type: replace-cross Abstract: Governments worldwide have responded to the rapid expansion of AI by publishing national and regional AI strategies. Comparing national and regional AI strategies to identify their convergences and divergences can uncover their common practic...