arXiv cs.LG - 2026-07-28 ​
488 items collected.
1. Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B ​
Author: Tejasvi C. Addagada
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CR
arXiv:2607.22545v1 Announce Type: new Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail addresses in a single ...
2. CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents ​
Author: Mingwei Zheng, David OBrien, Siwei Cui, Pardis Pashakhanloo, Rajdeep Mukherjee, Myeongsoo Kim, Sachit Kuhar
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SE
arXiv:2607.22711v1 Announce Type: new Abstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. However, the conventional append-only trajectory architecture found in practice tightly couples file-read ac...
3. CausalGate: Causal Importance Distillation for Transformer Module Pruning ​
Author: Kiran Nair, Smriti Regmi, Rodrigue Rizk
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CL, cs.CV, stat.ML
arXiv:2607.22720v1 Announce Type: new Abstract: Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitudes, to drop redundant modules. However, these correlation-based metrics often fail to capture subtle, ...
4. Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks ​
Author: Kaibing Yang, Guangfeng Cai, Shengtian Yang, Shuo He, Yu Li, Mengyi Liu, Pengwei Chen, Jun Xu, Lei Feng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22724v1 Announce Type: new Abstract: Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group. However, on difficult long-horizon tasks, this comparison can suffer fr...
5. QFedPolyp: A Communication- and Inference-Efficient Federated Learning Framework for Polyp Segmentation ​
Author: Madan Baduwal, Priyanka Paudel
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2607.22743v1 Announce Type: new Abstract: Background and Objective: Automatic polyp segmentation supports computer-aided diagnosis and early colorectal cancer detec- tion. Centralized deep learning requires hospitals to share sensitive medical data, while federated learning preserves privacy b...
6. Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence ​
Author: Zhaowen Fan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE
arXiv:2607.22748v1 Announce Type: new Abstract: Modern neural networks primarily adapt through parameter modification within predefined computational structures. While recent methods introduce modularity, conditional computation, and parameter-efficient adaptation, they generally do not distinguish ...
7. Hierarchical Grading in Large Language Models ​
Author: T. Shaska
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22757v1 Announce Type: new Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, and the training object...
8. An Integrated Deep Learning and Statistical Framework for Whole-Network Gene--Environment Association with Leaf Vascular Architecture ​
Author: Geran Zhao, Yangsheng Wang, Xiaotian Dai, Guifang Fu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, stat.ML
arXiv:2607.22763v1 Announce Type: new Abstract: Leaf veins exhibit remarkable diversity in architecture and patterning, yet existing gene--environment association studies have primarily quantified leaf venation using a small collection of low-dimensional summary traits, thereby discarding most of th...
9. Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation ​
Author: Yunting Song, Matthew Watson, Peter Grabowski, Jun Qin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22766v1 Announce Type: new Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumulate hidden structural contradictions, safety risks, and systemic human an...
10. DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning ​
Author: He Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22769v1 Announce Type: new Abstract: The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data. Existing dynamic data scheduling methods face critical limitations in industrial-scale pretraining and supervised fine-t...
11. Dementia Etiology Diagnosis via Collaborative Meta Knowledge Enhancement ​
Author: Siyuan Du, Mengxi Chen, Xinyang Jiang, Zilong Wang, Jiangchao Yao, Dongsheng Li, Ya Zhang, Lili Qiu, Yanfeng Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2607.22770v1 Announce Type: new Abstract: Although artificial intelligence (AI) has shown promising performance in several medical tasks, accurate dementia etiology diagnosis with AI remains challenging due to complex overlapping symptoms among diseases. Scaling up the dataset size by combinin...
12. CC-AOS: Cost- and Horizon-Conditioned Amortized Backward Induction for Finite-Horizon Optimal Stopping ​
Author: Tianwei Yu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22774v1 Announce Type: new Abstract: Finite-horizon optimal stopping is a central problem in early time-series classification, where a system must decide at each sequence prefix whether the expected benefit of another observation justifies its acquisition cost. Existing data-driven backwa...
13. Predicting the Outcome of rTMS Depression Therapy using EEG Signals and CNN ​
Author: Wael Korani, Md Fahimul Kabir Chowdhury, Sadam AlQadi, Priyan Malarvizhi kumar, Reza Rostami, Reza Kazemi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22776v1 Announce Type: new Abstract: Repetitive transcranial magnetic stimulation (rTMS) is a non invasive therapy for Major Depressive Disorder (MDD). In this study, we generate images using two time frequency methods to represent EEG signals: Fourier-Bessel Series Expansion with Euclide...
14. LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning ​
Author: Chen Wang, Boming Kang, Qinghua Cui
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22777v1 Announce Type: new Abstract: Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along amino-acid sequences, their training objectives do not explicitly constrain the model to learn three-dimensional re...
15. Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control ​
Author: Federico Del Pup, Elisa Tentori, Manfredo Atzori
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22779v1 Announce Type: new Abstract: Hand gesture recognition via surface electromyography (sEMG) is fundamental to prosthetic control. In this field, deep learning approaches have become the gold standard. However, current architectures struggle to scale; model performance typically decr...
16. What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation ​
Author: Minwoo Yu, Young-guk Ha
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22781v1 Announce Type: new Abstract: High task performance does not show whether a model retains prediction-relevant structural information in its internal representation. Temporal graph models, for example, can achieve high future-link AUC while basic graph statistics remain difficult to...
17. Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs ​
Author: Ilia Sobakinskikh, Paul Alexander Bilokon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.AR, cs.DC, cs.PF, stat.CO
arXiv:2607.22786v1 Announce Type: new Abstract: In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series. The financial time series are price series such as asset prices. Unf...
18. FMOPF: Latent Flow Matching with Constraint-Aware Interaction Priors for AC Optimal Power Flow ​
Author: Zhilin Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2607.22788v1 Announce Type: new Abstract: AC optimal power flow determines the minimum-cost generation dispatch under nonlinear power balance constraints and is solved thousands of times daily in electricity market operations. Learning a direct mapping from load conditions to OPF solutions can...
19. Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial Training ​
Author: Ali Tabaraei, Federico Simonetta, Stavros Ntalampiras
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.SD
arXiv:2607.22794v1 Announce Type: new Abstract: Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arising from inter-speaker variability. To address this critical issue, we present the first patient-independent multi...
20. Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis ​
Author: Yuntong Chen, Jianyu Liu, Guobin Zhao, Ziang Wang, Chao Chen, Ju Huang, Xitian Tian, Lijiang Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22797v1 Announce Type: new Abstract: Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against physical reality before it is acted upon. Current intelligent fault diagnosers fail this standard in tw...
21. LithoFormer: A Robust Framework for Stratigraphic Inference via Transformers ​
Author: Shwetha Salimath, Francesca Bugiotti, Sylvain Wlodarczyk, Sohaib Ouzineb
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22804v1 Announce Type: new Abstract: Accurate geological characterization of subsurface reservoirs from well log data is essential to support projects such as carbon capture and storage (CCS), geothermal development, and extraction of natural resources. Existing automated techniques for g...
22. OrchNAS: Orchestrated Neural Architecture Search Service for Personalised Federated Edge Intelligence ​
Author: Keya Patel, Sajib Mistry, Sheik Mohammad Mostakim Fattah, Aneesh Krishna
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22805v1 Announce Type: new Abstract: We propose OrchNAS, an energy-aware, personalised, federated edge intelligence framework that leverages a Neural Architecture Search Service to automatically design service-adaptive models for heterogeneous edge environments. The framework orchestrates...
23. SLA-Constrained Carbon-Aware Routing in Geo-Distributed Serverless Clouds ​
Author: Anmol Chaudhary (Department of Electronics,Computer Engineering, NIAMT Ranchi), Rahul Mishra (Department of Electronics,Computer Engineering, NIAMT Ranchi)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.DC
arXiv:2607.22806v1 Announce Type: new Abstract: Modern cloud deployments distribute applications across multiple geographic regions, yet standard routing mechanisms prioritize latency while ignoring the fluctuating carbon intensity of local power grids. Latency-driven routing incurs avoidable carbon...
24. From Hybrid Mechanistic--Data-Driven Modeling Toward Neuro-Symbolic AI: What, Why, and How ​
Author: Moein E. Samadi, Andreas Schuppert
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.LO, stat.ML
arXiv:2607.22811v1 Announce Type: new Abstract: Hybrid mechanistic/data-driven models, which combine first-principles with learned components, are increasingly used in process engineering and scientific machine learning. Common hybrid modeling designs are specified primarily through their architectu...
25. MEMENTO: Memory-Guided Memetic Code-as-Policy Evolution ​
Author: Alkis Sygkounas, Victor Aregbede, Amy Loutfi, Andreas Persson
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22832v1 Announce Type: new Abstract: Long-horizon embodied tasks require policies that execute many dependent actions before task success can be observed. Representing policies as executable control pro- grams (code-as-policy) enables their decision logic to be inspected and revised after...
26. Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias ​
Author: Ofek I. Cohen, Lior Shani, Aviv Rosenberg, Ankur Samanta, Tal Wagner, Yonathan Efroni
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.22837v1 Announce Type: new Abstract: Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns around sensitive data. However, such adaptation remains non-trivial: it often requires operationally ...
27. Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations ​
Author: Kazi Kamruzzaman Rabbi, Md. Zami Al Zunaed Farabe, M. Sohel Rahman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2607.22872v1 Announce Type: new Abstract: Post-training quantization (PTQ) has become a practical solution for deploying deep learning models on resource-constrained edge devices by compressing high-precision floating-point weights into low-precision representations without requiring retrainin...
28. Spatial Prediction of Soil Microplastics and Organic Matter Using Graph Attention Networks ​
Author: Anik Dev Nath, Md Al Amin, Bikash Kumar Paul
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22875v1 Announce Type: new Abstract: Accurate estimation of soil microplastics and organic matter is essential to assess ecosystem health and support sustainable land use. This study presents a graph-based deep learning approach using Graph Attention Networks (GATs) to model spatial depen...
29. Efficient Learning of Truncated Boolean Product Distributions: Influence to the Rescue ​
Author: Rohan Chauhan, Ioannis Panageas
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.DS, stat.ML
arXiv:2607.22889v1 Announce Type: new Abstract: Learning the natural parameters $z \in \mathbb{R}^n$ of discrete distributions $\mu_z$ from independent samples constrained to a subset $S \subseteq {0,1}^n$ is a foundational challenge in high-dimensional statistics. Existing methods for efficiently...
30. Learning from the Descent Direction: Adaptive Gradient Descent under One-Sided H\"older Regularity ​
Author: Arzu Ahmadova, Ismail Huseynov
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.OC
arXiv:2607.22906v1 Announce Type: new Abstract: We study adaptive gradient descent for continuously differentiable, possibly nonconvex objectives under one-sided H"older regularity. Unlike classical H"older- or Lipschitz-gradient assumptions, which control the full gradient variation, our conditio...
31. Beyond Directed Acyclic Graphs: Causal Zeros and Causal Differential Equations ​
Author: Sergei V. Kalinin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22910v1 Announce Type: new Abstract: Pearl's structural causal model (SCM) framework, built on directed acyclic graphs (DAGs) and the do-calculus, is the dominant formal language for causal reasoning. Yet it carries two structural restrictions: every relationship must be pre-specified as ...
32. Hidden Boundary Motion in Transformer Optimization: Function-Space Orthogonalization of Affine Weight and Bias Updates ​
Author: Zhang Gongyue, Sheng Yixuan, Liu donghan, Wang Zhiyong, Ren Weihong, Liu honghai
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22927v1 Announce Type: new Abstract: Weights and biases are normally optimized as separate parameter tensors, yet they do not represent separate functions when the input to an affine layer has nonzero mean. For an affine map $z=Wx+b$ with input mean $\mu$, a weight update contains a sampl...
33. Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety ​
Author: Domenic Rosati, Ali Dadsetan, Hong Huang, Xijie Zeng, Hassan Chowdhry, Subhabrata Majumdar, Hassan Sajjad, Frank Rudzicz
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22929v1 Announce Type: new Abstract: A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produce hate speech. Preventing such harmful fine-tuning while retaining benign adaptability remains diffi...
34. Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions ​
Author: Quyen Tran, Hai Nguyen, Quan Dao, Zhuowei Li, Nam Le, Trung Le, Dimitris Metaxas
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2607.22931v1 Announce Type: new Abstract: Analytic Continual Learning (ACL) offers a computationally efficient alternative to gradient-based approaches. Recent ACL methods are based on Recursive Least Squares (RLS) and have achieved the state-of-the-art results compared to other alternatives. ...
35. Discrepancy-Rounded Fair Bandits with Static and Time-Varying Exposure Floors ​
Author: Ibne Farabi Shihab, Joyanta Jyoti Mondal, Anuj Sharma
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22935v1 Announce Type: new Abstract: Minimum-exposure constraints arise in recommendation, content curation, and regulated allocation when each provider, arm, or group must receive guaranteed exposure inside a period rather than only in aggregate. We study stochastic bandits with exact ex...
36. Verbalized Particle Posterior: Bayesian Inference over Natural Language Hypotheses ​
Author: Yan Zhang, Shikan Lian, Shibo Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.22961v1 Announce Type: new Abstract: Verbalized Machine Learning (VML) parameterizes a model as a natural-language prompt that an LLM evaluates as f(x; theta). The framework is interpretable, but it commits to a single hypothesis with no measure of uncertainty, and that hypothesis varies ...
37. Learned Interventions in Lean 4 grind ​
Author: Evan Wang, Simon Chess, Sophie Szeto, Theodore Meek
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22972v2 Announce Type: new Abstract: Lean 4's grind tactic combines congruence closure, E-matching, and case-splitting into a single automated solver, and like any such solver, it relies on hand-tuned heuristics to decide what to instantiate and where to case-split. These heuristics are t...
38. Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram ​
Author: Ethan Davis
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22980v1 Announce Type: new Abstract: Brain-computer interfaces (BCIs) have long sought calibration-free operation, but classifiers are typically benchmarked by discrimination alone, blind to whether predicted probabilities are well calibrated - a meaningful gap given nonstationary electro...
39. Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes ​
Author: Asha Barua, Sajad Khodadadian
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.OC, stat.ML
arXiv:2607.22982v1 Announce Type: new Abstract: Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success....
40. Label-free Industrial Fault Detection via Adversarial Inverse Reinforcement Learning: A System for Run-to-Failure Prognostics ​
Author: Dhiraj Neupane, Mohamed Reda Bouadjenek, Richard Dazeley, Sunil Aryal
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.22987v1 Announce Type: new Abstract: Machinery fault detection (MFD) remains heavily reliant on supervised learning, which struggles with the scarcity of fault labels in real-world settings. While reinforcement learning (RL) offers a framework to model the sequential nature of degradation...
41. Recycling computational processes of dynamic programming for combinatorial optimization problems: a reservoir computing approach ​
Author: Sora Todaka, Akihiro Yamamoto, Nozomi Akashi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23009v1 Announce Type: new Abstract: Reusing previously computed results is a long-standing principle for reducing computational cost, but such reuse has largely been confined to a single problem's computation. Sharing computational processes across multiple simultaneously solved problems...
42. Mini-batch Noise Lowers Sharpness via Dominant-Subspace Fluctuations ​
Author: Junho So, Dongwook Shin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23012v1 Announce Type: new Abstract: During SGD training, the gradients often align strongly with the dominant subspace spanned by the top-$k$ eigenvectors of the Hessian of the loss. While this seems to naturally imply that loss reduction mainly occurs within this space, prior work has s...
43. All in One: Generative Modeling as Mean-Field Game Design ​
Author: Kun Zhao, Xu Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.23026v1 Announce Type: new Abstract: Mean-field games (MFGs) offer a unifying lens on continuous-time generative modeling: a cost tuple recovering twelve prominent models---Continuous Normalizing Flows, OT-Flow, Score-based Models, Schr"{o}dinger Bridges, and more---as special cases of o...
44. Multi-Agent Privacy Game in Federated Learning: A Unified Mean-Field View ​
Author: Kun Zhao, Xu Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.23029v1 Announce Type: new Abstract: Federated learning enables collaborative model training across distributed clients without centralising their data, yet privacy remains a persistent concern because the shared model updates can leak information about local datasets. Existing privacy-pr...
45. Online Policy Evaluation for MDPs with Dynamic UBSR Measures ​
Author: Weikai Wang, Erick Delage
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.OC
arXiv:2607.23030v1 Announce Type: new Abstract: Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning. Existing approaches either focus on restrictive classes of risk measures or rely on access to a simulator, limiti...
46. MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models ​
Author: Ashitabh Misra, Madhav Agrawal, Arham Jain, Tarek Abdelzaher
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.23047v1 Announce Type: new Abstract: Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but existing methods solve the allocation for a single fixed memory budget. In practice the budget varies across deploy...
47. The Entropic Bound for Transformers: Why Static Rank Fails and Attention-Native Rank Recovers ​
Author: Byeong Hoon Yoon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23050v1 Announce Type: new Abstract: Neural scaling laws describe how loss decreases as models, data, and compute grow, but they do not answer a prior question: for a fixed task, what is the minimum model capacity required to solve it? We study this through the Entropic Bound, a spectral ...
48. Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models ​
Author: Dhruvil S, Fenil Sojitra, Ravirajsinh Chauhan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.23054v1 Announce Type: new Abstract: Multi-head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key-value pairs through a shared low-rank bottleneck (cKV), achieving 81% KV-cache reduction during inference. Despite its adoption in massive production models, no prior work has...
49. Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation ​
Author: Shuai Wang, Daoan Zhang, Zhe Tang, Hao Cheng, Jiaheng Wei
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23125v1 Announce Type: new Abstract: Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks. Current post-training methods usually rely on human-annotated data, distillation from external models, reinforcement learning wit...
50. Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems ​
Author: Tanmay Khandait, Preetom Biswas, Hideki Okamoto, Bardh Hoxha, Georgios Fainekos, Giulia Pedrielli
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY
arXiv:2607.23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation. Existing falsification approaches rely on conditional sampling strategies that factor the joint distribution o...
51. Foundation Models and Fine-Tuning: Toward a New Generation of Models for Time Series Forecasting ​
Author: Morad Laglil, Bertrand Pracca, Emilie Devijver, Eric Gaussier
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23146v1 Announce Type: new Abstract: Inspired by recent breakthroughs in large language models for natural language processing, foundation models have emerged as a promising paradigm for zero-shot time series forecasting, enabling accurate predictions on datasets never seen during pre-tra...
52. XGRVFL-MV: Residual-Coupled Graph-Embedded Multi-View Random Vector Functional Link Network with FleXi Guardian Loss ​
Author: Yogesh Kumar, Mudasir Ganaie
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23149v1 Announce Type: new Abstract: Random Vector Functional Link (RVFL) networks provide an efficient randomized learning framework for classification. Existing multi-view RVFL methods utilize complementary information from multiple views. However, preserving view-specific geometric str...
53. In-Context Learning as Implicit Policy Gradient ​
Author: Masahiro Kaneko, Timothy Baldwin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.23153v1 Announce Type: new Abstract: Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their corresponding evaluation scores as in-context examples. Despite these empirical findings, the theoretical foundat...
54. Wrong Design Intent Is Worse Than None: A Derangement-Control Diagnosis of Header Conditioning in CAD Program Completion ​
Author: Yang Xiao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23191v1 Announce Type: new Abstract: Fine-tuned code LLMs can be conditioned on a lightweight design-intent header to steer parametric CAD generation, but whether the model actually reads the header's content has not been tested under a metric independent of the conditioning itself, nor w...
55. Domain-Prior-Regularized Graph Modeling for Anomaly Detection in Cyber-Physical Systems ​
Author: Youngseok Hwang, Joonsung Kwon, Geonwoo Lee, Hyunwoo Park
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.23197v1 Announce Type: new Abstract: Anomaly detection on multivariate sensor time series is critical for industrial monitoring of cyber-physical systems (CPS), where even subtle deviations from normal behavior can indicate process disruption. Recent graph-based approaches have made signi...
56. Variance-Preserving Orthogonal Selection (VPOS): Greedy Feature Selection via Orthogonal Deflation in PCA Loading Space ​
Author: Baran Koseoglu, Berrin Yanikoglu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23198v2 Announce Type: new Abstract: We propose Variance-Preserving Orthogonal Selection (VPOS), a greedy framework for unsupervised feature selection that operates in the weighted PCA loading space. After each selection, VPOS projects out the chosen feature's variance direction via null-...
57. ParasGB: A Graph Benchmark Suite for Parasitic Estimation on AMS Circuits ​
Author: Jiajun Zou, Jiawei Liu, Ao Liu, Junnong Tian, Yibin Zhang, Chengjie Liu, Yuxi Wang, Shan Shen, Wenhua Gu, Jun Yang, Wenjian Yu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23225v1 Announce Type: new Abstract: As chip manufacturing processes advance to deep submicron nodes, parasitic interconnect effects increasingly dominate the performance of analog and mixed-signal (AMS) circuits and often lead to costly layout iterations. This makes early-stage estimatio...
58. From Score Learning to Discretized Sampling: An End-to-End Generalization Analysis of Diffusion Models ​
Author: Jinshu Huang, Yiming Jiang, Chunlin Wu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.ST, stat.TH
arXiv:2607.23226v1 Announce Type: new Abstract: Despite the empirical success of score-based diffusion models, a complete theoretical understanding of how finite-sample learning, network parameterization, and numerical discretization jointly dictate generative quality remains underdeveloped. Existin...
59. Context-Aware Concept Distillation for Trustworthy Flood Prediction ​
Author: Eli Levinkopf, Efrat Morin, Claudia V. Goldman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.23237v1 Announce Type: new Abstract: Effective flood risk management relies on accurate forecasting, yet the "black box" nature of stateof-the-art Deep Learning models creates a barrier to trust and accountability in high-stakes public safety decisions. While existing Explainable AI (XAI)...
60. StageGuard: Physiologically Constrained Sleep Staging ​
Author: Juntang Wang, Yihan Wang, Hao Wu, Jiayu Gao, Shixin Xu, Dongmian Zou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, q-bio.QM
arXiv:2607.23284v1 Announce Type: new Abstract: Automated sleep staging is increasingly used in large-scale studies to derive sleep-architecture endpoints: total sleep time, REM latency, sleep efficiency, and bout-duration statistics. Deep learning models achieve epoch-level accuracy approaching int...
61. FILLER: Feature Imputation via Latent Location Exploration and Retrieval ​
Author: Santu Mondal, Chayan Maitra, Rajat K. De
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.23295v1 Announce Type: new Abstract: In real-world machine learning applications, incomplete observations create a fundamental challenge. Researchers have come up with several ideas to address this crucial problem. However, current models still face challenges in balancing scalability and...
62. AlloBench: Measuring Online Tool Allocation Capability in LLM Agents ​
Author: Daniel Wang, Andrew Xu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23332v1 Announce Type: new Abstract: Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer an agent that creates a small number of highly reusable tools, rather than many one-offs. We introd...
63. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex ​
Author: Chanwoo Park, Asuman Ozdaglar
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.23333v1 Announce Type: new Abstract: We revisit the regret loss framework introduced in Park et al. (2025), which uses decision-theoretic regret as a direct loss function for training models to make better decisions, through the lens of probability-simplex policies. Our first result shows...
64. Neural operator discovery from heterogeneous trajectories ​
Author: Zituo Chen, Qiaofeng Li, Jiaxin Hu, Sili Deng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23337v1 Announce Type: new Abstract: Neural operators provide data-driven mappings for modeling dynamical systems. Extending them to families of systems typically requires explicit conditioning variables such as physical parameters, geometries, or boundary conditions. In many real-world s...
65. Does Graph Compression Preserve Signal Propagation? ​
Author: Kawshik Banerjee, Khaled Mohammed Saifuddin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23338v1 Announce Type: new Abstract: Graph compression reduces the computational cost of graph learning, but its effect on signal propagation remains largely underexplored. Existing work evaluates compression through downstream task performance or structural preservation, neither of which...
66. SPRKD: Effective Knowledge Distillation for Deep Neural Networks via Saddle Region Approximation ​
Author: Aditya Dewan, Arjun Yogeswaran, Benjamin Fedoruk
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23346v1 Announce Type: new Abstract: Modern deep neural networks are potent catalysts for scientific and industrial impact, yet excessive parameter counts impede deployment in low-compute settings such as hospital equipment and energy infrastructure. Predominant knowledge distillation (KD...
67. Exploration of the generative capabilities of Boltzmann machines applied to social systems under the majority rule ​
Author: Mauricio A. Valle, Gonzalo A. Ruz
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CY
arXiv:2607.23349v1 Announce Type: new Abstract: We study the generative capabilities of Boltzmann machines to recover systems governed by the majority rule under critical conditions. To this end, we train deep belief networks (DBNs) with different configurations, where the first layer can use Gaussi...
68. On the Impossibility of Unbiased and Length-Invariant Policy Optimization with Outcome Rewards ​
Author: Fei Ding, Yongkang Zhang, Yuhao Liao, Zijian Zeng, Huiming Yang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23364v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models, notably adopted by DeepSeek-R1. The recent improvement Dr. GRPO (COLM 2025) identifies the response...
69. Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features ​
Author: Muhammad Abdullah Haroon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CE, econ.EM, stat.CO
arXiv:2607.23370v1 Announce Type: new Abstract: Bitcoin price prediction on sub-daily timescales is a hard open problem in computational finance. Bitcoin exhibits fat-tailed returns, non-stationary dynamics, and a price discovery process influenced by social discourse on Reddit and Twitter. Conventi...
70. Directional Influence Function: Estimating Training Data Influence in Constrained Learning ​
Author: Xin Wang (Jeff), R. Tyrrell Rockafellar (Jeff), Xuegang (Jeff), Ban
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.23388v2 Announce Type: new Abstract: As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints. Understanding how training samples in- fluence ...
71. When Can Depth Replace Precision? A Resource Theory of Quantized Neural Computation ​
Author: Mojtaba Soltanalian
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.OC
arXiv:2607.23390v1 Announce Type: new Abstract: When can additional low-bit residual computation replace missing numerical precision for a fixed input-output map? We model a quantized residual system over a fixed horizon as a pure schedule selecting fields from a declared low-bit operation library, ...
72. A Statistical Difference between Single-Layer Learning and Hierarchical Learning in Wide Neural Networks ​
Author: Sumio Watanabe
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.ST, stat.ML, stat.TH
arXiv:2607.23397v1 Announce Type: new Abstract: Hierarchical neural networks are widely used in artificial intelligence, yet their mathematical properties remain incompletely understood. In the infinite-width limit, two different theoretical frameworks have been proposed. One reduces deep learning t...
73. Transfer Learning Architectures for Scalable Multi-Fidelity Bayesian Optimization ​
Author: Jaewook Lee, Ethan Errington, Christian D. Lorenz, Miao Guo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23404v1 Announce Type: new Abstract: Self-driving laboratories increasingly rely on multi-fidelity Bayesian optimization (MFBO) to balance cheap, approximate evaluations against scarce, expensive ones, with a predictive surrogate at its core. Gaussian processes (GPs) are the default choic...
74. Blood Pressure Estimation from PPG: A Comparative Study of Direct and ECG-Mediated Deep Learning Pipelines ​
Author: Bo Wu, Haoling Wang, Zhuodiao Kuang, Kateryna Shapovalenko
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.ET, stat.AP
arXiv:2607.23406v1 Announce Type: new Abstract: Continuous cuffless blood pressure (BP) monitoring is essential for connected health systems and wearable devices, enabling early detection, longitudinal tracking, and personalized management of cardiovascular disease. Many prior approaches attempt to ...
75. Harmonized Interpretable ECG Waveform Features for Robust Cross-Dataset Clinical Prediction ​
Author: Jie Lin, Weijie Sun, Sunil V. Kalmady, Anita Khalafbeigi, Abram Hindle, Padma Kaul, Russell Greiner
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23412v1 Announce Type: new Abstract: Electrocardiograms (ECGs) are widely used for cardiovascular risk prediction, yet models often fail to transfer across hospitals because of protocol, population, and measurement differences. We benchmark cross-dataset generalization on three tasks - he...
76. Short-Term Pain for Long-Term Gain: Adaptive Experiment with Post-Commitment Reward Shift ​
Author: Puping Jiang, Wei Tang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23432v1 Announce Type: new Abstract: Decision-makers in learning environments face a dilemma when their short-term optimal actions may not favor their long-term benefits the most. To understand the fundamental tradeoff behind the dilemma, we study adaptive experimentation with post-commit...
77. PerturbPFN: Probing the Limits of Synthetic Priors in Drug Perturbation Modelling ​
Author: Yuche Gao, Jos'e Miguel Hern'andez-Lobato, Siyuan Guo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23447v1 Announce Type: new Abstract: Predicting cellular responses to unseen chemical perturbations is challenging due to unknown targets and mechanisms, high-dimensional expression responses, and limited experimental coverage of the large small-molecule design space. We propose PerturbPF...
78. Local Regularization Does Not Characterize Multiclass PAC Learnability ​
Author: Eric Hou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23449v1 Announce Type: new Abstract: Local regularization assigns each hypothesis a test-point-dependent score and predicts with a minimum-score hypothesis consistent with the sample. Asilis et al. asked whether this principle characterizes multiclass PAC learnability. We give a negative ...
79. Generalization bounds and sample complexity for remaining useful life prediction from complete degradation trajectories ​
Author: Huy Hoang Le, Kim-Anh Nguyen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23454v1 Announce Type: new Abstract: Data-driven remaining useful life (RUL) prediction requires complete degradation trajectories for training, yet such run-to-failure data are scarce and expensive. Practitioners currently lack principled guidance on how many failure examples suffice for...
80. Extending Fourier Neural Operators for Modeling Parameterized and Coupled PDEs ​
Author: Cheng Jing, Uvini Balasuriya Mudiyanselage, Abhishek Verma, Kallol Bera, Shahid Rauf, Kookjin Lee
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23466v1 Announce Type: new Abstract: Parameterized and coupled partial differential equations (PDEs) are central to modeling phenomena in science and engineering, yet neural operator methods that address both aspects remain limited. We extend Fourier neural operators (FNOs) with minimal a...
81. Learning to Optimize: Joint Routing and Flow Allocation on Sparse Non-Euclidean Networks ​
Author: Haomiao Sun, Fang He, Congyuan Ji, Xindi Tang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23467v1 Announce Type: new Abstract: We study an integrated pickup-and-delivery problem on sparse, non-Euclidean networks that jointly optimizes cyclic routing, cargo flow allocation, and cross-cycle service. The tight coupling of these operational constraints creates a complex discrete-c...
82. Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning ​
Author: Minh Vu, Konstantinos Slavakis
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23474v1 Announce Type: new Abstract: This paper develops an online, off-policy policy-iteration framework for reinforcement learning (RL), based on sparse Gaussian-mixture-model Q-functions (S-GMM-QFs). The framework reconciles streaming, non-stationary data with the Riemannian structure ...
83. A Multi-stage Constrained Optimization Framework for Data-driven Problems ​
Author: Ye Shi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.OC, stat.ML
arXiv:2607.23480v1 Announce Type: new Abstract: Variational autoencoders (VAEs) transform high-dimensional, often noisy data into a compact latent representation, making downstream optimization more tractable. Three challenges persist in VAE-based constrained optimization: (i) sampling effectively w...
84. Charging Phase Health Indicators for Battery State-of-Health Estimation: A Systematic Comparison of CC, CV, and Combined Approaches under Cross-Battery Validation ​
Author: Huy Hoang Le, Kim-Anh Nguyen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23482v1 Announce Type: new Abstract: Accurate State-of-Health estimation is essential for safe battery operation and cost-effective maintenance. Although numerous health indicators have been derived from constant-current (CC) and constant-voltage (CV) charging phases, their effectiveness ...
85. An adaptive multi-fuzzy logic model for diagnosing transformer faults using dynamic weight optimization ​
Author: Kim-Anh Nguyen, Huy Hoang Le, Ba Tu Phung
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23486v1 Announce Type: new Abstract: Dissolved gas analysis (DGA) is crucial for diagnosing early power transformer failures. Traditional DGA interpretation methods like Duval Triangle, IEC ratio, Roger ratio, Doernenburg ratio and Key Gas are inconsistent and vary in accuracy, especially...
86. Learning Sampling Parameters for Diffusion Models ​
Author: Arisrei Lim, Yossi Gandelsman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2607.23488v1 Announce Type: new Abstract: Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These parameters are typically manually chosen once and then held fixed across pro...
87. Physics-Informed Neural Networks for Discovering Periodic Orbits in the Gravitational Three-Body Problem ​
Author: Nikolaos Kollias, Nikolaos Matzakos
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, nlin.CD, physics.comp-ph
arXiv:2607.23501v1 Announce Type: new Abstract: Locating periodic solutions of chaotic dynamical systems normally requires an initial guess close enough to the target orbit for numerical continuation or gradient-based search to converge. We show that Physics-Informed Neural Networks (PINNs) trained ...
88. Impute On-Demand: Adaptive Correlated Time Series Imputation for Changing Environments ​
Author: Zhichen Lai, Huan Li, Dalin Zhang, Dong Gong, Lina Yao, Christian S. Jensen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DB
arXiv:2607.23503v1 Announce Type: new Abstract: Internet of Things (IoT) applications generate vast amounts of Correlated Time Series (CTS) data that often contain missing values and require imputation. Existing methods emphasize accuracy but often lack adaptability to changing IoT environments: the...
89. Topological Data Analysis and Graph-Theoretic Approaches for Tennis Match Prediction ​
Author: Jake Schwaderer, Alexander Bastien, Omid Khormali, Alejandro Navarrete, Mia Pesavento, Angelika Elderbrook
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.AT, stat.ML
arXiv:2607.23509v1 Announce Type: new Abstract: We present two approaches for predicting tennis match outcomes using topological data analysis and graph theory on ATP singles matches from 2000-2025. The first method applies lower-star filtration to player competitive networks, extracting topological...
90. Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling ​
Author: Hengyuan Cao, Shizhuo Cheng, Mingxuan Liu, Weicheng Huang, Yunhong Lu, Chenxi Cai, Yan Zhang, Min Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, q-bio.BM
arXiv:2607.23518v1 Announce Type: new Abstract: The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in structural biology, by facilitating end-to-end generation via joint sequence-structure modeling or hallucination. However, existing approac...
91. Neonatal Hypoxic-ischaemic Encephalopathy Classification from the EEG and HRV Signals Using a Conformer based Masked Autoencoder ​
Author: Shuwen Yu, William P Marnane, Geraldine B. Boylan, Gordon Lightbody
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.SP
arXiv:2607.23554v1 Announce Type: new Abstract: In this paper, we propose the MAEConformer, a novel self-supervised learning framework that combines the Conformer architecture with the Masked Autoencoder (MAE) paradigm for large-scale representation learning from unlabelled electroencephalography (E...
92. Random Forest-Based Prediction of Bone Volume Fraction and Fracture Position from S-Parameters ​
Author: Jianhe Li, Jinsui Meng, Yida Zhao, Zihe Wang, Liaoran Sun, Tao Shan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23563v1 Announce Type: new Abstract: In this paper, we propose a method for predicting bone volume fraction (BVF) and fracture position by constructing a random forest model based on multichannel S-parameters. A nine-antenna microwave scanning system is designed and fabricated to acquire ...
93. MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model ​
Author: Xin Zhao, Yumin Liu, Zhuo Li, Weichu Zheng, Feng Zhu, Xiaokang Yang, Yaohui Jin, Yanyan Xu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.23607v1 Announce Type: new Abstract: Molecular structure elucidation from tandem mass spectra (MS/MS) is a central inverse problem in analytical chemistry. Most existing approaches to MS/MS identification remain tied to reference libraries or predefined candidate sets, whereas de novo met...
94. Restoration Flow Matching-Based Channel Refinement and Equalization Correction for MIMO Semantic Communications ​
Author: Wenkai Liu, Nan Ma, Jianqiao Chen, Xiaodong Xu, Meixia Tao, Ping Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23615v1 Announce Type: new Abstract: In multiple-input multiple-output (MIMO) semantic communication, imperfect channel state information (CSI) and equalization mismatch can seriously degrade semantic reconstruction quality. To address this issue, we propose a unified restoration flow mat...
95. Optimal Reward Shaping: Autonomous Car Parking Case Study ​
Author: Emre "Ozkaya, Nicolas R. Gauger
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.OC
arXiv:2607.23617v1 Announce Type: new Abstract: Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulting in severe local minima such as policy paralysis or over-conservative hazard avoidance. In this wo...
96. Variational-Ising-Attention (VIA):TailoredAttentionMattersfor Science ​
Author: Rui Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.chem-ph
arXiv:2607.23634v1 Announce Type: new Abstract: Attention enables context modeling via query-key scoring with softmax normalization. Driven by industrial long-context demands, mainstream research has converged toward sparsity and efficiency--yet softmax's independence assumption persists. For scient...
97. CALMRec: Causally Aligned Language Memory for Long-Horizon Recommendation ​
Author: Gengyu Zhan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.23647v1 Announce Type: new Abstract: Large language models (LLMs) can summarize heterogeneous user evidence in natural language, but current LLM recommenders often collapse enduring preferences, transient intent, and exposure-induced behavior into one profile. This makes recommendation vu...
98. DP-IVON-Gradsq: Differentially Private Squared-Gradient Improved Variational Online Newton ​
Author: Nour Jamoussi, Ikram Dridi, Giuseppe Serra, Marios Kountouris
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23649v1 Announce Type: new Abstract: Differential privacy provides formal privacy guarantees for training neural networks on sensitive data, while Bayesian deep learning offers a principled framework for uncertainty-aware prediction. Combining these two objectives remains challenging, as ...
99. Breaking the Total Variance Barrier: Sharp Sample Complexity for Linear Heteroscedastic Bandits with Fixed Action Set ​
Author: Heyang Zhao, Tianyuan Jin, Weixin Wang, Vincent Y. F. Tan, Pan Xu, Quanquan Gu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, stat.ML
arXiv:2607.23679v1 Announce Type: new Abstract: Recent years have witnessed increasing interests in tackling heteroscedastic noise in bandits and reinforcement learning. In these works, the cumulative variance of the noise $\Lambda = \sum_{t=1}^T \sigma_t^2$, where $\sigma_t^2$ is the variance of th...
100. Extreme Volatility Warning under Label Scarcity via Multi-Source Anomaly Fusion ​
Author: Jin Qian, Zhangzhi Xiong, Mingrui Li, Zhen Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23682v1 Announce Type: new Abstract: Early warning of extreme market volatility is central to financial risk management, but actionable events are rare, nonstationary, and often triggered by exogenous information shocks. In our CSI~300 setting, only $\sim$80 positive samples are observed ...
101. The Intruder Threshold: A Spectral Law for LoRA Fine-Tuning ​
Author: Peng Xie
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, stat.ML
arXiv:2607.23711v1 Announce Type: new Abstract: LoRA fine-tuning can create intruder dimensions: new leading singular vectors of the updated weight matrix $W+BA$ that are nearly orthogonal to all pretrained singular vectors and that drive catastrophic forgetting. Since their discovery, no theory has...
102. Outcome-Confounded Local Supervision in On-Policy Distillation ​
Author: Guoqing Ma
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23731v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own trajectories while a teacher supplies dense token-level likelihoods at student-visited prefixes. These likelihoods are often read locally: agreement appears safe to imitate, whereas disagreement ...
103. Source-Free Controlled Adaptation of Teachers for Continual Test-Time Adaptation ​
Author: Anurag Roy, Riddhiman Moulick, Vinay Kumar Verma, Saptarshi Ghosh, Abir Das
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2607.23735v1 Announce Type: new Abstract: In many real-world scenarios, encountering continual shifts in domain during inference is very common. Consequently, continual test-time adaptation (CTTA) techniques leveraging a teacher-student framework have gained prominence, allowing models to adap...
104. Soft-Constrained Optimization of Latent Space in Variational Autoencoders ​
Author: Ye Shi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, stat.ML
arXiv:2607.23751v1 Announce Type: new Abstract: The usefulness of a variational autoencoder (VAE) depends on two properties of its latent space that are hard to obtain together: high encoding capacity in the individual latent variables, and a low-dimensional, disentangled organization of those varia...
105. On the post-hoc Evaluation of PDE Discovery: A Multifaceted Challenge of Scientific Advancement ​
Author: Baptiste Mathevon, Farah Cherfaoui, Amaury Habrard, Marc Sebban
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23753v1 Announce Type: new Abstract: Partial differential equation (PDE) discovery aims to identify from data the governing law of a physical system. Constituting a cornerstone of scientific advancement, it has become during the past decade a major line of research in the rapidly evolving...
106. WISERouter: LLM Routing with Workload Budget Constraint ​
Author: Yifei Li, Zihui Gao, Laks V. S. Lakshmanan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.23765v1 Announce Type: new Abstract: Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is prohibitive at scale. LLM routing exploits diversity in model capability and cost by assigning each query to a suit...
107. A Comparison of Data Augmentation Methods for Training Deep Neural Networks on Synthetic Aperture Sonar ​
Author: C. J. Moore, Gregory D. Vetaw, Jordan Malof
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23770v2 Announce Type: new Abstract: In this work we study Automatic Target Recognition (ATR) for Synthetic Aperture Sonar (SAS) data with a focus on deep neural networks (DNNs). The main challenge in training DNNs for SAS-ATR arises from the limited quantity of labeled target examples du...
108. Scale Weight Decay and Train Better ​
Author: Anuj Apte
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.dis-nn, cs.AI, math.OC
arXiv:2607.23777v1 Announce Type: new Abstract: The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data. This is typically done with a constant decoupled weight decay which causes the network weights to shrink steadily over the course of training. T...
109. SCTA: An Agentic Framework for Stable and Interpretable Target Gene Discovery from Single-Cell RNA Sequencing ​
Author: Shuyu Chen, Chen Zhu, Ye Zhang, Yang Li, Qiqi Xie, Haohan Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, q-bio.GN
arXiv:2607.23821v1 Announce Type: new Abstract: Identifying therapeutic target genes from single-cell RNA sequencing (scRNA-seq) data remains a fundamental challenge in translational biology. Unlike bulk assays, scRNA-seq captures heterogeneous cellular states and rare subpopulations, but this same ...
110. DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification ​
Author: Yuhang Wang, Lingyao Li, Hao Zhou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23822v1 Announce Type: new Abstract: Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so...
111. Latent-LoRA: Compact Latent-Space Adapters with Gradient-Free Routing for Continual Learning ​
Author: Reza Rahimi Azghan, Gautham Krishna Gudur, Giulia Pedrielli, Pavan Turaga, Hassan Ghasemzadeh
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.23837v1 Announce Type: new Abstract: Large language models generalize well to individual tasks but lack an inherent mechanism for learning them sequentially, leading to catastrophic forgetting. To mitigate this, LoRA-based continual learning methods allocate a separate low-rank adapter pe...
112. Covariance Last-Layer Ensembles: Function-Space Diversity for Efficient Uncertainty Quantification ​
Author: H. Martin Gillis, Isaac Xu, Gabriel Spadon, Thomas Trappenberg
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23856v2 Announce Type: new Abstract: A Last-Layer Ensemble (LLE), $K$ linear units on one shared frozen feature map, is an efficient single-pass approach to the disagreement-based epistemic uncertainty for out-of-distribution (OOD) detection. Its weakness is that members share the backbon...
113. Controllable Diversity in Normalization-Based Implicit Ensembles via Softmax-Temperature Modulation ​
Author: Mihai Suteu, Ovidiu Serban
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2607.23860v1 Announce Type: new Abstract: Deep ensembles provide the most reliable uncertainty estimates in deep learning, but their cost grows linearly with the number of members. Implicit ensembles lower this cost by sharing a single backbone across members. Member diversity is a primary det...
114. XMix: Combating Extremely Noisy Labels via Local Smoothness in Self-Supervised Feature Space ​
Author: Chengqi Li, Yangdi Lu, Zhihao Shi, Wenbo He, Chamseddine Talhi, Nadjia Kara
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23865v1 Announce Type: new Abstract: Supervised deep learning models rely on large, accurately labeled datasets, yet noisy annotations are often unavoidable and can severely degrade performance under high noise levels. Recent state-of-the-art methods tackle this by using sample selection ...
115. A Coulomb Particle Model for Learning Kernel Attention in Transformers ​
Author: Masoud Badiei Khuzani, Sharath Honnaiah, Atiq Islam, Alex Cozzi, Abraham Bagherjeiran
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.23869v1 Announce Type: new Abstract: Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution. We propose a particle-based method that learns this distribution by optimizing kernel-target alignmen...
116. Flash-CNNCap: Capacitance Extraction via Image Mapping ​
Author: Hector R. Rodriguez, Jiechen Huang, Wenjian Yu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23877v1 Announce Type: new Abstract: We present Flash-CNNCap, a CNN-based capacitance extractor that reformulates full-matrix capacitance prediction as image-to-image regression over spatial contribution maps. Prior scalar CNN-based extractors require $O(n^2)$ forward passes to recover al...
117. Physics-Informed Neural Networks for Predicting Nitrous Oxide Flux ​
Author: Freddy Yu, Jashanjeet Kaur Dhaliwal, Subhadeep Chakraborty
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML
arXiv:2607.23880v1 Announce Type: new Abstract: Nitrous oxide (N$_2$O) is the dominant ozone-depleting substance emitted in the 21st century, and the third largest contributor to anthropogenic greenhouse gases due to its high potency and long atmospheric lifetime, with more than 70% of N$_2$O emissi...
118. ADVERSARIAL: And-Inverter Graph-Assisted Hardware Trojan Detection At Scale ​
Author: Yaroslav Popryho, Debjit Pal, Inna Partin-Vaisband
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AR, cs.CR
arXiv:2607.23882v1 Announce Type: new Abstract: Modern System-on-Chip (SoCs) often contain hundreds of millions to tens of billions of gates, making existing Hardware Trojan (HT) detection methods impractical due to their immense scale. The proposed approach incorporates symbolically enabled learnin...
119. WorldDiT: A Unified Diffusion Architecture for World and Action Modeling ​
Author: Sen Wang, R. Gnana Praveen, Bidhan Roy, Marcos Villagra
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.RO
arXiv:2607.23909v1 Announce Type: new Abstract: Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world mode...
120. Greedy dynamical meta-learning ​
Author: Aria Yom
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23925v1 Announce Type: new Abstract: Gradient descent scales well to large models, but becomes unstable over long time horizons. Gradient-free optimizers can scale to arbitrary timespans, but are hobbled by high dimensions. Since learning occurs in large models over long timescales, neith...
121. DECAF: De-Clustering for Adaptive Representational Unlearning ​
Author: Anjie Le, Can Peng, Hongcheng Guo, J. Alison Noble
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23934v1 Announce Type: new Abstract: Machine unlearning, which aims to remove the influence of specific training data from a trained model, is a key requirement for privacy, accountability, and adaptive deployment. We argue that many unlearning methods are vulnerable to a simple clusterin...
122. Variational Boosting for Physics-Informed Neural Networks ​
Author: Pavlos Protopapas, Kaylee Vo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23940v1 Announce Type: new Abstract: Physics-Informed Neural Networks (PINNs) solve differential equations by minimizing the residual of a nonlinear operator over a neural parameterization of the solution. However, monolithic PINNs often suffer from ill-conditioning, spectral bias, and op...
123. Joint Flow Matching for Generator-Consistent Classification ​
Author: Hayden McAlister, Lech Szymanski
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.23946v1 Announce Type: new Abstract: We introduce Joint Flow Matching (JFM), a training framework for continuous normalising flows over multiple variables. Standard flow matching transports variables from noise to data simultaneously, offering no natural mechanism for forward and reverse ...
124. Understanding Machine Unlearning Through the Lens of Mode Connectivity ​
Author: Jiali Cheng, Hadi Amiri
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.23970v1 Announce Type: new Abstract: Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress, the loss landscape and optimization geometry of unlearning are poorly understood. In this paper, we study machine...
125. Adaptive Data Admission and Retention for Streaming Federated Learning ​
Author: Zhuoyi Zhao, Ben Liang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NI
arXiv:2607.23987v1 Announce Type: new Abstract: We study streaming federated learning with limited client memory, where newly generated training data incur time-varying sampling costs and must be selectively admitted and retained over time. We consider a joint server-side admission and client-side m...
126. When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost ​
Author: Pin Qian, Su Wang, Chong Peng, Junxian You, Lifei Liu, Haoran Yu, Yihang Chen, Xiaochong Jiang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24010v1 Announce Type: new Abstract: Active RAG systems decide when to retrieve external knowledge during generation, making them a budget-sensitive case of agentic RAG and self-adaptive retrieval. Yet evaluations often leave the operating point underspecified: two systems may both claim ...
127. Capacity-Aware Deep Learning for Generalizable Traffic Volume Estimation Across Links and Cities ​
Author: L'eo Hein, Giovanni De Nunzio, Aur'elie Pirayre, Laurent Najman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.24056v1 Announce Type: new Abstract: Network-wide traffic volume estimation typically relies on propagating measurements from fixed sensors, making performance highly dependent on sensor density and limiting deployment in sparsely instrumented networks. We propose a link-level learning fr...
128. Constrained Reinforcement Learning Using Successor Representations ​
Author: Michael Girstl, Alexander Mattick, Christopher Mutschler
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24057v1 Announce Type: new Abstract: Real-world Reinforcement Learning depends on the ability to formulate safety constraints into a policy. A common way to model such constraints is to introduce an additional cost signal in the Markov Decision Process, which notifies the agent of unwante...
129. ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning ​
Author: Wenwu Fan, Qihong Lin, Zhijie Xia, Zhuo Zheng, Sihao Wang, Qiang Chen, Liangsheng Zhu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.24062v1 Announce Type: new Abstract: Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between training and inference. This training-inference discrepancy stems from two primary factors: an architectural separation ...
130. Learning Reusable Hybrid Motion Priors for Humanoid Locomotion from Motion Imitation ​
Author: Valerio Belli (UNIROMA, UCL), Valerio Modugno (UCL), Enrico Mingo Hoffman (HUCEBOT), Fabio Amadio (HUCEBOT)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.RO
arXiv:2607.24083v1 Announce Type: new Abstract: Reinforcement learning can produce robust humanoid controllers, but each new task is typically trained as a separate policy with its own reward design and training process. Motion imitation provides an alternative source of motor competence by training...
131. MAPLE: Efficient and Diverse Multi-Alpha Generation for Portfolio Construction ​
Author: Yu-Chen Den, Kuan-Yu Chen, Kendro Vincent, Tien-Hao Chang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CE
arXiv:2607.24131v1 Announce Type: new Abstract: Classical alpha mining achieves strong risk-adjusted returns by combining many low-correlated predictive signals, yet deep learning stock-ranking methods typically produce a single alpha per stock, rely on increasingly complex architectures with dimini...
132. An Empirical Study of Feature Selection Granularity ​
Author: Muhammad Rajabinasab, Arthur Zimek
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24145v1 Announce Type: new Abstract: Feature selection aims to identify the most informative and relevant features for a given dataset, either in terms of capturing the underlying data structure and distribution better, or with respect to the performance on a downstream task. Existing res...
133. Calibrated Tree-Neural Fusion for Fine-Grained Vegetation Community Classification ​
Author: Dristi Datta, Md Khalid Hasan Sakib, Manoranjan Paul
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24160v1 Announce Type: new Abstract: Accurate vegetation-community classification is essential for ecological monitoring, habitat assessment, and evidence-based environmental management in heterogeneous landscapes. Existing studies often rely on standalone tree ensembles or generic neural...
134. Forecasting the Emergence and Evolution of Crash Hotspots: A Unified Deep Learning Framework for Proactive Traffic Safety ​
Author: Jingwen Zhu, Keshu Wu, Pei Li, Steven T. Parker, Bin Ran, David A. Noyce
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.ET
arXiv:2607.24168v1 Announce Type: new Abstract: Road crashes remain among the gravest threats to public safety, and preventing them is a defining task of transportation systems worldwide. Much of that harm concentrates at hotspots, yet a hotspot is less a place than an episode; it emerges quietly at...
135. Monitoring Post-Disaster Urban Recovery Using High-Resolution SAR Time Series and Unsupervised Learning: Evidence from the 2023 T\"urkiye-Syria Earthquake ​
Author: Luigi Russo, Deodato Tapete, Silvia Liberata Ullo, Paolo Gamba
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.24180v1 Announce Type: new Abstract: Monitoring post-disaster recovery is essential for understanding how urban systems rebuild and progressively return to functionality. However, tracking reconstruction remains difficult because reliable ground-truth information is often scarce and recov...
136. Every Client Is an Environment: Federated De-confounding for Spatio-Temporal Forecasting ​
Author: Qingxiang Liu, Anqi Liang, Heng Wang, Yuxuan Liang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.24218v1 Announce Type: new Abstract: Federated learning has emerged as a promising paradigm for spatio-temporal forecasting (STF), enabling collaborative model training without sharing raw observations. Existing federated STF methods primarily regard cross-client heterogeneity as an optim...
137. Why does Greedy Search produce Optimal Clustering Outcomes? A Fixed-Core Assignment Theory ​
Author: Kai Ming Ting, Kaifeng Zhang, Sanjay Chawla
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24237v1 Announce Type: new Abstract: Many existing clustering methods are designed based on a set-oriented definition---a cluster is a set of similar points---relying a point-to-point similarity function to find similar points. This works well for compact clusters, but clustering performa...
138. ML-based Predictive Models for Power Consumption in Virtualised O-RANs ​
Author: Rishu Raj, Genevieve Akude, Urooj Tariq, Daniel Kilper
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.24256v1 Announce Type: new Abstract: As communication networks adopt virtualized and disaggregated architectures, achieving energy efficiency has become increasingly important for both economic and environmental reasons. Traditional methods for power modeling are inadequate in these dynam...
139. KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems ​
Author: Shuo Wang, Fang Xi, Wenyuan Huang, Qing Wang, Junming Su
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24260v1 Announce Type: new Abstract: Modern LLM systems increasingly rely on knowledge-selection processes that produce high-value structured priors, such as ranked evidence, graph topology, multimodal alignment, and confidence signals. Yet LLM serving remains fundamentally oblivious to t...
140. Physics-Guided Generative AI for Property-Targeted 3D Porous Media Design ​
Author: Peng Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.24274v1 Announce Type: new Abstract: Inverse design of three-dimensional porous media is central to applications in filtration, catalysis, energy storage, fuel cells, thermal management, and biomedical scaffolds, but remains challenging because many distinct pore geometries can share simi...
141. MEGA-CL: A Molecular Foundation Model for Generalizable ADMET Prediction through Graph External Attention and Contrastive Learning ​
Author: Tinghui Jin, Kedu Jin, Ying Li, Guanghui Ren, Jingzhi Xue, Shiyu Zhou, Xiaoli Dai, Li-bin Wei, Xijing Chen, Di Zhao, Jinfeng Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24314v1 Announce Type: new Abstract: Predicting the absorption, distribution, metabolism, excretion and toxicity (ADMET) properties of small molecules remains a major challenge in drug discovery. Here, we present MEGA-CL, a foundation graph neural network framework for universal molecular...
142. DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation ​
Author: Tan T. Nguyen, Quan V. Dang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24331v1 Announce Type: new Abstract: As the inference phase of Large Language Models (LLMs) requires handling long context windows, the Key-Value (KV) cache initially appears to address this challenge but eventually becomes a significant bottleneck as the context window continues to grow....
143. Unsupervised Graph Representation Learning with Complementary View Alignment ​
Author: Zengyi Wo, Shiyu Zhang, Qiyao Peng, Tianpeng Li, Xuan Guo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24338v1 Announce Type: new Abstract: Unsupervised graph representation learning aims to derive meaningful node embeddings by capturing both structural and attribute information without relying on labeled data. Existing methods, such as GAEs, have demonstrated effectiveness but typically r...
144. Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls ​
Author: Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Khandaker Rifah Tasnia, Sifat Rahman Ahona, Juena Ahmed Noshin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.24343v1 Announce Type: new Abstract: Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body but should not determine a recipient, account, command, or credential. Existing statistical methods typi...
145. Perturbative-NeuSA: A Structured Spectral Framework for Time-Dependent PDEs ​
Author: Xianli Zhu, Jia Yin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.NA
arXiv:2607.24345v1 Announce Type: new Abstract: Neural spectral PDE solvers often learn an entire unresolved vector field even when an inexpensive approximate model can already capture most of the trajectory. Here we introduce Perturbative-NeuSA, a residual formulation that decomposes the target sol...
146. MobiWave: Dispatch-Oriented Graph Wavelets and Drift-Guided Selective Optimization for Autonomous Fleet Rebalancing ​
Author: Xiao Han, Pinbo Wang, Yuanshao Zhu, Guojiang Shen, Xiangjie Kong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24365v1 Announce Type: new Abstract: Autonomous fleets enable mobility platforms to coordinate idle vehicles directly, making fleet-wide rebalancing possible. However, two obstacles limit reliable deployment: overlapping regional and local traffic patterns can hide roads that remain usefu...
147. MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention ​
Author: Jianlin Yu, Jing Lin, Linghui Kong, Aiyue Chen, Weiyi Sun, Chenyu Zeng, Wangli Lan, Jinxi Li, Zhuo Zheng, Ziyang Yue, Danning Ke, Fei Yi, Tianchi Hu, Yuan Ding, Yiwu Yao, Junsong Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2607.24377v1 Announce Type: new Abstract: The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numerical...
148. K-Survival Means ​
Author: Abdallah Alabdallah
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24405v1 Announce Type: new Abstract: In this work, we propose K-SurvMeans, a novel extension of K-Means for clustering survival data. The method explicitly uses the survival outcome in the clustering process to optimize cluster centers, thereby maximizing pairwise survival differences bet...
149. Context Is King: How In-Context Specification Shapes the Geometry of Concepts ​
Author: Elad David, Max Fomin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24425v1 Announce Type: new Abstract: Large language models place structured concepts on geometrically faithful manifolds: weekdays lie on a circle, months on another, usually taken to be a fixed world-model the network stores and looks up. We show that context is king: the structure a mod...
150. DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference ​
Author: Dengke Han
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.24434v1 Announce Type: new Abstract: Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user ...
151. What do Reward Models Memorize? ​
Author: Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutova
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.24484v1 Announce Type: new Abstract: This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dat...
152. UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective ​
Author: Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.24507v1 Announce Type: new Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling. However, adapting AR checkpoints across cor...
153. Physics Transformer: Tailoring Transformer for General PDE Prediction ​
Author: Guoze Sun, Rui Zhang, Jiankai Tang, Mengtao Yan, Runze Mao, Zhi X. Chen, Hao Sun
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24513v1 Announce Type: new Abstract: Transformer architectures have attracted increasing attention for solving partial differential equations (PDEs), owing to their flexibility in handling irregular discretizations and their ability to capture long-range physical dependencies. However, un...
154. Low-Rank Dependence Decomposition via Accelerated Symmetric Non-negative Matrix Factorization ​
Author: Lavinia Ghita, Dhruv Desai, Jake Goldberg, Roman Yokunda Enzmann
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.NA
arXiv:2607.24518v1 Announce Type: new Abstract: Symmetric non-negative matrix factorization (SymNMF) recovers latent group structure from a dependence matrix, but its dense, quadratic-memory objective has confined prior work to moderate sizes. We present a large-scale GPU study of seven algorithm fa...
155. Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls ​
Author: Marzieh Zare
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NE
arXiv:2607.24519v1 Announce Type: new Abstract: Pretrained EEG foundation models are increasingly proposed for clinical decoding, but their transfer across populations and robustness to negative controls remain unclear. We benchmark six models (LaBraM, EEGMamba, CBraMod, REVE, BENDR, and BIOT) on fi...
156. FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models ​
Author: Kaiyang Ye, Yuan Ge, Junxiang Zhang, Bei Li, Ziming Zhu, Haishu Zhao, Xiaoqian Liu, Chenglong Wang, Jingbo Zhu, Zhengtao Yu, Tong Xiao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2607.24522v1 Announce Type: new Abstract: While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored. To this end, we propose Flow Continuous Trajectory Supervision (FlowCTS...
157. From Machine Learning to Large-Scale EO Products: Best Practices for Making Maps ​
Author: Ghjulia Sialelli, Robin Young, Yuchang Jiang, Cesar Aybar, Linus Scheibenreif, Damien Robert, Clemens Mosig, Adam J. Stewart, Jan D. Wegner, Aleksis Pirinen, Olof Mogren, Konrad Schindler
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24532v2 Announce Type: new Abstract: Recent years have seen a rapid expansion in the production of large-scale geospatial maps derived from Earth observation (EO) data, driven largely by advances in machine learning (ML) and large computing infrastructure. Although the barrier to generati...
158. The K-SCAN Clustering Algorithm ​
Author: Filip Kosiorowski, Grzegorz Sroka
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.NE
arXiv:2607.24537v1 Announce Type: new Abstract: In the Big Data era, the scalability of clustering algorithms constitutes a key challenge. Traditional density-based methods (e.g., DBSCAN) offer robustness to noise and the ability to detect non-linear clusters, yet their quadratic time complexity $O(...
159. EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings ​
Author: Xiaocheng Fang, Jieyi Cai, Guangkun Nie, Haoyu Wang, Jiarui Jin, Yujie Xiao, Bo Liu, Chenyang He, Qinghao Zhao, Gaofeng Cheng, Hongyan Li, Shenda Hong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.24553v1 Announce Type: new Abstract: Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardiac findings. Global ECG--text alignment may entangle modality-specific factors, while long-tailed finding distrib...
160. LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding ​
Author: Junsung Hwang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.24555v1 Announce Type: new Abstract: Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read in full at every decode step. Attention keys are locally low-rank though globally high-rank: shared low-rank bases discard page-specific directions...
161. BettiSplit: Topology-Guided Privacy-Aware Split Learning Against Feature Inversion and Gradient Leakage ​
Author: Akarsh K. Nair, Muhammad Arifur Rahman, David Brown, Mufti Mahmud
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR
arXiv:2607.24556v1 Announce Type: new Abstract: Split learning enables collaborative model training by partitioning neural networks across clients and servers. However, improper split placement can lead to severe privacy leakage through intermediate representations. In this work, we propose a topolo...
162. Bit-Accurate FPGA Evaluation of Learned Feature Gating in a Fixed-Point Fourier-Feature Automatic Modulation Classifier ​
Author: Gawthaman Senthilvelan, Luthira Abeykoon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24568v1 Announce Type: new Abstract: Learned feature reweighting can improve automatic modulation classification (AMC) in software, but the same operation introduces additional arithmetic and latency when implemented on an FPGA. This work measures that trade-off in a compact fixed-point c...
163. Evaluating Fuzz Testing for Reinforcement Learning Agents ​
Author: Zhibin Kang, Hanmo You, Dong Wang, Haiming Zheng, Junjie Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.SE
arXiv:2607.24577v1 Announce Type: new Abstract: Reinforcement Learning (RL) agents are increasingly deployed in safety-critical domains such as robotics, autonomous driving, and drone control, where unexpected behaviors may lead to severe real-world consequences. Fuzz testing has recently emerged as...
164. PYPM-GGD: Pitman-Yor Process Mixture with Generalized Gaussian Density using ADAM ​
Author: Kart-Leong Lim
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24583v1 Announce Type: new Abstract: Large scale Bayesian nonparametrics (BNP) learner such as Stochastic Variational Inference (SVI) can handle datasets with large class number and large training size at fractional cost. Like its predecessor, SVI rely on the assumption of conjugate varia...
165. Attribution and Uncertainty Behavior of Learned Residual Gyro Correction for Gyro-Stellar Estimation ​
Author: Mariela De Lucas 'Alvarez, Melvin Laux, Arthur de Freitas Precht, Maurice Martin, Edoardo Caroselli, Frank Kirchner, Alexander Fabisch
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24608v1 Announce Type: new Abstract: This work investigates uncertainty decomposition and explainability in a deep learning-based framework for gyroscope bias correction. A 1-D Convolutional Neural Network is trained to predict residual angular rate corrections from multi-sensor inputs, i...
166. Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects ​
Author: Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty, Iryna Gurevych, Subhabrata Dutta
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.24645v1 Announce Type: new Abstract: The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can va...
167. When Can You Correct Distribution Drift in Temporal Graph Generation? A Sharpening--Drift Tension and an Impossibility for Observation-Based Correction ​
Author: Tianpeng Li, Xuan Guo, Wenjun Wang, Wang Zhang, Pengfei Jiao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.SI
arXiv:2607.24662v1 Announce Type: new Abstract: Generative models of temporal graphs are trained on one stretch of an evolving network and deployed on the next, and they degrade badly in the gap. We show this degradation is derivable, general, and not fixable from observations. The masked flow-match...
168. Explainable Reinforcement Learning via Physics-Aware Policy Distillation ​
Author: Shaker Al-Tamari, Waled Kadour
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.24672v1 Announce Type: new Abstract: In safety-critical sectors such as robotics and automotive engineering, the deployment of Deep Reinforcement Learning (DRL) is often hindered by the black-box nature of deep neural networks. This lack of transparency poses significant challenges for re...
169. Causal-TS: A Python Library for Causal Discovery in High-Dimensional and Nonstationary Time Series ​
Author: Mohammad Fesanghary
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, stat.ME
arXiv:2607.24673v1 Announce Type: new Abstract: We describe Causal-TS, an open-source Python library for causal discovery in high-dimensional and nonstationary multivariate time series. Causal-TS provides four specialized algorithms-CDNOTS, CDNOTS+, CEDAR, and GRACE-along with wrappers for GES, Gran...
170. Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs ​
Author: Justin Sirignano, Konstantinos Spiliopoulos, Samuel Cohen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.NA
arXiv:2607.24726v1 Announce Type: new Abstract: The Deep Galerkin Method (DGM) and Physics Informed Neural Networks (PINNs) have become widely-used methods for solving partial differential equations (PDEs) in the rapidly growing field of scientific machine learning. In these methods, a neural networ...
171. Face Recognition with Machine Learning in OpenCV_ Fusion of the results with the Localization Data of an Acoustic Camera for Speaker Identification ​
Author: Johannes Reschke, Armin Sehr
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:1707.00835v1 Announce Type: cross Abstract: This contribution gives an overview of face recogni-tion algorithms, their implementation and practical uses. First, a training set of different persons' faces has to be collected and used to train a face recognizer. The resulting face model can be u...
172. SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series ​
Author: Giovanni B. Esposito, Francesco Antici, Daniele Cesarini, Andrea Bartolini
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.PF
arXiv:2607.22548v1 Announce Type: cross Abstract: Data centers and their compute nodes require accurate and flexible digital twins capable of modeling the complex interplay of workloads, environmental parameters, and physical metrics. Current machine learning approaches for HPC and its telemetry typ...
173. Learning to Optimize at Scale: A Benders Decomposition-TransfORmers Framework for Stochastic Combinatorial Optimization ​
Author: Seung Jin Choi, Kimiya Jozani, Josh Cooper, Esra Buyuktahtakin Toy
Published: 7/28/2026, 4:00:00 AM
Categories: math.OC, cs.LG
arXiv:2607.22550v1 Announce Type: cross Abstract: We propose a learning-augmented Benders decomposition framework to solve large-scale two-stage stochastic mixed-integer programs. We focus on the two-stage stochastic capacitated lot-sizing problem (TSSCLSP) under demand uncertainty. Our method accel...
174. MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities ​
Author: Nicolas Sibuet, Horacio Saggion, Riccardo Rossi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.LG, cs.SE
arXiv:2607.22552v1 Announce Type: cross Abstract: The automatic translation of mathematical expressions in scientific literature into executable symbolic code (a process we refer to as Formula Formalization) is hindered by a severe scarcity of high-quality, ground-truth datasets specialized for tech...
175. Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy ​
Author: Kazem Faghih, Yize Cheng, Shoumik Saha, Mobina Pournemat, Armin Gerami, Soheil Feizi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2607.22554v1 Announce Type: cross Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers chang...
176. Codifying the Judge: Scalable Evaluation via Program Distillation ​
Author: Tzu-Heng Huang, Shengqi Qiu, Frederic Sala
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.22561v1 Announce Type: cross Abstract: LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alterna...
177. DSTFView: Multi-View Cloud-Edge Workload Forecasting with Dual-Input Spatio-Temporal-Frequency Modeling ​
Author: Qingzhong Li, Hui Ma, Yajun Zhang, Qingchang Ma, Zhou Long
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.22565v1 Announce Type: cross Abstract: With the widespread deployment of edge-side AI inference, edge platforms are increasingly required to support latency-sensitive, highly concurrent, and reliability-critical applications. However, existing methods often struggle to balance multidimens...
178. DRP-FLR: Data-Driven Assessment of Demand Response Potential for Flexible Load Regulation in Smart Grids ​
Author: Yunhao Yao, Siyu Jing, Yang Yang, Qiang Xu, Changqi Weng, Xiang-Yang Li
Published: 7/28/2026, 4:00:00 AM
Categories: math.OC, cs.LG
arXiv:2607.22590v1 Announce Type: cross Abstract: The rapid growth of AI workloads and renewable energy resources exacerbates supply-demand imbalance in power systems, making traditional load regulation designed for efficient allocation inadequate and motivating demand response (DR) mechanisms to en...
179. Quotient Tree Arithmetic: Deferred-Division Computation with Bounded Symbolic Depth and Cross-Subtree Cancellation ​
Author: Gregory Magarshak
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SC, cs.AI, cs.LG, cs.MS
arXiv:2607.22612v1 Announce Type: cross Abstract: We introduce Quotient Tree Arithmetic (QTA), a computational substrate in which values are represented as deferred quotient pairs (N, D) whose ratio is evaluated lazily at a designated materialization boundary. The framework applies to any domain: IE...
180. Masked Autoencoders Learn Perception-Relevant Representations from Resting State Neural Data ​
Author: Aleksandr Kovalev, Antonio Lozano, Fabrizio Grani, Cristina Soto Sanchez, Leili Soo, Roc'io L'opez-Peco, Adrian Villamarin-Ortiz, Roberto Moroll'on Ruiz, Mar'ia del Mar Ayuso Arroyave, Alfonso Rodil, Eduardo Fern'andez
Published: 7/28/2026, 4:00:00 AM
Categories: q-bio.NC, cs.AI, cs.LG, cs.NE
arXiv:2607.22615v1 Announce Type: cross Abstract: Clinical neuroprosthetics face a data bottleneck: labeled perception trials are scarce while hours of spontaneous neural activity are largely underutilized. Here, we test whether self-supervised learning can use these unlabeled datasets to improve pe...
181. A Resolution of the SS--RS--GD Inequalities ​
Author: Binghui Peng
Published: 7/28/2026, 4:00:00 AM
Categories: math.OC, cs.LG, stat.ML
arXiv:2607.22620v1 Announce Type: cross Abstract: Yun, Sra, and Jadbabaie (COLT 2021, open question) conjectured the SS--RS--GD inequalities: for well-conditioned symmetric matrices $A_1,\dots,A_n$, the operators $W_{ss}$, $W_{rs}$, and $W_{gd}$ that encode the expected iterate of single-shuffle SGD...
182. TokenMem: Faithful Knowledge Injection for Frozen LLMs ​
Author: Chengzhang Yu, Chenyang Zheng, Zening Lu, Yingru He, Yutong Huang, Yiming Zhang, Yue Xu, Zhanpeng Jin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.22625v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge, but suffers from knowledge conflicts: when retrieved information contradicts parametric memory, the shared self-attention pathway produces unpredictab...
183. AI-Assisted Causal Inference and Mediation Analyses of Environmental and Psychosocial Determinants of Subjective Cognitive Difficulties in the All of Us Research Program ​
Author: Cong Cao, Shuangge Ma
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG
arXiv:2607.22640v1 Announce Type: cross Abstract: Short-term environmental exposures have been linked to cognitive and behavioral outcomes, although many reported associations may reflect broader geographic and contextual differences. Using longitudinal data from the All of Us Research Program (2018...
184. AutoCluster, AutoTopicModeling, AutoTrendAnalysis: A Complete AutoML Pipeline for Predicting Emerging Trends ​
Author: Ahmed Abolfadl, Marwa Mahmoud Abla, Mervat Abu-Elkheir, Maggie Mashaly
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DL, cs.AI, cs.LG
arXiv:2607.22641v1 Announce Type: cross Abstract: Predicting emerging trends is vital for businesses, researchers, and policymakers; yet traditional approaches often lack scalability and adaptability. This paper presents a trend prediction framework based on Automated Machine Learning (AutoML), desi...
185. CRAFT: Learn the Schema, Execute the Plan ​
Author: Aakash Kolekar, Sahika Genc, Shahriar Shariat, Bunyamin Sisman, Tibor Mezi, Barbara Poblete, Shree Vandana Kachroo, Calvin Chi, Parth Parmar, Ari Singer, Prayaas Jain, Cindy Barker, Benoit Dumoulin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.MA, cs.SE
arXiv:2607.22642v1 Announce Type: cross Abstract: Enterprise coding agents translate natural-language analytical requests into executable code over proprietary APIs, schemas, and metric definitions. Yet the prevailing deployment pattern injecting exhaustive schema and tool documentation into each pr...
186. DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification ​
Author: Mohammed Yousif, Prabhjot Singh, Arjun Pankajakshan, Madhu Reddiboina
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2607.22644v1 Announce Type: cross Abstract: Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or type. This leads to inefficient use of compute and human resources: simple documents are over-process...
187. Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models ​
Author: Yijia Dai, Zhaolin Gao, Yahya Sattar, Jennifer J. Sun, Sarah Dean
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.22646v1 Announce Type: cross Abstract: Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but the algorithm underlying this capability remains undetermined: prior work has proposed several ca...
188. Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance ​
Author: Andrei Starodubov, Yaqub Aris Prabowo, Andreas Hadjipieris, Roberto Galeazzi, Ioannis Kyriakides
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.IT, cs.LG, cs.RO, cs.SY, eess.SP, eess.SY, math.IT
arXiv:2607.22667v1 Announce Type: cross Abstract: This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-vessel tracking in heterogeneous maritime sensor networks. The proposed approach is motivated by information-theoretic sensor management: inst...
189. A Vocabulary for Multi-Agent Automated Research Systems ​
Author: Bardiya Akhbari
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2607.22682v1 Announce Type: cross Abstract: We introduce a vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare. The vocabulary specifies 1) who the agents are, 2) what operations are available in the system, 3) who...
190. Test-Time Coverage: Test-Conditioned Data Curation for Deployment-Aware Learning ​
Author: Nadine Chang, Maying Shen, Shizhe Diao, Jialiang Wang, Jingde Chen, Thomas Breuel, Pavlo Molchanov, Rafid Mahmood, Jose M. Alvarez
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2607.22697v1 Announce Type: cross Abstract: Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution. However, standard data curation methods score training-side criteria rather than directly optimizing deployme...
191. MIME: Multimodal Interactive Motion Encoder ​
Author: Addison Zucek, Prerit Gupta, Kamila Kuatova, Aniket Bera
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.22702v1 Announce Type: cross Abstract: Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI. These settings require representations that align language with both individual actor dynamics and the...
192. Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model ​
Author: Xiao Wang, Hao Si, Qiang Chen, Yu-Xiang Zhang, Beihe Zhang, Jianhua Yang, Qingquan Yang, Dengdi Sun, Wanli Lyu, Guosheng Xu, Jin Tang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.22704v1 Announce Type: cross Abstract: Nuclear fusion has made significant progress in recent years and is expected to become one of the most important pathways to addressing global energy challenges. This paper focuses on observing plasma using visible-light cameras, analyzing its spatio...
193. EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations ​
Author: Anuraag Gadehothur Karnam, Tarunesh Sathish
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.22705v1 Announce Type: cross Abstract: Object-centric learning aims to represent scenes as objects whose properties can be reused in new combinations. Existing evaluations usually score segmentation, single-image factor prediction, or downstream accuracy, but these tests do not directly a...
194. Visual Token Compression Enhances Robustness of MLLMs ​
Author: Shishen Gu, Jiequan Cui, Wenbo Hu, Zenglin Shi, Zhenzhen Hu, Richang Hong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.22716v1 Announce Type: cross Abstract: In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbreak attacks and hallucinations. Given that vision and language modalitie...
195. A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection ​
Author: Ali Borji
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.22722v1 Announce Type: cross Abstract: Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible perturbation that causes the model to keep its original, correct prediction, even though a human would no longer...
196. Structural Preservation Governs Data Augmentation in Deep Learning-Based Laser Speckle Material Classification ​
Author: Mohamed Abdallah Salem, Nourhan Zein Diab
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.22725v1 Announce Type: cross Abstract: Data augmentation is routinely used to improve generalization in image classification, but the assumptions underlying standard policies are poorly matched to coherent imaging. Laser speckle patterns are not generic textures; they arise from coherent ...
197. Trustworthy Medical Segmentation: Uncertainty-Aware U-Net Evaluation Under Clinical Image Degradation ​
Author: Pranav Kaliaperumal, Manisha Kaliaperumal
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.22727v1 Announce Type: cross Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may...
198. Generative Augmentation for EEG Motor Imagery Classification: A Class-Conditional VAE with Cycle-Consistent Decoder Refinement ​
Author: Matei Moldoveanu, Alain Sirois, Claire Ben Ali, Fabien Lotte, Florian Yger
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.22733v1 Announce Type: cross Abstract: We investigate whether a generative model can supply useful synthetic motor-imagery (MI) electroencephalography (EEG) trials that improve the accuracy of independent downstream classifiers. We train a class-conditional variational autoencoder (CVAE) ...
199. Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features ​
Author: Dzmitry Malyshau
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.22739v1 Announce Type: cross Abstract: We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcement learning or explicit memory. Cortex is a compact Quake policy with 10.98 million trainable parameters in a ...
200. Beyond Error-vs-Discard Characteristic: Toward Stable and Reliable Evaluation for Face Image Quality Assessment ​
Author: Bhavesh Wani, \v{Z}iga Babnik, Vitomir \v{S}truc, Philipp Terh"orst
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.22752v1 Announce Type: cross Abstract: Face Image Quality Assessment (FIQA) aims to estimate the utility of facial images for reliable recognition. The evaluation of FIQA methods is predominantly based on the Error-versus-Discard Characteristic (EDC), which evaluates performance by progre...
201. Benchmarking LLMs for Verilog Design Flows ​
Author: Angshuman Chakravertty, Rahul Koshti, Buddhi Prakash Sharma, Vinay Chamola
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AR, cs.LG
arXiv:2607.22759v1 Announce Type: cross Abstract: Large language models (LLMs) show promise in code generation, but their capabilities to produce correct, synthesizable hardware description language (HDL) code still remain to be properly benchmarked. Existing evaluations are primarily relying on pas...
202. DRC-Aid: Design-Rule Correction via Agentic Framework utilizing Inference-Time Large Language Models ​
Author: Anushka Mukherjee, Kang He, Kaushik Roy
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AR, cs.LG
arXiv:2607.22761v1 Announce Type: cross Abstract: Resolving Design Rule Violations (DRVs) in layouts entails an iterative loop of geometric edits and verification. We present DRC-Aid, a closed-loop agentic framework that automates local DRC repair by formulating it as verification-in-the-loop search...
203. TLRNet: Estimating Individual Treatment Effect based on Local Information and Single Learner Structure ​
Author: Ali Haghpanah Jahromi, Mohammad Taheri, Zohreh Azimifar
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.22762v1 Announce Type: cross Abstract: Causal inference has become a central issue across various fields, including computer science, statistics, economics, education, healthcare, and medicine. The broad applicability of this discipline has garnered increased research funding and attentio...
204. Subject-Level Heterogeneity in EEG Motor Imagery Decoding: A Large-Scale Benchmark and Portfolio-Based Reduction of the Search Space ​
Author: Paul Barbaste, Olivier Oullier, Xavier Vasques
Published: 7/28/2026, 4:00:00 AM
Categories: q-bio.NC, cs.LG
arXiv:2607.22778v1 Announce Type: cross Abstract: Robust EEG motor imagery decoding remains limited by strong inter-individual variability, making it difficult to identify pipelines that generalize across users. We present a large-scale, standardized within-session benchmark of decoding pipelines ac...
205. FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon ​
Author: Om Mohite
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AR, cs.DC, cs.LG, cs.PF
arXiv:2607.22785v1 Announce Type: cross Abstract: Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units. Prior attempts, including our own, failed or pro...
206. Practical advantage beyond the quadratic speedup limit with fully-quantum walks ​
Author: Massimiliano Incudini, Guglielmo Mazzola
Published: 7/28/2026, 4:00:00 AM
Categories: quant-ph, cond-mat.dis-nn, cs.LG
arXiv:2607.22818v1 Announce Type: cross Abstract: We introduce a new class of fully-quantum Metropolis walks in which both the proposal and acceptance steps are intrinsically quantum. Unlike standard quantum walks obtained by quantizing classically efficient Markov chains, our algorithm employs Hami...
207. Queryable Self-Organizing Maps: A Database Abstraction for Topology-Driven Data Exploration ​
Author: Denis Mayr Lima Martins, Gottfried Vossen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DB, cs.LG
arXiv:2607.22843v1 Announce Type: cross Abstract: Self-Organizing Maps (SOMs) have long been used as exploratory tools for high-dimensional data: they organize objects into a two-dimensional topology that reveals clusters, gradients, sparse regions, dense regions, and boundaries. Yet, in modern data...
208. AI-interpreted Optical Scattering for Robust and Focal Depth-Aware Imaging ​
Author: Eunji Ko, Patrick Ross, Corey Hart, Wolfgang Losert
Published: 7/28/2026, 4:00:00 AM
Categories: physics.optics, cs.AI, cs.LG
arXiv:2607.22867v1 Announce Type: cross Abstract: Optical scattering has conventionally been regarded as an impediment in imaging research due to the degradation of image quality during reconstruction. Nevertheless, this study explores two cases in which optical scattering may serve a beneficial rol...
209. What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents ​
Author: Shawn Ray
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.LG
arXiv:2607.22868v1 Announce Type: cross Abstract: Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whether intervention changes future behavior. We separate three questions. First, relative to fixed o...
210. Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study) ​
Author: Junda Zhao, Shurui Zhou, Eldan Cohen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2607.22880v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have driven growing interest in using LLMs to automate test generation. Prior work commonly evaluates generated test suites using proxy metrics such as code coverage and mutation score. However, studies...
211. Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests ​
Author: Junda Zhao, Shurui Zhou, Eldan Cohen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG
arXiv:2607.22883v1 Announce Type: cross Abstract: While Large Language Models (LLMs) show great promise for automating unit test generation, recent studies suggest that the quality of generated tests can be negatively impacted when models are prompted with buggy code. This paper presents a new metri...
212. Meshless Domain Randomization via Explicit Parameter Perturbation of 3D Gaussian Splatting ​
Author: Felipe Nunes Carbone de Carvalho, Joyce de Morais Souza, Alan de Aguiar, Charles Morphy D. Santos, Jo~ao Paulo Gois
Published: 7/28/2026, 4:00:00 AM
Categories: cs.GR, cs.CV, cs.LG
arXiv:2607.22890v1 Announce Type: cross Abstract: Domain Randomization (DR) is a standard technique for closing the Sim-to-Real gap, yet traditional DR pipelines rely on classical computer graphics rendering driven by polygon meshes. For complex organic subjects, such as insect specimens, extracting...
213. Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Coding Agent Teams ​
Author: Shouren Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2607.22917v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have significantly improved coding and programming workflows. Claude Code, in particular, is one of the most powerful LLM coding agents and is capable of conducting complex coding tasks. However, several drawbacks ca...
214. Not All LLM Reasoning is Visible in the Chain-of-Thought ​
Author: Vatsal Baherwani, Tom Goldstein, Ashwinee Panda
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.22925v1 Announce Type: cross Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens...
215. Amortized Bayesian Causal Discovery of Extended Factor Graphs ​
Author: Yichen Gu, Yuxuan Song, Weizhou Qian, Yixin Wang, Joshua Welch
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, q-bio.MN
arXiv:2607.22934v1 Announce Type: cross Abstract: Learning causal graphs from interventional data is a challenging problem with broad applications. In molecular biology, for example, a central goal is to uncover gene regulatory networks from large-scale perturbation data. An ideal algorithm for this...
216. Invariant Discovery for Networked Systems ​
Author: Hongyu H`e, Alexander Krentsel, Sylvia Ratnasamy, Maria Apostolaki
Published: 7/28/2026, 4:00:00 AM
Categories: cs.NI, cs.AI, cs.LG, cs.SC
arXiv:2607.22944v1 Announce Type: cross Abstract: Invariants, the relations expected to hold among measured signals of a network, underpin applications from verification to traffic generation, telemetry imputation, and input validation, yet writing them by hand demands rare expertise in both formal ...
217. Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model ​
Author: Robab Aghazadeh Chakherlou, Siddartha Khastgir, Peter Popov, Xingyu Zhao
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2607.22951v1 Announce Type: cross Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified operational profile. Conventional benchmark-based evaluation, often summarized by aggregate accuracy, p...
218. On the Order-Conditional Optimality of Gaffke's Bound ​
Author: George Bissias, Erik Learned-Miller
Published: 7/28/2026, 4:00:00 AM
Categories: math.ST, cs.LG, math.PR, stat.ML, stat.TH
arXiv:2607.22971v2 Announce Type: cross Abstract: Let $X = (X_1, \ldots, X_n)$ be a random vector from any Borel probability law on $\mathbb{R}_+^n$. We revisit the problem of deriving a lower confidence bound (LCB) on a scalar parameter of that law. We recast classical work, beginning with Buehler,...
219. Variable Importance Identification Through Lazy Training for Binary Classification ​
Author: Anand Singh, Luke Pennella, Eshan Kabir, Xiaoxi Shen
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.22979v1 Announce Type: cross Abstract: Deep neural networks have been widely used in many applications (e.g., computer vision and natural language processing); however, understanding their explainability remains a challenging task. Recently, substantial research has been devoted to improv...
220. Metamorphic Testing for Clinical ML Models: A Framework Proposal and Pilot Study ​
Author: Jie JW Wu, Feiyu E, Bo Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.LG
arXiv:2607.22984v1 Announce Type: cross Abstract: Machine learning models for clinical prediction tasks, such as in-hospital mortality and sepsis onset, routinely achieve high AUROC scores. However, AUROC measures ranking performance rather than clinical sensibility. A model may rank patients correc...
221. Robust Conformalized Selection with Noisy Responses ​
Author: Chengyao Yu, Hongxin Wei, Bingyi Jing
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.ME
arXiv:2607.22985v1 Announce Type: cross Abstract: Conformalized selection has been widely applied to select high-quality candidates from large datasets with rigorous uncertainty quantification, such as reliable labeling, drug discovery, and the alignment of large language models. Nevertheless, exist...
222. Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning ​
Author: Tao Zhang, Qixuan Fan, Yiyuan Liang, Yanjie Wang, Song Yan, Tian Tian, Jiahuan Zhou, Luxin Yan, Sheng Zhong, Xu Zou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.22994v1 Announce Type: cross Abstract: Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting. While exemplar replay is effective, it raises concerns regarding privacy and storage. Thus, generative replay has emerged a...
223. Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline ​
Author: Qing Yang, Xun Wang, Ziguan Wang, Zhenjiang Li, Hongqiang Wang, Dongdong Weng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.GR, cs.LG
arXiv:2607.22997v1 Announce Type: cross Abstract: Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next major frontier for AI, echoed by industry leaders such as Jensen Huang (``the next big thing is Phy...
224. WCM: World-Cognition Model for Generalizable Human-Robot Interaction ​
Author: Yuzhen Chen, KC Zhou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.HC, cs.LG
arXiv:2607.22999v1 Announce Type: cross Abstract: Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tasks. Current robot-control paradigms, including vision-language-action policies and world-model-based planners, ...
225. Nesterov acceleration in optimizing over probability measures ​
Author: Jiaqi Tang, Qin Li, Wilfrid Gangbo
Published: 7/28/2026, 4:00:00 AM
Categories: math.OC, cs.LG
arXiv:2607.23008v1 Announce Type: cross Abstract: Optimization over probability measures has become an increasingly important paradigm in modern machine learning, scientific computing, and uncertainty quantification. Motivated by Nesterov's accelerated gradient method in Euclidean space, we develop ...
226. Covariance-Boosted Gaussian Processes for Spatiotemporal Irregularities ​
Author: Jeremy Ovadia
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, physics.space-ph, stat.ME
arXiv:2607.23018v1 Announce Type: cross Abstract: Nonstationary Gaussian process (GP) models are powerful tools for capturing input-dependent variability by adapting to observed data. However, with limited sampling and highly parameterized covariance structure, they are often prone to overfitting an...
227. When Less Is More: A Controlled Benchmark of Lightweight CNNs for Satellite Land-Cover Segmentation on DeepGlobe ​
Author: Atiq Ur Rehman, Joseph Michael Donovan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG, stat.AP
arXiv:2607.23024v1 Announce Type: cross Abstract: High-resolution satellite imagery is the backbone of good land-cover classification, and without that, environmental monitoring, urban planning, and sustainable resource management all fall short. Deep learning architectures perform well in semantic ...
228. Characterizing Arbitrary Lindbladian Dynamics with a Few Pauli Measurements ​
Author: Taiqi Zhou, Weiyuan Gong
Published: 7/28/2026, 4:00:00 AM
Categories: quant-ph, cs.IT, cs.LG, math.IT
arXiv:2607.23044v1 Announce Type: cross Abstract: Quantum devices are open systems whose dynamics interleave coherent evolution with dissipation, and benchmarking, error mitigation, and error correction all rest on a faithful model of both. Existing characterization protocols either assume prior kno...
229. Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models ​
Author: Sultan Alshehri, Zhantao Yang, Han Zhang, Marios Savvides
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.CL, cs.LG
arXiv:2607.23052v1 Announce Type: cross Abstract: Dual-encoder vision-language models (VLMs) expose a similarity interface that enables zero-shot retrieval but fails compositional constraints: queries like "umbrella and no person" retrieve images containing both, even when concept detection is relia...
230. Neural Network-Driven Volatility Drag Mitigation under Aggressive Leverage ​
Author: Christian Bongiorno, Efstratios Manolakis, Rosario Nunzio Mantegna
Published: 7/28/2026, 4:00:00 AM
Categories: q-fin.PM, cs.LG
arXiv:2607.23068v1 Announce Type: cross Abstract: This paper introduces a compact reformulation of a modular end-to-end neural network for global minimum-variance portfolio optimization that decouples model complexity from both look-back window length and universe size. A five-parameter hyperbolic w...
231. Traceable LLM Reasoning for Fake-Order Fraud Detection ​
Author: Siqi You, Bingsong Xu, Zhixian Zheng, Xinjian Peng, Yang Xie, Ying Wang, Jiarong Xu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2607.23075v1 Announce Type: cross Abstract: Detecting fake-order fraud at scale remains a critical challenge for large online-to-offline (O2O) service platforms, as existing approaches often rely on expert-designed features, produce black-box decisions, and provide limited interpretability. To...
232. Operator Neural Jump ODEs: $L^2$-optimal prediction in function spaces ​
Author: Florian Krach, Oliver L"othgren, Josef Teichmann
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.23110v1 Announce Type: cross Abstract: In this paper, we study the extension of Neural Jump ODEs to infinite-dimensional function spaces. In particular, the underlying process $X$ now takes values in $L^2(\Xi, \mathbb{R}^{d_X})$ instead of $\mathbb{R}^{d_X}$ and the Operator NJ-ODE approx...
233. Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs ​
Author: Zhihao Xu, Hao Zhong, Zeting Zhou, Yuhang Xu, Haoyu Tong, Wei Wang, Jinshan Chen, Keqiang He, Chong Zhu, Shengzhong Liu, Fan Wu, Guihai Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DC, cs.LG
arXiv:2607.23115v1 Announce Type: cross Abstract: This paper aims to enable computation- and communication-efficient GPU sharing across devices within local area networks (LANs), facilitating ubiquitous AI inference on heterogeneous personal devices. We achieve distributed task offloading via CUDA A...
234. SMART: LLM-Augmented Hybrid Retrieval for Dynamic Product Ads ​
Author: Congfei Zhang, Jingxiao Ma, Xiaodong Liu, Hsiang-wei Chao, Siman Wang, Ge Liu, Shantanu Aggarwal, Vincent Zhang, Meghana Missula, Rachel Liao, Zichu Li, Xiao Bai, Yunzhi Zhou, Yajun Wang, Zhe Liu, Jinchao Li, Yu Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.CL, cs.LG
arXiv:2607.23121v1 Announce Type: cross Abstract: Dynamic Product Ads (DPA) require retrieving relevant items from multi-million product catalogs, balancing two competing objectives: retargeting (re-surfacing known interests) and prospecting (discovering new categories). While Large Language Models ...
235. Adaptive Multi-Scale Forecasting and Gate-Localized Conformal Prediction for Multivariate Nonstationary Time Series ​
Author: Ziling Ma, Junshu Jiang, 'Angel L'opez-Oriona, Ying Sun, Hernando Ombao
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.AP, stat.ME
arXiv:2607.23165v1 Announce Type: cross Abstract: We propose ABF-T-GLCP, a model-agnostic framework for forecasting and uncertainty quantification in nonstationary multivariate time series. The central idea is to learn an adaptive predictive state representation for point forecasting and reuse it fo...
236. Beyond ICA: Identifiability by Symmetry Breaking ​
Author: Pengzhou Wu
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.ST, stat.TH
arXiv:2607.23182v1 Announce Type: cross Abstract: We prove the identifiability of deep generative models (DGMs) with piecewise-affine (PWA) decoders and Gaussian mixture model (GMM) priors, in a purely unsupervised setting. We introduce three algebraic contrast principles for symmetry breaking: doma...
237. Data-Driven Diffusion Processes on Differential Forms via the Projected Ambient Connection Laplacian ​
Author: Alvaro Almeida Gomez, Jorge Duque Franco
Published: 7/28/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA
arXiv:2607.23192v1 Announce Type: cross Abstract: We develop a data-driven approximation of the projected ambient connection Laplacian acting on differential forms over smooth Riemannian manifolds sampled by point clouds. The proposed construction extends the classical framework of diffusion maps an...
238. FedSLIM: Privacy-Preserving Federated MDL-Based Descriptive Pattern Mining Across Data Silos ​
Author: Samar Samir Khalil, Noha S. Tawfik, Marco Spruit
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG
arXiv:2607.23236v1 Announce Type: cross Abstract: Federated learning has achieved considerable success for predictive modelling, yet federated descriptive analytics remains largely unexplored. Existing federated pattern mining approaches are predominantly support-based and do not optimise a principl...
239. IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages ​
Author: Sahil Deepak Gawande, Mayank Singh
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.LG
arXiv:2607.23242v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language ...
240. FedTaste: Topology-Aware Structural Transfer for Multimodal Federated Learning with Missing Modalities ​
Author: Haochen Liang, Jie Zhang, Hideya Ochiai
Published: 7/28/2026, 4:00:00 AM
Categories: cs.MM, cs.LG
arXiv:2607.23245v1 Announce Type: cross Abstract: Multimodal Federated Learning is often challenged by arbitrary modality missingness and Non-IID data distributions, which lead to severe representation drift and hinder effective collaboration across clients. Existing methods typically rely on genera...
241. CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics ​
Author: Xinhong Xu, Yimeng Zhang, Yuanlong Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.23258v1 Announce Type: cross Abstract: Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central question remains unresolved: \textbf{whether a model pretrained on one collection of recordings can generalize to ne...
242. Photonic reservoir computing with complex networks ​
Author: Sion Park, Kohei Watabe, Satoshi Sunada, Tomoki Yamagami, Atsushi Uchida
Published: 7/28/2026, 4:00:00 AM
Categories: physics.optics, cs.LG, nlin.CD
arXiv:2607.23285v1 Announce Type: cross Abstract: Photonic reservoir computing has attracted increasing attention as a fast and low-cost approach for time-series prediction. Photonic reservoir computing utilizes the high speed, broad bandwidth, and spatial parallelism of light. However, the effect o...
243. TopoFE: topology-aware LLM-guided Automated Feature Engineering ​
Author: Sha Li, Naren Ramakrishnan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.23286v1 Announce Type: cross Abstract: Automatic feature engineering (AutoFE) for tabular learning can be naturally formulated as a program synthesis problem, where the objective is to discover predictive feature transformations from an exponentially large search space. Recent advances in...
244. Learning Asymptotics with Convergence-Rate Guarantees using Linear Least Squares ​
Author: Christos N. Efrem
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, cs.NA, math.CO, math.NA
arXiv:2607.23287v1 Announce Type: cross Abstract: We introduce a new research area that is called Asymptotics Learning Theory (ALT) and combines optimization with asymptotic analysis. In particular, ALT provides a unified approach for computing unknown constants/parameters in proven asymptotic expan...
245. Approximate reservoir computing with a semiconductor laser for reducing energy consumption ​
Author: Tatsuki Ito, Kazutaka Kanno, Satoshi Kawakami, Atsushi Uchida
Published: 7/28/2026, 4:00:00 AM
Categories: physics.optics, cs.LG, nlin.CD
arXiv:2607.23288v1 Announce Type: cross Abstract: Photonic reservoir computing is a promising physical machine-learning technique for predicting time-series data. The quantization of the response signal from the reservoir is required for the implementation of photonic reservoir computing, and the nu...
246. Continuous surrogates versus threshold Boolean networks for modeling Arabidopsis ISR gene regulation ​
Author: Gonzalo A. Ruz
Published: 7/28/2026, 4:00:00 AM
Categories: q-bio.MN, cs.LG, cs.NE
arXiv:2607.23289v1 Announce Type: cross Abstract: Gene regulatory network modeling often requires balancing predictive accuracy and mechanistic interpretability. In this work, we compare continuous surrogate models and a discrete mechanistic model on the same \textit{Arabidopsis thaliana} induced sy...
247. PathRIR: Physics-Guided Acoustic Path Selection and Late-Tail Compensation for Fast Room Impulse Response Simulation ​
Author: Shaoheng Xu, Chunyi Sun, Jihui Zhang, Amy Bastine, Prasanga N. Samarasinghe, Thushara D. Abhayapala
Published: 7/28/2026, 4:00:00 AM
Categories: eess.AS, cs.LG, cs.SD
arXiv:2607.23293v1 Announce Type: cross Abstract: Image-source-method (ISM)-based room impulse response (RIR) simulation is a useful and physically interpretable tool for acoustic scene modeling, but full-order ISM becomes computationally expensive as the reflection order and room complexity increas...
248. Context-Adaptive Inference: A Unified Statistical and Foundation-Model View ​
Author: Yue Yao, Caleb N. Ellington, Jingyun Jia, Baiheng Chen, Dong Liu, Rikhil Rao, Jiaqi Wang, Samuel Wales-McGrath, Yixin Yang, Zhiyuan Li, Eric P. Xing, Ben Lengerich
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.ME
arXiv:2607.23304v1 Announce Type: cross Abstract: Modern predictive systems are expected to adapt their behavior to the specific situation they are facing. A clinical model should not treat every patient the same; a retrieval-augmented model should change its answer when given different evidence; a ...
249. Online Fair Division with Budget Constraints ​
Author: Saar Cohen, Nicholas Teh, Paul W. Goldberg, Michael J. Wooldridge
Published: 7/28/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.LG, cs.MA, econ.TH
arXiv:2607.23310v1 Announce Type: cross Abstract: We study an online variant of discrete fair division under generalized assignment budget constraints. Goods arrive one at a time and must be assigned irrevocably to a feasible agent or to charity, which holds all unallocated goods, while fairness is ...
250. BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis ​
Author: Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.CY, cs.ET, cs.LG
arXiv:2607.23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages. Sanskrit, Tamil, and ot...
251. IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems ​
Author: Shwetha Singaravelu, Gayathri Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.CY, cs.ET, cs.LG
arXiv:2607.23322v1 Announce Type: cross Abstract: Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language general-knowledge tasks and lack coverage of specialized pedagogical doma...
252. BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi ​
Author: Hariom Ingle, Ronit Ghode, Ishwari Gondkar, Jidnyasa Harad, Raviraj Joshi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.LG
arXiv:2607.23344v1 Announce Type: cross Abstract: Named Entity Recognition (NER) for low-resource languages such as Marathi remains a challenging task due to limited annotated resources and linguistic complexity. Although recent Large Language Models (LLMs) have demonstrated strong performance acros...
253. Logit-Coordinate Generative Models for Mixed Continuous-Categorical Tabular Data ​
Author: Yuefei Shen, Xiaotong Shen
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.23348v1 Announce Type: cross Abstract: Mixed continuous--categorical data pose a representation problem for continuous generative models. Flow Matching and Gaussian diffusion operate in Euclidean spaces, whereas categorical laws lie on probability simplices and may be highly imbalanced. W...
254. Hallucination Rates in Language Generation ​
Author: Debmalya Panigrahi, Fan Wei, Ian Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DS, cs.CL, cs.LG
arXiv:2607.23361v1 Announce Type: cross Abstract: Language generation in the limit is an elegant model introduced by Kleinberg and Mullainathan [KM24] to formally study language generation by an algorithm that learns solely based on example strings. In this model, an algorithm is said to correctly g...
255. Investigating the Visual Cues of CNNs for Vascular Segmentation: A Case Study in Microscopy and Fundus Imaging ​
Author: Weslley dos Santos Silva, Cesar Henrique Comin
Published: 7/28/2026, 4:00:00 AM
Categories: eess.IV, cs.CV, cs.LG
arXiv:2607.23371v1 Announce Type: cross Abstract: Vascular segmentation is a standard procedure for clinical diagnosis, yet the specific visual features determining model decisions remain poorly understood. This paper investigates the visual cues Convolutional Neural Networks (CNNs) use to segment b...
256. Rendering on Real Silicon: GPU Render-Timing as a Passive, AI-Resistant CAPTCHA Signal ​
Author: David Noever, Forrest McKee
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.LG
arXiv:2607.23389v1 Announce Type: cross Abstract: Conventional CAPTCHAs pose puzzles that modern AI systems increasingly solve, while behavioral and cryptographic-attestation defenses carry privacy or enrollment costs. We investigate an orthogonal signal: the physical timing behavior of a client's G...
257. Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models ​
Author: Roman Solovyev, Ilya Kiselev, Alexander Stempkovskiy, Tatiana Gabruseva
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SD, cs.LG, eess.AS
arXiv:2607.23395v1 Announce Type: cross Abstract: Music Source Separation (MSS), the task of recovering individual sound components (stems) from a polyphonic mixture, is central to applications ranging from karaoke and remixing to audio restoration and content production. The separation quality depe...
258. Two-Timescale Hierarchical Reinforcement Learning for Resilient Operations ​
Author: Young Hyun Cho, Franz Stoll, Will Wei Sun, Guang Lin, Stephan Biller
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.ME
arXiv:2607.23434v1 Announce Type: cross Abstract: Unexpected shocks recur in global operations, requiring decision rules that adapt as market and operating conditions change. Many operational systems also have hierarchical structures in which long-term and short-term decisions pursue a shared object...
259. Neural Representation of Minimal Surfaces ​
Author: Jiayin Sun, Albert Chern
Published: 7/28/2026, 4:00:00 AM
Categories: cs.GR, cs.LG
arXiv:2607.23437v1 Announce Type: cross Abstract: We propose a neural representation for minimal surfaces. Unlike prior approaches based on discretization or Physics-Informed Neural Networks (PINNs), where meshes or neural fields are optimized to approximate the governing equations, our method build...
260. Constraint-Bound Agnostic Bayesian Optimization: One Model for All Thresholds ​
Author: Jin Wang, Xi Lin, Handing Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.NE, cs.AI, cs.LG
arXiv:2607.23448v1 Announce Type: cross Abstract: Expensive constrained optimization problems in real-world industry design often involve constraint thresholds that are difficult to determine in advance. Engineers may need to adjust constraint thresholds to explore different feasibility-performance ...
261. When Every Simulation Counts: Value-Based Reinforcement Learning for Accelerated Photonics Inverse Design ​
Author: Longying Wen, Feiyang Wu, Jinglin Yu, Chongxian Yuan, Renjie Li, Zhaoyu Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: physics.optics, cs.AI, cs.LG, physics.app-ph
arXiv:2607.23469v1 Announce Type: cross Abstract: Photonic-crystal surface-emitting lasers (PCSELs) can combine high-power operation with narrow-divergence surface emission, but optimizing coupled parameters requires costly full-wave simulations. Deep Q-network (DQN) optimization can reuse simulated...
262. ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour ​
Author: Jianhang Xie, Sicheng Tan, Vishnu Naresh Boddeti, Zhichao Lu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG
arXiv:2607.23478v1 Announce Type: cross Abstract: Fully homomorphic encryption (FHE) provides strong cryptographic guarantees for private inference, but deploying transformer models under FHE remains prohibitively expensive. A key bottleneck is that non-linear operations such as softmax, normalizati...
263. To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion ​
Author: Shaswati Saha, Rajasekhar Anguluri, Manas Gaur
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.23492v1 Announce Type: cross Abstract: Concept erasure techniques (CETs) edit text-to-image diffusion models to erase undesired targets such as NSFW content or copyrighted styles, while preserving model utility on benign concepts. Current CETs face a trade-off between erasure robustness a...
264. Learning switched non-linear dynamical systems from a single trajectory ​
Author: Sunny G. W. Wang, Hemant Tyagi
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.ST, stat.TH
arXiv:2607.23502v1 Announce Type: cross Abstract: We study empirical risk minimization for learning non-linear dynamical systems whose transition dynamics may switch over time. Under stability assumptions, and i.i.d switching over a set of $K$ modes, we derive non-asymptotic bounds on the prediction...
265. An Unofficial FastLAS Tutorial: A Programmer's Guide ​
Author: Fabio Aurelio D'Asaro
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LO, cs.AI, cs.LG
arXiv:2607.23557v1 Announce Type: cross Abstract: FastLAS is a scalable system for Inductive Logic Programming (ILP): you give it some background knowledge, a language bias, and a set of examples, and it searches for a set of logic program rules (a hypothesis) that explains the examples. These notes...
266. Anticipatory Risk-Guided Reinforcement Learning for Safe Flight Through Dynamic Clutter ​
Author: Yuchao Mei, Guohao Zhang, Luxia Ai, Haopeng Chen, Wenbing Tao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.LG
arXiv:2607.23565v1 Announce Type: cross Abstract: Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perception, but more critically on anticipating collision risks induced by relative motion. Conventional modular pipelines frequently suffer f...
267. Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models ​
Author: Liangyu Li, Qingwen Liu, Mingqing Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2607.23602v1 Announce Type: cross Abstract: Controllers based on sampling and latent world models assign a predicted terminal cost to each candidate action sequence, choose the minimum, execute its first action block, and replan. This rule can fail even when the terminal cost perfectly and acc...
268. DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory ​
Author: Xingyang Yu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG, hep-th
arXiv:2607.23614v1 Announce Type: cross Abstract: We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a boun...
269. Order in Desbordante: Techniques for Efficient Implementation of Order Dependency Discovery Algorithms ​
Author: Yakov Kuzin, Dmitriy Shcheka, Michael Polyntsov, Kirill Stupakov, Mikhail Firsov, George Chernishev
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.LG, cs.PF
arXiv:2607.23632v1 Announce Type: cross Abstract: Science-intensive data profiling focuses on discovery and validation of various patterns in datasets. This study considers discovery of one such pattern - order dependency (OD). Simply put, OD states that some list of columns is ordered according to ...
270. Extending Desbordante with Probabilistic Functional Dependency Discovery Support ​
Author: Ilia Barutkin, Maxim Fofanov, Sergey Belokonny, Vladislav Makeev, George Chernishev
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DB, cs.AI, cs.CE, cs.LG
arXiv:2607.23636v1 Announce Type: cross Abstract: Data profiling aims to extract complex patterns from data for further analysis and use that data in domains such as data cleaning, data deduplication, anomaly detection, and many more. Functional dependencies (FDs) are one of the most well-known patt...
271. Distributed Convolutional Rank Regression over Decentralized Networks ​
Author: Chunjing Li, Tiange Zhao, Xiaohui Yuan
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ME, cs.LG
arXiv:2607.23639v2 Announce Type: cross Abstract: This paper studies convolution rank regression (CRR) over decentralized distributed learning networks. We propose a novel decentralized CRR framework, in which estimators are obtained by solving consensus-constrained optimization with kernel-smoothed...
272. When Rates Are Geometric: Rate-Certificate Transfer for Contact Splittings in Optimization ​
Author: George A Kevrekidis
Published: 7/28/2026, 4:00:00 AM
Categories: math.OC, cs.LG, math.DS
arXiv:2607.23642v1 Announce Type: cross Abstract: Discrete optimization algorithms are often analyzed through continuous-time limiting ODEs, but a convergence certificate for the ODE is not automatically one for the discrete algorithm. We develop contact Hamiltonian systems as a setting where the tr...
273. No Free Lunch in Flow Surrogates under Time-Varying Boundary Conditions: A Two-Regime Study ​
Author: Georg Winkler, Martin Stoll
Published: 7/28/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA, physics.flu-dyn
arXiv:2607.23667v1 Announce Type: cross Abstract: A flow surrogate validated on a simple regime is often taken as evidence that the approach will carry to a richer one. We test this assumption on two transient flows under time-varying boundary conditions emulating the process startup: the three-dime...
274. Offline-to-Online Creative Optimization with Generative Models and Adaptive Testing ​
Author: Kevin Lee, Benjamin Letham, Zhiyuan Jerry Lin, Elodie Samson, Eric Onofrey, Poppy Zhang, Shawndra Hill, Eytan Bakshy
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.23696v1 Announce Type: cross Abstract: Ad creative optimization is increasingly constrained by evaluation rather than generation. Generative models can produce many plausible creatives, but reliable evaluation requires online experiments, in which only a limited slate can be tested. We st...
275. The Illusion of Secure LLM Code: Closing the Security Gap via Iterative Reprompting ​
Author: Ishpuneet Singh, Shreyas Mahajan, Gurjot Singh, Maninder Singh
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.HC, cs.LG, cs.MA
arXiv:2607.23710v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly integrated into software development workflows, yet their ability to autonomously generate secure authentication code remains uncertain. This paper evaluates the security architecture of authentication sy...
276. Distributional Split Criteria for Random Forests: Extensions, Shrinkage, and the Robustness of Mean Splitting ​
Author: Silas Koemen
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.ME
arXiv:2607.23721v1 Announce Type: cross Abstract: Distributional random forests replace mean-based CART splitting with criteria that compare the full conditional response distribution in candidate children. We implement and systematically study a family of such criteria inside a single honest-forest...
277. Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning ​
Author: Zahra Abdalla Elashaal, Afef Hfaiedh, Nahla Khraief, Issmail Ellabib, Giansalvo Cirrincione
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.LG
arXiv:2607.23726v1 Announce Type: cross Abstract: Exploration in sparse-reward long-horizon tasks poses significant challenges for reinforcement learning. To address these challenges, we propose a two-level Hierarchical Reinforcement Learning (HRL) framework. The first level handles high-level strat...
278. AI Strategy: How to Choose What AI Product to Implement ​
Author: Foster Provost, Panos Ipeirotis
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG, econ.GN, q-fin.EC, stat.AP
arXiv:2607.23733v1 Announce Type: cross Abstract: Firms struggle to choose AI projects that pay off: two projects can look equally promising to smart, motivated stakeholders and yet deserve opposite decisions. At the residential real-estate brokerage Compass, one AI product (Likely-to-Sell recommend...
279. TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs ​
Author: Muhammad Umar Farooq Qaisar, Lin Zhang, Zhen Chen, Wajdy Othman, Shehzad Ashraf Chaudhry, Chang Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.NI, cs.LG
arXiv:2607.23734v1 Announce Type: cross Abstract: Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environment...
280. GNN-based Multi-Agent Control of Traffic Shockwaves in Sparse Vehicular Ad-hoc Networks ​
Author: Prachi Nandi, Madhuri Malakar, Sonakshi Satpathy, Pabitra Mohan Khilar
Published: 7/28/2026, 4:00:00 AM
Categories: cs.NI, cs.LG, cs.NA, math.NA
arXiv:2607.23792v1 Announce Type: cross Abstract: Traffic shockwaves are stop-and-go waves that propagate upstream through the streams of vehicles and are one of the major causes of traffic congestion, fuel inefficiency, and increased accident rates in modern transportation systems. Although Connect...
281. A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever ​
Author: Sietse Schelpe (Corbenic AI)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR, cs.LG, cs.PF
arXiv:2607.23806v1 Announce Type: cross Abstract: Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a ...
282. TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation ​
Author: Susil Kumar Mohanty, Rohit Patel, Kosuru Yuvaraj, Jeenal Chaudhary, Disha Singhania
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG
arXiv:2607.23838v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base at query time. This makes RAG useful for private data, fast-changing information, and reducing hallucination, ...
283. Long-Tailed Medical Image Classification ​
Author: Nathanael Ren, Saagar Arya
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.23883v1 Announce Type: cross Abstract: In this paper, we examine the difficulties of using standard techniques for medical image classification due to long-tailed distributions (wherein rarer conditions have very few samples) resulting in bias towards diagnosing common diseases and away f...
284. SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception ​
Author: Goodarz Mehr, Sepideh Gohari, Montasir Abbas, Azim Eskandarian
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG, cs.MA, cs.RO
arXiv:2607.23910v1 Announce Type: cross Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicles, such as occlusions and limited sensor range. However, the development of robust V2X algorithms,...
285. SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving ​
Author: Yihui Zhang (Beihang University), Tianyu Wo (Beihang University), Jinghao Wang (Beihang University), Xiaoyang Sun (University of Leeds), Menghao Zhang (Beihang University), Cangzhou Yuan (Beihang University), Li Li (Beihang University), Chunming Hu (Beihang University), Albert Y. Zomaya (The University of Sydney), Renyu Yang (Beihang University)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG, cs.PF
arXiv:2607.23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency. Persistent long-li...
286. Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries ​
Author: Taeyoung Kim
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.23967v1 Announce Type: cross Abstract: Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time relaxation mechanism for grokking in linear models trained with full-batch heavy-ball optimization and weight ...
287. Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and Gender ​
Author: Liu He, Yuanchao Li, Yin-Long Liu, Rui Feng, Yiming Wang, Jiaxin Chen, Yizhe Wang, Jiahong Yuan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SD, cs.LG
arXiv:2607.23977v1 Announce Type: cross Abstract: Acoustic biomarkers show promise for detecting Alzheimer's Disease (AD), yet whether the cues driving diagnostic AI align with those salient to human listeners is underexplored across languages and genders, where pathological markers and perceptual s...
288. HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows ​
Author: Qingyi Yang, Siqian Qiu, Bing Li, Xu Shan, Jia Feng, Shunan Zhou, Xudong Zhou, Tiantian Xing, Jiale Guo, Xiaoyi Dong, Gaoyu Liu, Xiaohuan Liu, Haiqing Pu, Qingwen Deng, Xun Zhang, Zhongrun Xiang, Haiyang Qian, Ying Yan, Yongkang Xu, Nuo Lei, Tianlong Jia, Baoying Shan, Carlo De Michele
Published: 7/28/2026, 4:00:00 AM
Categories: physics.geo-ph, cs.LG
arXiv:2607.23983v1 Announce Type: cross Abstract: Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer. Although artificial intelligence methods have advanced flood prediction and model-error correction, most existing studies have no...
289. Smooth Learning with Hard Constraints via Legendre-Regularized Policies ​
Author: Zikun Lin, Rui Chen, Yijie Wang
Published: 7/28/2026, 4:00:00 AM
Categories: math.OC, cs.LG
arXiv:2607.24007v1 Announce Type: cross Abstract: We revisit contextual optimization from the perspective of policy class design. A desirable policy class should be expressive enough to learn rich context-decision relationships, should enforce hard feasibility constraints rather than soft penalty te...
290. SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation ​
Author: Yu Cui, Yi Xu, Jiahao Wang, Hao Zhang, Yu Zhang, Xiaoyi Zeng, Can Wang, Jinxin Hu, Jiawei Chen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.LG
arXiv:2607.24025v1 Announce Type: cross Abstract: Transformer architectures have achieved remarkable success across diverse domains; however, directly applying their standard self-attention mechanism to recommendation often yields suboptimal performance, sometimes even trailing behind well-designed ...
291. Beyond Local Inspection: Global, Guideline-Grounded Evaluation of Post-hoc XAI Methods for ECG Classification ​
Author: Nils Gumpfer, Michael Guckert, Samuel Sossalla, Birgit A{\ss}mus, Jennifer Hannig
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.LG
arXiv:2607.24035v1 Announce Type: cross Abstract: Explainable AI (XAI) is used to assess whether artificial intelligence models rely on meaningful patterns, yet explanations that appear plausible for individual predictions may systematically misrepresent model behavior. This is particularly problema...
292. The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression ​
Author: Kevin Han Huang, Haoyu Ye, Somak Laha, Morgane Austern
Published: 7/28/2026, 4:00:00 AM
Categories: math.ST, cs.LG, stat.ML, stat.TH
arXiv:2607.24041v1 Announce Type: cross Abstract: Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance matrices are non-degenerate. In this paper, we relax both assumptions...
293. Variational Quantum Conditional Boltzmann Machines for Time-Series Forecasting: Architectures, Symmetric Hyperparameter Evaluation, and a Nonlinear Benchmark ​
Author: Gerhard Hellstern, Danyal Maheshwari, Martin Zaefferer, Martin Braun, Tanja D"ohler
Published: 7/28/2026, 4:00:00 AM
Categories: quant-ph, cs.LG, q-fin.ST
arXiv:2607.24065v1 Announce Type: cross Abstract: In this study, we developed and evaluated four conditional energy-based forecasting architectures: a classical Gaussian-Bernoulli CRBM, a hybrid quantum-classical QCRBM, a full-register QQRBM, and a lag-feature QFeatureQRBM with complete derivations ...
294. When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents ​
Author: Marina Gardella (CB), Camilo Mari{~n}o (UDELAR, CB), Diego Belzarena (UDELAR, CB), Ignacio Ram{'i}rez (UDELAR), Gregory Randall (UDELAR), Jean-Michel Morel (LU - Hong Kong)
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG, eess.IV
arXiv:2607.24077v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to traditional OCR systems, achieving state-of-the-art performance on stand...
295. BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion ​
Author: Minchong Chen, Xiaoyun Yuan, Minyu Cao, Jianing Zhang, Jun Zhang, Shuyang Liu, Xiaokang Yang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG, physics.optics
arXiv:2607.24110v1 Announce Type: cross Abstract: Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sensor misalignment caused by different optics, viewpoints, fields of view, and exposure timings...
296. On Non-Stationary Dynamic Pricing: Adaptivity and Optimality ​
Author: Feiyu Jiang, Zifeng Zhao
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.24115v1 Announce Type: cross Abstract: We study the contextual dynamic pricing problem under non-stationarity, where a firm sells products to $T$ sequentially arriving consumers that behave according to an unknown demand model that can change over time. The demand model is assumed to be a...
297. TEmBed-T: A Multi-Dimensional Benchmark for Table-Level Embeddings ​
Author: Ayeen Poostforoushan, Liane Vogel, Carsten Binnig
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DB, cs.LG
arXiv:2607.24130v1 Announce Type: cross Abstract: Tabular data is the dominant structured-data modality, and learning table representations has become a core research direction. Table-level embeddings in particular underpin a wide range of applications, including table retrieval, data lake discovery...
298. Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness ​
Author: Yang Li, Hai Liu, Dian Shao, Yu Wang, Xiyu Chen, Sergey Volkov, Bozhi Wang, Ziyu Sun, Sihang Liu, Ye Luo, Xiaowei Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.24162v1 Announce Type: cross Abstract: Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete component choices under tight evaluation budgets. Existing approaches - heuristic search, black-box optimizati...
299. Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization ​
Author: Alexey Gavrilov, Alan-Barsag Gazzaev, Sergey Muravyov
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.LG
arXiv:2607.24176v1 Announce Type: cross Abstract: Compressed short-text generators can fail in two different places: the codec may discard information before generation starts, or the latent generator may produce weak codes. Without separating these failure modes, researchers can spend compute impro...
300. EXE-Bench: Ranking the Tradeoffs of AI-based Windows Malware Detectors for Real-World Usability ​
Author: Andrea Ponte, Daniel Gibert, Matous Kozak, Dmitrijs Trizna, Maura Pintor, Battista Biggio, Fabio Roli, Luca Demetrio
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.LG
arXiv:2607.24177v1 Announce Type: cross Abstract: Due to the lack of systematic evaluations, we are not yet able to determine which AI-based Windows malware detector to deploy in production, since existing evaluations (i) differ in terms of data used for both training and testing; (ii) do not consid...
301. LLM-based Source Code Compression via Thresholded Symbol Ranking ​
Author: Angelo Nardone, Paolo Ferragina
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IT, cs.CL, cs.LG, math.IT
arXiv:2607.24192v1 Announce Type: cross Abstract: We study the problem of lossless compression of source code, motivated by the storage demands of large-scale software archives, such as Software Heritage (https://www.softwareheritage.org/). General-purpose compressors (e.g., zstd, bzip2) offer a goo...
302. Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD ​
Author: Jose Cribeiro-Ramallo, Florian Kalinke, Zolt'an Szab'o
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.ST, stat.TH
arXiv:2607.24235v1 Announce Type: cross Abstract: Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others...
303. Decision trees, Frobenius traces, and Weierstrass coefficients of elliptic curves ​
Author: Barinder S. Banwait, Xiaoyu Huang, Kyu-Hwan Lee, Seewoo Lee, Thomas Oliver, Alexey Pozdnyakov
Published: 7/28/2026, 4:00:00 AM
Categories: math.NT, cs.LG
arXiv:2607.24251v1 Announce Type: cross Abstract: We investigate the extent to which the reduced minimal Weierstrass coefficients of an elliptic curve over $\mathbb{Q}$ may be computed from it's Frobenius traces. Decision tree models reveal that the first two reduced minimal Weierstrass coefficients...
304. Catalyst Diffusion Transformer: Generative Inverse Design of Heterogeneous Catalysts ​
Author: Hayoung Doo, Dong Hyeon Mok, Seoin Back, Jonggeol Na
Published: 7/28/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.LG
arXiv:2607.24272v1 Announce Type: cross Abstract: The vast chemical design space and complex, interdependent design variables make catalyst discovery for targeted properties highly labor- and resource-intensive. Although generative models have emerged as a promising solution, existing approaches are...
305. The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages ​
Author: Priyansh Srivastava
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.24276v1 Announce Type: cross Abstract: Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are trained predominantly on English-centric corpora, they introduce a systematic and often overlooked dis...
306. Cross-Attention Calibrated Deduplication for Retrieval-Augmented Generation System ​
Author: Phuong Le Huy, Nam H. Nguyen, Quan V. Dang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.LG
arXiv:2607.24332v1 Announce Type: cross Abstract: Common chunking strategies in Retrieval-Augmented Generation (RAG) systems often create redundant chunks. These redundant chunks make the vector database bigger and slow down retrieval. A common fix is cosine-similarity thresholding. This method redu...
307. When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs ​
Author: Tong Zhang, Zexin Li, Simin Chen, Yun Peng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CR, cs.LG
arXiv:2607.24392v1 Announce Type: cross Abstract: Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performance impact,...
308. Stochastic Counterdiabatic Driving via Biorthogonal Liouvillian Eigenmodes ​
Author: Sandeep Suresh Cranganore, Sebastian Lehner, Johannes Brandstetter, Max Welling
Published: 7/28/2026, 4:00:00 AM
Categories: physics.comp-ph, cond-mat.stat-mech, cs.LG
arXiv:2607.24393v1 Announce Type: cross Abstract: Finite-time driving of stochastic systems generates excess dissipation, causing the evolving probability distribution to lag behind the instantaneous equilibrium, and consequently degrading the convergence of nonequilibrium free energy estimators bas...
309. Multivariate Time Series Forecasting with Adaptive Non-Local Observables ​
Author: Yu-Ting Lee, Huan-Hsin Tseng, Samuel Yen-Chi Chen
Published: 7/28/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG
arXiv:2607.24399v1 Announce Type: cross Abstract: Multivariate time series forecasting (MTSF) predicts future values of multiple variables from historical data. While quantum neural networks have been increasingly applied to this task, they typically rely on fixed local measurements, which restrict ...
310. proxymate: Diagnosis and Adjustment of Proxy Estimates for Reliable Inference ​
Author: Alexandra N. M. Darmon, Deeksha Sinha, Steve Wilkins-Reeves, Caner Gocmen
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.24401v1 Announce Type: cross Abstract: Proxy outcomes (such as short-term behavioral signals, model predictions, or surrogate endpoints) are frequently used in place of primary outcomes that are too slow to mature, rare, or challenging to measure directly. But valid inference on a proxy d...
311. Frequency-Based Reservoir computing ​
Author: Arthur S Powanwe
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.24420v1 Announce Type: cross Abstract: Reservoir computing has emerged as an efficient machine learning framework for predicting time series generated by dynamical systems. In contrast to other machine and deep learning approaches, a reservoir computing trains only the output layer via li...
312. Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation ​
Author: M M Asif Ferdous
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.CL, cs.LG
arXiv:2607.24440v1 Announce Type: cross Abstract: Vision-language models (VLMs) deployed on consumer hardware must decide when to answer and when to defer, and that decision depends on having a confidence signal that tracks correctness. A practitioner with a fixed memory budget faces a choice betwee...
313. ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image ​
Author: Mingzhi Xu, Yizhe Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.24453v1 Announce Type: cross Abstract: Learning from minimal human supervision is a long-standing goal in medical image analysis, where dense expert annotations are costly. We study retinal vessel segmentation in an extreme semi-supervised setting with one annotated image and a pool of un...
314. Self-Attention Dynamics with Rotary Position Embeddings: Twisted States and Explicit Consensus Rates on the Sphere ​
Author: Hao Ye (Xi'an Institute of Optics,Precision Mechanics, Chinese Academy of Sciences, University of Chinese Academy of Sciences)
Published: 7/28/2026, 4:00:00 AM
Categories: math.DS, cs.LG
arXiv:2607.24502v1 Announce Type: cross Abstract: Rotary position embeddings (RoPE) modify attention scores through position-dependent rotations, but their effect on normalized token dynamics is not captured by the vanilla spherical self-attention model. We study the continuous-time dynamics obtaine...
315. The balance between compactness and forecast accuracy of data-driven latent-space reduced-order models in controlled wake flows ​
Author: Alberto Solera-Rico, Patricia Garc'ia-Caspue~nas, Carlos Sanmiguel Vila, Stefano Discetti
Published: 7/28/2026, 4:00:00 AM
Categories: physics.flu-dyn, cs.LG
arXiv:2607.24569v1 Announce Type: cross Abstract: Model-based active flow control requires predictive models that are accurate, stable, and fast enough for real-time optimisation. In controlled wake flows, this is often achieved through Reduced-Order Models (ROMs) that first compress high-dimensiona...
316. A Model for Imbalanced Label Aggregation: A Focus on Minority-Class Detection ​
Author: Gabriel Singer, Samuel Gruffaz, Olivier Vo Van, Nicolas Vayatis, Argyris Kalogeratos
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG
arXiv:2607.24622v1 Announce Type: cross Abstract: We study imbalanced crowdsourcing with a focus on class-dependent annotator accuracy, a setting that, to the best of our knowledge, remains relatively underexplored despite its importance in real-world inspection systems where the labels of greatest ...
317. Efficiency Matters in Autonomous Research ​
Author: Haiqian Yang, Yuan Cao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2607.24647v1 Announce Type: cross Abstract: AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome. In this paper, we argue that the efficiency of...
318. Kimi K3: Open Frontier Intelligence ​
Author: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, Yuxuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Guanduo Chen, Guangyu Chen, Guanzheng Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kexin Chen, Peng Chen, Ruijue Chen, Wentao Chen, Xin Chen, Yang Chen, Yanru Chen, Yifei Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Dazhi Cheng, Yean Cheng, Jialei Cui, Jingbing Cui, Anqi Dai, Jiaqi Deng, Hao Ding, Rui Ding, Shaofeng Ding, Mengfan Dong, Mengnan Dong, Yuhao Dong, Yuxin Dong, Angang Du, Chenzhuang Du, Dikang Du, Jusen Du, Yulun Du, Yu Fan, Jing Feng, Qiulin Feng, Yichen Feng, Kelin Fu, Qiang Fu, Fuxuan Gao, Hongcheng Gao, Jingyue Gao, Tong Gao, Weijia Gao, Shangyi Geng, Jie Gong, Linhu Gong, Shengao Gong, Xiaochen Gong, Qizheng Gu, Yicheng Gu, Shuhao Guan, Haiqing Guo, Shiqi Guo, Xiang Guo, Zhengyan Guo, Beixi Hao, Wenxin Hao, Xiaoru Hao, Dailan He, Haotian He, Lehan He, Qi He, Weiran He, Xinran He, Xinyi He, Yibo He, Yunjia He, Chao Hong, Tiange Hong, Hao Hu, Jiaxi Hu, Ruikun Hu, Weiming Hu, Yangyang Hu, Zhenxing Hu, Liang Hua, Jinbin Huang, Ke Huang, Ruiyuan Huang, Siying Huang, Weixiao Huang, Yan Huang, Zhengjie Huang, Zhiqi Huang, Yulong Hui, Chaobo Jia, Yutong Jiang, Zhejun Jiang, Zuoyou Jiang, Wenyi Jin, Xinyi Jin, Yu Jing, Huanjun Kong, Guokun Lai, Aidi Li, Cheng Li, Chengyuan Li, Cong Li, Fang Li, Guanyu Li, Haoyang Li, Jia Li, Junxiong Li, Lei Li, Letian Li, Lincan Li, Weihong Li, Wentao Li, Xintong Li, Yang Li, Yishen Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zhaoxi Li, Zheming Li, Zhengxiao Li, Zhiyuan Li, Jiawei Lin, Xiaohan Lin, Yibo Lin, Zichao Lin, Ziyan Lin, Bill Liu, Boxiao Liu, Chuan Liu, Liang Liu, Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Weizhou Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yipeng Liu, Zhengying Liu, Zhiheng Liu, Enzhe Lu, Haoyu Lu, Linqiang Lu, Tingzhan Lu, Zhiyuan Lu, Aotian Luo, G. Luo, Junyu Luo, Yifan Luo, B. Lyu, Wenzhou Lyu, Shaoguang Mao, Yuan Mei, Xin Men, Minqing Ni, Yixuan Niu, Siyuan Pan, Shujun Peng, Zhangyang Qi, Ruoyu Qin, ZeChao Qin, Zeyu Qin, Haiquan Qiu, Jianxin Qiu, Jiezhong Qiu, Bowen Qu, Yuhao Qu, Zeyu Shang, Youbo Shao, Han Shen, Jincheng Shi, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Wingchun Siu, Pengwei Song, Xiaoxi Song, Jianlin Su, Yunfeng Su, Zhaochen Su, Lin Sui, Jingsong Sun, Junyao Sun, Shaoning Sun, Shuzhe Sun, Tongyu Sun, Yujun Sun, Yunpeng Tai, Chuning Tang, Heyi Tang, Sirui Tang, Zecheng Tang, Chaoran Tian, Rongpeng Tian, Yu Tian, Wei Tu, Chensi Wang, Chuang Wang, Chunjie Wang, Dinglu Wang, Feng Wang, Hailong Wang, Haiming Wang, Hao Wang, Hao Wang, Huaqing Wang, Hui Wang, Jiayi Wang, Jinglong Wang, Jinhong Wang, Jiuzheng Wang, Linian Wang, Shaobo Wang, Shenzhi Wang, Shuyi Wang, Si Wang, Siyuan Wang, Tianfu Wang, Wenjue Wang, Xingran Wang, Xinmei Wang, Xinyuan Wang, Xusheng Wang, Yalin Wang, Yangkun Wang, Yao Wang, Yaoyu Wang, Yejie Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhenhao Wang, Zhongsheng Wang, Zifan Wang, Chu Wei, Ming Wei, Shouxin Wei, Zichen Wen, Fan Wu, Haoning Wu, Rucong Wu, Wenhao Wu, Xiaoxue Wu, Yingcong Wu, Yongqi Wu, Yuxin Wu, Zijian Wu, Xinglang Xian, Chenxuan Xiang, Yuye Xiang, Bocheng Xiao, Chenjun Xiao, Xin Xiao, Jin Xie, Xiaotong Xie, Yifeng Xie, Zhe Xie, Bowei Xing, Yiming Xiong, Baosheng Xu, Boyu Xu, Jiale Xu, Jianfan Xu, Jing Xu, Jinjing Xu, L. H. Xu, Qingtao Xu, Shuyao Xu, Suting Xu, Tiantian Xu, Tianxiang Xu, Weixin Xu, Xinran Xu, Yangchuan Xu, Ye Xu, Yueni Xu, Ziyao Xu, Haonan Xue, Junjie Yan, Yaoyao Yan, Fan Yang, Guangyao Yang, Hao Yang, Junwei Yang, Ruoyu Yang, Wenjie Yang, Xiaofei Yang, Xinyu Yang, Yi Yang, Yiling Yang, Ying Yang, Yuchen Yang, Zhen Yang, Zhilin Yang, Zian Yang, Zuhao Yang, Haotian Yao, Dan Ye, Haoran Ye, Wenjie Ye, Zhanbo Ye, Bohong Yin, Haoxiang Yin, Xietong Yin, Chengzhen Yu, Haozhen Yu, Longhui Yu, Shengnan Yu, Shuying Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Tongtian Yue, Wei Yue, Yang Yue, Dunyuan Zha, Haobing Zhan, B. H. Zhang, Dehao Zhang, Fei Zhang, Hao Zhang, Haoyuan Zhang, Huanyu Zhang, Jiapei Zhang, Jiaxuan Zhang, Jin Zhang, Kaiyi Zhang, Miaozhen Zhang, Puqi Zhang, Qinglei Zhang, Rong Zhang, Rui Zhang, Shaoshuai Zhang, Shiyi Zhang, Xiaobin Zhang, Xiaoyun Zhang, Y. Zhang, Yangkun Zhang, Ye Zhang, Yichi Zhang, Yikun Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Zijing Zhang, Bin Zhao, Chenguang Zhao, Feifan Zhao, Jinglun Zhao, Jinxiang Zhao, Shuai Zhao, Wenshuo Zhao, Xiangyu Zhao, Xuanle Zhao, Yikai Zhao, Zijia Zhao, Haozhi Zheng, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Haofeng Zhong, Lei Zhong, Longguang Zhong, M. Zhou, Qiankang Zhou, Runjie Zhou, Ruozhang Zhou, Xinyu Zhou, Yiqiao Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yangjunfeng Zhu, Yuxuan Zhu, Zhen Zhu, Chen Zhuang, Weiyu Zhuang, Xinxing Zu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.LG
arXiv:2607.24653v1 Announce Type: cross Abstract: We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve i...
319. MMOE: Modernizing Diffusion Transformers with Efficient Expert Design ​
Author: Yanhao Jia, Jiepeng Wang, Haibin Huang, Chi Zhang, Erik Cambria, Xuelong Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.GR, cs.LG
arXiv:2607.24665v1 Announce Type: cross Abstract: Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation Models (AFMs), especially diffusion-transformer backbones, have begun ...
320. Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification ​
Author: Francisco Mena, Dino Ienco, Roberto Interdonato, Cassio F. Dantas, Simon Besnard
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.24683v1 Announce Type: cross Abstract: Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real-world scenarios subject to operational constraints, such as sensor failures or privacy restrictions, lead to i...
321. Stacking the Deck: Tunable Trainability in Stacked LCUs ​
Author: Nikhil Khatri, Stefan Zohren, Gabriel Matos
Published: 7/28/2026, 4:00:00 AM
Categories: quant-ph, cs.LG
arXiv:2607.24686v1 Announce Type: cross Abstract: Variational quantum circuits have been central to many proposed near-term applications of quantum computing, but a growing body of evidence suggests that trainability and quantum advantage are fundamentally at odds: ans"atze expressive enough to res...
322. Beyond Scale and Generation: Understanding Language Model-based Entity Matching ​
Author: Zeyu Zhang, Xue Li, Iacer Calixto, Paul Groth, Sebastian Schelter
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DB, cs.CL, cs.LG
arXiv:2607.24688v1 Announce Type: cross Abstract: Entity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher architectures. However, prior studies often conflate matcher architectu...
323. The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation ​
Author: Tianyi Men, Zhuoran Jin, Kang Liu, Jun Zhao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.24720v1 Announce Type: cross Abstract: Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning abilit...
324. Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation ​
Author: Bingnan Li, Haozhe Wang, Haozhong Xiong, Fangtai Wu, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2607.24731v1 Announce Type: cross Abstract: On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains po...
325. Learning Distributions from Multiple Data Providers ​
Author: Jon Kleinberg, Amin Saberi, Xizhi Tan, Grigoris Velegkas
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DS, cs.GT, cs.LG, stat.ML
arXiv:2607.24732v1 Announce Type: cross Abstract: Motivated by learning from heterogeneous and overlapping data providers, we study a stylized model of distribution learning from restricted conditional samples. The goal is to learn an unknown distribution $p$ on a finite domain $[n]$. The learner is...
326. Certified Parallel-in-Time Sinkhorn for Dynamic Entropic Optimal Transport ​
Author: Xinyang Wen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DC, cs.LG
arXiv:2607.24741v1 Announce Type: cross Abstract: Dynamic applications, including optimal-transport Flow Matching, repeatedly solve related entropic optimal transport problems, yet conventional distributed Sinkhorn processes frames sequentially and synchronizes after every iteration. We present Temp...
327. Hierarchical Reinforcement Learning with Optimal Level Synchronization Based on Flow-Based Deep Generative Model ​
Author: JaeYoon Kim, Junyu Xuan, Christy Liang, Farookh Hussain
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2107.08183v2 Announce Type: replace Abstract: High-dimensional state and action spaces combined with sparse reward structures in reinforcement learning (RL) environments typically require advanced control architectures. Hierarchical Reinforcement Learning (HRL) demonstrates superior performanc...
328. On a linear fused Gromov-Wasserstein distance for graph structured data ​
Author: Dai Hai Nguyen, Koji Tsuda
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2203.04711v2 Announce Type: replace Abstract: We present a framework for embedding graph structured data into a vector space, taking into account node features and topology of a graph into the optimal transport (OT) problem. Then we propose a novel distance between two graphs, named linearFGW,...
329. Fairness Interventions in Classification: A Study on AI Explainability ​
Author: Thomas Souverain, Paul 'Egr'e
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CY
arXiv:2407.14766v4 Announce Type: replace Abstract: This paper presents a philosophical and experimental study of fairness interventions in AI classification, centered on the explainability and transparency of corrective methods, and on the opposition between two fairness criteria, namely Demographi...
330. Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks ​
Author: Rui Hu, Yifan Zhang, Zhuoran Li, Longbo Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2410.02596v2 Announce Type: replace Abstract: Generative Flow Networks (GFlowNets) are a novel class of generative models designed to sample from unnormalized distributions and have found applications in various important tasks, attracting great research interest in their training algorithms. ...
331. CausAdv: A Causal-based Framework for Detecting Adversarial Examples ​
Author: Hichem Debbi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, stat.ME, stat.ML
arXiv:2411.00839v4 Announce Type: replace Abstract: Deep learning has led to tremendous success in computer vision, largely due to Convolutional Neural Networks (CNNs). However, CNNs have been shown to be vulnerable to crafted adversarial perturbations. This vulnerability of adversarial examples has...
332. Fisher Information based Stochastic Gradient Ascent for Online Learning of Dirichlet Process Mixture and Theory ​
Author: Kart-Leong Lim, Xudong Jiang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, stat.ML
arXiv:2412.08951v3 Announce Type: replace Abstract: Scalable algorithms of posterior approximation allow Bayesian nonparametrics such as Dirichlet process mixture to scale up to larger dataset at fractional cost. Recent algorithms, notably the stochastic variational inference performs local learning...
333. Nearly Tight Bounds for Cross-Learning Contextual Bandits with Graphical Feedback ​
Author: Ruiyuan Huang, Zengfeng Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2502.04678v2 Announce Type: replace Abstract: Repeated first-price auctions are contextual decision problems with censored but reusable feedback: after submitting a bid, a learner can infer the outcomes of related bids and evaluate them under different private values. This structure motivates ...
334. Sign-Symmetry Learning Rules are Robust Fine-Tuners ​
Author: Aymene Berriche, Mehdi Zakaria Adjal, Riyadh Baghdadi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2502.05925v2 Announce Type: replace Abstract: Backpropagation (BP) has long been the predominant method for training neural networks due to its effectiveness. However, numerous alternative approaches, broadly categorized under feedback alignment, have been proposed, many of which are motivated...
335. A Survey of Graph Transformers: Architectures, Theories and Applications ​
Author: Chaohao Yuan, Kangfei Zhao, Ercan Engin Kuruoglu, Liang Wang, Tingyang Xu, Wenbing Huang, Deli Zhao, Hong Cheng, Yu Rong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2502.16533v3 Announce Type: replace Abstract: Graph Transformers (GTs) have demonstrated a strong capability in modeling graph structures by addressing the intrinsic limitations of graph neural networks (GNNs), such as over-smoothing and over-squashing. Recent studies have proposed diverse arc...
336. Sampling Decisions: Exact Path-Space Correction, Prior Cancellation and Local-Boltzmann Guidance ​
Author: Michael Chertkov, Sungsoo Ahn, Hamidreza Behjoo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.stat-mech, cs.AI, cs.SY, eess.SY, stat.ML
arXiv:2503.14549v3 Announce Type: replace Abstract: How can a cheap but biased sequential, finite-horizon sampler over a discrete space be corrected so that its terminal output follows a prescribed Gibbs distribution? We formulate Sampling Decisions as a path-space relative-entropy projection on a g...
337. Gradient-Free Continual Learning ​
Author: Grzegorz Rype's'c
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2504.01219v2 Announce Type: replace Abstract: Neural networks are notorious for forgetting old skills when taught new ones - a problem known as catastrophic forgetting. Standard continual learning techniques try to fix this by saving old data or relying on complex gradient updates, but these m...
338. Analyzing the Importance of Blank for CTC-Based Knowledge Distillation ​
Author: Benedikt Hilmes, Nick Rossenbach, Ralf Schl"uter
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.SD, eess.AS
arXiv:2506.01503v2 Announce Type: replace Abstract: With the rise of large pre-trained foundation models for automatic speech recognition new challenges appear. While the performance of these models is good, runtime and cost of inference increases. One approach to make use of their strength while re...
339. Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions ​
Author: Haotian Jiang, Zeyu Bao, Shida Wang, Qianxiao Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2506.05678v3 Announce Type: replace Abstract: The evolution of sequence modeling architectures, from recurrent neural networks and convolutional models to Transformers and structured state-space models, reflects ongoing efforts to address the diverse temporal dependencies inherent in sequentia...
340. Exact Evaluation of the Accuracy of Diffusion Models for Inverse Problems with Gaussian Data Distributions ​
Author: Emile Pierret, Bruno Galerne
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2507.07008v2 Announce Type: replace Abstract: Used as priors for Bayesian inverse problems, diffusion models have recently attracted considerable attention in the literature. Their flexibility and high variance enable them to generate multiple solutions for a given task, such as inpainting, su...
341. Reconstruction of SINR Maps from Sparse Measurements using Group Equivariant Non-Expansive Operators ​
Author: Lorenzo Mario Amorosa, Francesco Conti, Nicola Quercioli, Flavio Zabini, Tayebeh Lotfi Mahyari, Yiqun Ge, Patrizio Frosini
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.NI
arXiv:2507.19349v3 Announce Type: replace Abstract: As sixth generation (6G) wireless networks evolve, accurate signal-to-interference-noise ratio (SINR) maps are becoming increasingly critical for effective resource management and optimization. However, acquiring such maps at high resolution is oft...
342. Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers ​
Author: Xingyue Huang, Rishabh, Gregor Franke, Ziyi Yang, Jiamu Bai, Weijie Bai, Jinhe Bi, Zifeng Ding, Yiqun Duan, Chengyu Fan, Wendong Fan, Xin Gao, Ruohao Guo, Yuan He, Zhuangzhuang He, Xianglong Hu, Neil Johnson, Bowen Li, Fangru Lin, Siyu Lin, Tong Liu, Yunpu Ma, Hao Shen, Hao Sun, Beibei Wang, Fangyijie Wang, Hao Wang, Haoran Wang, Yang Wang, Yifeng Wang, Zhaowei Wang, Ziyang Wang, Yifan Wu, Zikai Xiao, Chengxing Xie, Fan Yang, Junxiao Yang, Qianshuo Ye, Ziyu Ye, Guangtao Zeng, Yuwen Ebony Zhang, Zeyu Zhang, Zihao Zhu, Bernard Ghanem, Philip Torr, Guohao Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2509.03059v2 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can be significantly improved through Reinforcement Learning with Verifiable Reward (RLVR), particularly in domains like mathematics and programming, where...
343. Neural Message-Passing on Attention Graphs for Hallucination Detection ​
Author: Fabrizio Frasca, Guy Bar-Shalom, Yftah Ziser, Haggai Maron
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2509.24770v2 Announce Type: replace Abstract: Large Language Models (LLMs) often generate incorrect or unsupported content, known as hallucinations. Existing detection methods rely on heuristics or simple models over isolated computational traces such as activations, or attention maps. We unif...
344. Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling ​
Author: Alexandru Meterez, Depen Morwani, Jingfeng Wu, Costin-Andrei Oncescu, Cengiz Pehlevan, Sham Kakade
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC, stat.ML
arXiv:2510.14717v2 Announce Type: replace Abstract: Increasing the batch size during training -- a ''batch ramp'' -- is a promising strategy to accelerate large language model pretraining. While for SGD, doubling the batch size can be equivalent to halving the learning rate, the optimal strategy for...
345. Continual Knowledge Consolidation LORA for Domain Incremental Learning ​
Author: Naeem Paeedeh, Mahardhika Pratama, Weiping Ding, Jimmy Cao, Wolfgang Mayer, Ryszard Kowalczyk, Ary Shiddiqi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2510.16077v2 Announce Type: replace Abstract: Domain Incremental Learning (DIL) is a sub-branch of continual learning that aims to address the never-ending arrival of new domains without catastrophic forgetting. Despite the advent of parameter-efficient fine-tuning (PEFT) approaches, prior wor...
346. On-Device Inference versus Wireless Streaming: Energy-Efficient Multi-Modal Deep Learning for Wearable Cardiovascular Patches ​
Author: Mustafa Fuad Rifet Ibrahim, Tunc Alkanat, Felix Manthey, Maurice Meijer, Alexander Schlaefer, Peer Stelldinger
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2510.18668v4 Announce Type: replace Abstract: Wearable cardiovascular sensor patches promise continuous, unobtrusive monitoring, but their tight energy, memory, and compute budgets make it unclear whether physiological signals should be analyzed on the device or streamed to the cloud for proce...
347. Forgetting is Everywhere ​
Author: Ben Sanati, Thomas L. Lee, Trevor McInroe, Aidan Scannell, Esmeralda S. Whitammer, David Abel, Amos Storkey
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, stat.ML
arXiv:2511.04666v4 Announce Type: replace Abstract: A fundamental challenge in developing general learning algorithms is their tendency to forget past knowledge as they adapt to new data. Addressing this problem requires a principled understanding of forgetting. Yet, despite decades of study, no uni...
348. Uncertainty Modeling for Multi-Objective RTA Interception with Distillation Acceleration ​
Author: Gaoxiang Zhao, Ruinan Qiu, Xiaoting Wang, Pengpeng Zhao, Rongjin Wang, Xiaoting Wang, Zhangang Lin, Xiaoqiang Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.GT
arXiv:2511.05582v3 Announce Type: replace Abstract: Real-Time Auction (RTA) interception decides which incoming advertising requests reach downstream systems, and therefore controls the quality of the data those systems learn from. At JD.com, off-site advertising produces on the order of hundreds of...
349. Copula Based Fusion of Clinical and Genomic Machine Learning Risk Scores for Breast Cancer Risk Stratification ​
Author: Agnideep Aich, Sameera Hewage, Md Monzur Murshed
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, stat.ME, stat.ML
arXiv:2511.17605v2 Announce Type: replace Abstract: Clinical and gene-expression models predict breast cancer outcomes, but simple linear fusion ignores dependence between their risk scores. Using METABRIC, we tested whether modeling the joint distribution of clinical and gene-expression scores impr...
350. Quantum Safe-Set Bayesian Optimization for Quality Improvement in Fuselage Assembly ​
Author: Jiayu Liu, Chong Liu, Trevor Rhone, Yinan Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2511.22090v2 Announce Type: replace Abstract: Recent efforts in smart manufacturing have enhanced aerospace fuselage assembly processes, particularly by innovating shape adjustment techniques to minimize dimensional gaps between assembled sections. Existing approaches have shown promising resu...
351. A Simple, Optimal and Efficient Algorithm for Online Exp-Concave Optimization ​
Author: Yi-Han Wang, Peng Zhao, Zhi-Hua Zhou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.OC, stat.ML
arXiv:2512.23190v3 Announce Type: replace Abstract: Online eXp-concave Optimization (OXO) is a fundamental problem in online learning, where the goal is to minimize regret when loss functions are exponentially concave. The standard algorithm, Online Newton Step (ONS), guarantees an optimal $O(d \log...
352. AutoFed: Personalized Federated Traffic Prediction via Adaptive Prompt ​
Author: Zijian Zhao, Yitong Shang, Sen Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2512.24625v3 Announce Type: replace Abstract: Accurate traffic prediction is essential for Intelligent Transportation Systems, including ride-hailing, urban road planning, and vehicle fleet management. However, due to significant privacy concerns surrounding traffic data, most existing methods...
353. Infinite-Precision Autoregressive Modeling for Vector Graphics and Layouts ​
Author: Yeonsang Shin, Insoo Kim, Bongkeun Kim, Keonwoo Bae, Bohyung Han
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV
arXiv:2601.05680v2 Announce Type: replace Abstract: While Transformer-based autoregressive models excel in data generation, their token discretization strategy inherently limits their precision in continuous domains. We analyze the scalability limitations of existing discretization-based approaches ...
354. A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning ​
Author: Jiali Cheng, Ziheng Chen, Chirag Agarwal, Hadi Amiri
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, cs.CV
arXiv:2601.09624v2 Announce Type: replace Abstract: Machine unlearning is becoming essential for building trustworthy and compliant language models. Yet unlearning success varies considerably across individual samples: some are reliably erased, while others persist despite the same procedure. We arg...
355. Ordering-based Causal Discovery via Generalized Score Matching ​
Author: Vy Vo, He Zhao, Trung Le, Edwin V. Bonilla, Dinh Phung
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.16249v3 Announce Type: replace Abstract: Learning DAG structures from purely observational data remains a long-standing challenge across scientific domains. An emerging line of research leverages the score of the data distribution to initially identify a topological order of the underlyin...
356. Physics-Encoded Inverse Modeling for Arctic Snow Depth Estimation ​
Author: Akila Sampath, Vandana P. Janeja, Jianwu Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.17074v5 Announce Type: replace Abstract: Accurate estimation of unobserved quantities in time-varying inverse problems remains challenging when observations are sparse and only indirectly related to the target variable. In Arctic climate applications, snow depth over sea ice is not direct...
357. Action-Sufficient Goal Representations ​
Author: Jinu Hyeon, Woobin Park, Hongjoon Ahn, Taesup Moon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2601.22496v2 Announce Type: replace Abstract: In offline goal-conditioned reinforcement learning (GCRL), hierarchical approaches decompose long-horizon tasks into high-level subgoal prediction and low-level action execution. A critical design choice in such architectures is the goal representa...
358. Plain Transformers are Surprisingly Powerful Link Predictors ​
Author: Quang Truong, Yu Song, Donald Loveland, Mingxuan Ju, Tong Zhao, Neil Shah, Jiliang Tang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.01553v3 Announce Type: replace Abstract: Link prediction is a core challenge in graph machine learning, demanding models that capture rich and complex topological dependencies. While Graph Neural Networks (GNNs) are the standard solution, state-of-the-art pipelines often rely on explicit ...
359. Asymmetric Hierarchical Anchoring for Robust Audio-Visual Cross-Modal Generalization ​
Author: Bixing Wu, Yuhong Zhao, Zongli Ye, Jiachen Lian, Xiangyu Yue, Gopala Anumanchipalli
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2602.03570v2 Announce Type: replace Abstract: Audio-visual joint representation learning under Cross-Modal Generalization (CMG) aims to transfer knowledge from a labeled source modality to an unlabeled target modality through a unified discrete representation space. Existing symmetric framewor...
360. Towards Isolated Interventions via Almost Orthogonal Features in Language Models ​
Author: Moritz Miller, Florent Draye, Bernhard Sch"olkopf
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2602.04718v3 Announce Type: replace Abstract: A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not substa...
361. Compressing LLMs with MoP: Mixture of Pruners ​
Author: Bruno Lopes Yamamoto, Lucas Lauton de Alcantara, Victor Zacarias, Leandro Giusti Mugnaini, Keith Ando Ogawa, Lucas Pellicer, Rosimeire Pereira Costa, Edson Bollis, Anna Helena Reali Costa, Artur Jordao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2602.06127v2 Announce Type: replace Abstract: The high computational demands of Large Language Models (LLMs) motivate methods that reduce parameter count and accelerate inference. In response, model pruning emerges as an effective strategy, yet current methods typically focus on a single dimen...
362. Deep learning approaches show promise for predicting childhood malnutrition: A comparative study with traditional machine learning methods using survey data ​
Author: Deepak Bastola, Yang Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2602.10381v2 Announce Type: replace Abstract: Childhood malnutrition remains a major public health concern in Nepal and other low-resource settings, while conventional case-finding approaches are labor-intensive and frequently unavailable in remote areas. This study provides one of the first a...
363. LLM-Based Scientific Equation Discovery via Physics-Informed Token-Regularized Policy Optimization ​
Author: Boxiao Wang, Kai Li, Tianyi Liu, Chen Li, Junzhe Wang, Yifan Zhang, Jian Cheng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.10576v2 Announce Type: replace Abstract: Symbolic regression aims to distill mathematical equations from observational data. Recent approaches have successfully leveraged Large Language Models (LLMs) to generate equation hypotheses, capitalizing on their vast pre-trained scientific priors...
364. Additive Control Variates Dominate Self-Normalisation in Off-Policy Evaluation ​
Author: Olivier Jeunen, Shashank Gupta
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.IR
arXiv:2602.14914v3 Announce Type: replace Abstract: Off-policy evaluation (OPE) is essential for assessing ranking and recommendation systems without costly online interventions. Self-Normalised Inverse Propensity Scoring (SNIPS) is a standard tool for variance reduction in OPE, leveraging a multipl...
365. AIFL: A Global Daily Streamflow Forecasting Model Using a Deterministic LSTM Pre-trained on ERA5-Land and Fine-tuned on IFS ​
Author: Maria Luisa Taccari, Kenza Tazi, Ois'in M. Morrison, Andreas Grafberger, Juan Colonese, Corentin Carton de Wiart, Christel Prudhomme, Cinzia Mazzetti, Matthew Chantry, Florian Pappenberger
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.app-ph
arXiv:2602.16579v2 Announce Type: replace Abstract: Reliable global streamflow forecasting is essential for flood preparedness and water resource management, yet data-driven models often suffer from a performance gap when transitioning from historical reanalysis to operational forecast products. Thi...
366. Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting ​
Author: Xinghong Fu, Yanhong Li, Georgios Papaioannou, Yoon Kim
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2602.17634v2 Announce Type: replace Abstract: Learning time series foundation models has been shown to be a promising approach for zero-shot time series forecasting across diverse time series domains. Insofar as scaling has been a critical driver of performance of foundation models in other mo...
367. Exact and Asymptotically Complete Robust Verifications of Neural Networks via Ising Solvers ​
Author: Wenxin Li, Wenchao Liu, Weihao Li, Chuan Wang, Qi Gao, Yin Ma, Hai Wei, Kai Wen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, physics.optics, quant-ph
arXiv:2603.00408v2 Announce Type: replace Abstract: We present an Ising-compatible framework for formal neural-network robustness verification under bounded input perturbations. For piecewise-linear activations, the Exact Logarithmic PWL Model (Log-PWL) provides an exact, sound, and complete formula...
368. Designing Service Systems from Textual Evidence ​
Author: Ruicheng Ao, Hongyu Chen, Siyang Gao, Hanwei Li, David Simchi-Levi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC, stat.ML
arXiv:2603.10400v2 Announce Type: replace Abstract: Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing policy, or the most effective quality control procedure. In many service systems, the primary evidence of perfor...
369. A General Deep Learning Framework for Wireless Resource Allocation under Discrete Constraints ​
Author: Yikun Wang, Yang Li, Yik-Chung Wu, Rui Zhang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.IT, math.IT
arXiv:2603.19322v2 Announce Type: replace Abstract: While deep learning (DL)-based methods have achieved remarkable success in continuous wireless resource allocation, efficient solutions for problems involving discrete variables remain challenging. This is primarily due to the zero-gradient issue i...
370. Can an Actor-Critic Optimization Framework Improve Analog Design? ​
Author: Sounak Dutta, Fin Amin, Sushil Panda, Jonathan Rabe, Yuejiang Wen, Paul Franzon
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY
arXiv:2603.24714v2 Announce Type: replace Abstract: Analog design often slows down because even small changes to device sizes or biases require expensive simulation cycles, and high-quality solutions typically occupy only a narrow part of a very large search space. While existing optimizers reduce s...
371. From Physics to Surrogate Intelligence: A Unified Electro-Thermo-Optimization Framework for TSV Networks ​
Author: Mohamed Gharib, Leonid Popryho, Inna Partin-Vaisband
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AR
arXiv:2603.29268v2 Announce Type: replace Abstract: High-density through-substrate vias (TSVs) enable 2.5D/3D heterogeneous integration but introduce significant signal-integrity and thermal-reliability challenges due to electrical coupling, insertion loss, and self-heating. Conventional full-wave f...
372. Sheaf-Laplacian Obstruction and Projection Hardness for Cross-Modal Compatibility on a Modality-Independent Site ​
Author: Tibor Sloboda
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.07632v2 Announce Type: replace Abstract: Cross-modal representations vary in how easily they can be aligned, and compatibility is generally non-transitive: two modalities may align through an intermediate modality at lower complexity than through a direct map. We introduce a reference for...
373. Gradient-Variation Regret Bounds for Unconstrained Online Learning ​
Author: Yuheng Zhao, Andrew Jacobsen, Nicol`o Cesa-Bianchi, Peng Zhao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, stat.ML
arXiv:2604.11151v2 Announce Type: replace Abstract: We develop parameter-free algorithms for unconstrained online learning with regret guarantees that scale with the gradient variation $V_T(u) = \sum_{t=2}^T |\nabla f_t(u)-\nabla f_{t-1}(u)|^2$. For $L$-smooth convex losses, we provide fully-adapt...
374. How Transformers Learn to Plan via Multi-Token Prediction ​
Author: Jianhao Huang, Zhanpeng Zhou, Renqiu Xia, Baharan Mirzasoleiman, Weijie Su, Wei Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2604.11912v2 Announce Type: replace Abstract: While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global structure in reasoning tasks. Multi-token prediction (MTP) has recently emerged as a promising alternative, yet its...
375. Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus ​
Author: Zijian Zhao, Jing Gao, Sen Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA
arXiv:2604.13472v2 Announce Type: replace Abstract: Cooperative multi-agent reinforcement learning (MARL) is widely used to address large joint observation and action spaces by decomposing a centralized control problem into multiple interacting agents. However, such decomposition often introduces ad...
376. MolCryst-MLIPs: A Machine-Learned Interatomic Potentials Database for Molecular Crystals ​
Author: Adam Lahouari, Shen Ai, Jihye Han, Jillian Hoffstadt, Philipp Hoellmer, Charlotte Infante, Pulkita Jain, Sangram Kadam, Maya M. Martirossyan, Amara McCune, Hypatia Newton, Shlok J. Paul, Willmor Pena, Jonathan Raghoonanan, Sumon Sahu, Oliver Tan, Andrea Vergara, Jutta Rogal, Mark E. Tuckerman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, physics.comp-ph
arXiv:2604.13897v2 Announce Type: replace Abstract: We present an open Molecular Crystal (MC) database of Machine-Learned Interatomic Potentials (MLIP) called MolCryst-MLIPs. The first release comprises fine-tuned MACE models for nine molecular crystal systems---Benzamide, Benzoic acid, Coumarin, Du...
377. Corner Reflector Array Jamming Discrimination Using Multi-Dimensional Micro-Motion Features with Frequency Agile Radar ​
Author: Jie Yuan, Lei Wang, Yanhao Wang, Yimin Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2604.16008v3 Announce Type: replace Abstract: This paper introduces a robust discrimination method for distinguishing real ship targets from corner-reflector-array jamming with frequency-agile radar. The key idea is to exploit the multidimensional micro-motion signatures that separate rigid sh...
378. FRIGID: Scaling Diffusion-Based Molecular Generation from Mass Spectra at Training and Inference Time ​
Author: Montgomery Bohde, Hongxuan Liu, Mrunali Manjrekar, Magdalena Lederbauer, Shuiwang Ji, Runzhong Wang, Connor W. Coley
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, q-bio.QM
arXiv:2604.16648v2 Announce Type: replace Abstract: Tandem mass spectrometry is prominent in scientific discovery workflows for identifying unknown small molecules, yet high-throughput structural elucidation remains challenging. While recent autoregressive and graph diffusion models have shown promi...
379. Fourier Weak SINDy: Spectral Test Function Selection for Robust Model Identification ​
Author: Zhiheng Chen, Urban Fasel, Anastasia Bizyaeva
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.DS
arXiv:2604.20141v2 Announce Type: replace Abstract: We introduce Fourier Weak SINDy, a minimal noise-robust and interpretable derivative-free equation learning method that combines weak-form sparse equation learning with spectral density estimation for data-driven test function selection. By using o...
380. Accelerating Frequency Domain Diffusion Models with Error-Feedback Event-Driven Caching ​
Author: Dong Liu, Yanxuan Yu, Ying Nian Wu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2604.22901v2 Announce Type: replace Abstract: Diffusion models achieve remarkable success in time series generation. However, slow inference limits their practical deployment. We propose E$^2$-CRF (Error-Feedback Event-Driven Cumulative Residual Feature caching) to accelerate frequency domain ...
381. Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Dynamics Models ​
Author: Julia Berger, Bernd Frauenknecht, Sebastian Trimpe, Bastian Leibe
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2604.25416v2 Announce Type: replace Abstract: Model-based reinforcement learning distinguishes between dynamics models operating on proprioceptive states and latent dynamics models typically operating on high-dimensional image observations. Among the latter, Dreamer's Recurrent State Space Mod...
382. Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis ​
Author: Michael A. Riegler, Birk Sebastian Frostelid Torpmann-Hagen
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2605.03160v2 Announce Type: replace Abstract: The standard protocol for interpreting sparse-autoencoder (SAE) features labels each feature from its top-activating contexts and validates the label by steering that single feature at a typical magnitude. We argue that this inspects one cell of a ...
383. Are Flat Minima an Illusion? ​
Author: Michael Timothy Bennett
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.05209v2 Announce Type: replace Abstract: Flat minima are an account of why deep networks generalise. However flatness is a matter of form (parameters), while generalisation is of function. The same function can be a result of many different parameterisations. I demonstrate this by rescali...
384. Autoregressive One-Step Generative Modeling for Dynamical System Forecasting ​
Author: Tianyue Yang, Xiao Xue
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, physics.flu-dyn
arXiv:2605.05540v2 Announce Type: replace Abstract: Fast surrogate modeling for high-dimensional physical dynamics requires more than low short-term error: useful models must roll out efficiently while preserving the statistical structure of long trajectories. Neural operators provide inexpensive au...
385. Emergent Symbolic Structure in Health Foundation Models: Extraction, Alignment, and Cross-Modal Transfer ​
Author: Gajendra Katuwal, Advait Koparkar, Salar Abbaspourazad, Anshuman Mishra, Sarvesh Kirthivasan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2605.07407v2 Announce Type: replace Abstract: We show that information can be transferred post-hoc across independently trained health foundation models (FMs), each pretrained on ~20M minutes of wearable sensor data from ~172K participants, by aligning their data-dependent coordinate systems. ...
386. Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory ​
Author: Daniel Goldstein, Navneel Singhal, Eugene Cheah
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2605.09877v4 Announce Type: replace Abstract: Recall presents a difficult choice: transformers have a linearly growing memory that slows each successive token, while linear RNNs typically have fixed costs but limited recall. We present Key-Value Means ("KVM"), a novel block-recurrence for atte...
387. Predicting Channel Closures in the Lightning Network with Machine Learning ​
Author: Simone Antonelli, Vincent Davis, Harrison Rush, Anthony Potdevin, Jesse Shrader, Vikash Singh, Emanuele Rossi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.SI
arXiv:2605.12759v2 Announce Type: replace Abstract: The Lightning Network (LN) is a second-layer protocol for Bitcoin designed to enable fast and cost-efficient off-chain transactions. Channels in the LN can be closed either by mutual agreement or unilaterally through a forced closure, which locks t...
388. Efficient Online Conformal Selection with Limited Feedback ​
Author: Sreenivas Gollapudi, Kostas Kollias, Kamesh Munagala, Ali Sinop
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2605.14953v3 Announce Type: replace Abstract: We address the problem of conformal selection, where an agent must select a low-cost subset of options to ensure that at least one "success" is identified at a pre-specified target rate $\phi$. While traditional online conformal prediction focuses ...
389. Distance-Matrix Wasserstein Statistics for Scalable Gromov--Wasserstein Learning ​
Author: Ao Xu, Tieru Wu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2605.14981v2 Announce Type: replace Abstract: Gromov--Wasserstein (GW) distances compare graphs, shapes, and point clouds through internal distances, without requiring a common coordinate system. This invariance is powerful, but discrete GW is a nonconvex quadratic optimal transport problem an...
390. WeCon: An Efficient Weight-Conditioned Neural Solver for Multi-Objective Combinatorial Optimization Problems ​
Author: Xuan Wu, Jinbiao Chen, Yang Li, Lijie Wen, Chunguo Wu, Yuanshu Li, Yubin Xiao, Chunyan Miao, You Zhou, Di Wang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2605.22876v2 Announce Type: replace Abstract: Existing neural solvers for Multi-Objective Combinatorial Optimization Problems (MOCOPs) commonly adopt decomposition-based strategies that scalarize a MOCOP into multiple subproblems associated with distinct weight vectors. However, they either in...
391. When Good Equations Get Bad Scores: Improving Symbolic Regression Through Better Parameter Optimization ​
Author: Boxiao Wang, Kai Li, Zhiwei Chen, Yang Huang, Runxiang Wang, Ziwen Zhang, Yifan Zhang, Jian Cheng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2605.23272v3 Announce Type: replace Abstract: Symbolic Regression (SR) plays a central role in scientific knowledge discovery by distilling mathematical equations from observational data. Most existing SR methods function within a bi-level optimization framework: an outer loop that searches fo...
392. Kan Extension Transformers: A Categorical Unification of Attention, Diffusion, and Predict-Detach Self-Conditioning ​
Author: Sridhar Mahadevan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2605.27259v2 Announce Type: replace Abstract: We propose Kan Extension Transformers (KETs) as a categorical design language for a diverse group of Transformer implementations. A layer can be viewed generally as a weighted structured extension operator: attention uses token neighborhoods, geome...
393. The Interplay Between Interpolation and Aggregation in Regression: Optimal Sample Complexity ​
Author: Mikael M{\o}ller H{\o}gsgaard, Kasper Green Larsen, Liang-Yu Zou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2605.29819v2 Announce Type: replace Abstract: This work investigates theoretically the interplay between interpolation and aggregation in regression. We establish that the $\gamma$-graph dimension characterizes learnability for a broad class of natural aggregation procedures. Furthermore, we p...
394. Universal Decision Learners ​
Author: Sridhar Mahadevan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2605.30694v2 Announce Type: replace Abstract: Many theories of decision making -- planning, reinforcement learning, causal intervention, online learning, and game-theoretic equilibrium -- turn local information into globally coherent behavior. This paper proposes a common categorical formulati...
395. Graph Neural Networks for Predicting Solvability of Finite Groups ​
Author: Tal Weissblat
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, math.GR
arXiv:2606.07619v3 Announce Type: replace Abstract: We present a Graph Neural Network (GNN) framework for the classification of finite groups according to their solvability. Using undirected Cayley graph representations, the proposed framework learns to distinguish solvable and non-solvable groups d...
396. From inverse problems to neural operators: prediction, mechanism, and generalization of data-driven models ​
Author: Conor Rowan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2606.08956v3 Announce Type: replace Abstract: Scientists have historically relied on mathematical models based on differential equations to relate system inputs -- forces, fluxes, or heat sources -- to outputs, such as displacement, velocity, concentration, and temperature. These models rely o...
397. daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization ​
Author: Dayuan Fu, Mohan Jiang, Tongyu Wang, Dian Yang, Jiarui Hu, Liming Liu, Jinlong Hou, Pengfei Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2606.16497v3 Announce Type: replace Abstract: GPU kernel optimization represents a paradigm where functional correctness is assumed and execution efficiency is the objective. We present daVinci-kernel, a reinforcement learning framework that couples skill discovery with skill exploitation thro...
398. Latent Confounded Causal Discovery via Lie Bracket Geometry ​
Author: Sridhar Mahadevan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.19610v2 Announce Type: replace Abstract: We study causal discovery from observational and interventional regimes when latent variables may affect the measured system. Our first algorithm, BRIDGE (Bracket Residuals for Interventional Discovery and Geometric Estimation), combines a density-...
399. When Average Calibration Fails: Site-Conditional Federated Conformal Risk Control ​
Author: Nafis Fuad Shahid
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2606.20115v3 Announce Type: replace Abstract: Conformal risk control (CRC) provides distribution-free segmentation guarantees by calibrating a prediction-set threshold on held-out data. In federated deployments, the standard approach pools calibration scores into a single threshold. We quantif...
400. Inverse Reinforcement Learning for Interpretable Keystroke Biomarkers in Parkinson's Disease ​
Author: Navin Bondade
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2606.25270v3 Announce Type: replace Abstract: Keystroke dynamics offer a passive window into motor function, but existing work extracts aggregate typing statistics and trains classifiers for PD/control discrimination, foregoing interpretability and rarely reporting reliability. We instead appl...
401. NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning ​
Author: Tianlin Pan, Lianyu Pang, Cheng Da, Huan Yang, Changqian Yu, Kun Gai, Wenhan Luo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2606.27771v3 Announce Type: replace Abstract: Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that are not captured by the reward proxy. We identify a simple structural signature of this drift: acro...
402. Accelerating Hierarchical Sparse Predictive Coding with Hybrid Amortized Inference ​
Author: Kazuhisa Fujita
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2606.27802v2 Announce Type: replace Abstract: Hierarchical predictive coding provides an interpretable framework for perception as error-driven inference in multi-layer models, while sparse coding imposes parsimonious latent representations through explicit sparsity constraints. Their combinat...
403. PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF ​
Author: Doo Hwan Hwang, Kee-Eung Kim
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2606.29758v2 Announce Type: replace Abstract: Reinforcement Learning from Human Feedback (RLHF) for Large Language Models increasingly relies on critic-free methods as a practical alternative to actor--critic training. Despite their simplicity, existing critic-free approaches propagate a traje...
404. QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting ​
Author: Shah Nawaz Haider, Steve Austin, Arnab Barua, Sarowar Morshed Shawon, Hadaate Ullah
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.02632v2 Announce Type: replace Abstract: Time-series forecasting supports decisions in finance, en-ergy, transportation, public health, and industrial monitoring. Recent foundation models improve transfer across forecast-ing tasks, but many depend on centralized data and Trans-former atte...
405. Beyond travel mode: urban context shapes active mobility's mental health effects over time ​
Author: Shujuan Chen, Yue Li, Ying Jin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.04520v2 Announce Type: replace Abstract: Active mobility is widely promoted for sustainable and healthier living, but whether it translates into equitable mental health benefits across individuals and places over time remains unknown. Using causal machine learning and causal deep learning...
406. Multi-Turn On-Policy Distillation with Prefix Replay ​
Author: Baohao Liao, Hanze Dong, Christof Monz, Xinxing Xu, Li Dong, Furu Wei
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL, stat.ML
arXiv:2607.04763v3 Announce Type: replace Abstract: We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each upda...
407. Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics ​
Author: Byung Gyu Chae
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.10923v2 Announce Type: replace Abstract: Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood. Cognitive Field Theory predicts that learning organizes collective dynamics through the infrared ...
408. Multi-dimensional training-priority weighting based on physical information propagation paths: a unified residual-weighting framework for physics-informed neural networks ​
Author: Zhangyi Lian, Xinda Dong, Wenxuan Huo, Weifeng Huang, Greg Zhu, Qiang He
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.11094v2 Announce Type: replace Abstract: Physics-informed neural networks (PINNs) have shown promise for solving partial differential equations (PDEs); however, their synchronous optimization treats residuals of different regions and constraints equally, which is inconsistent with the pro...
409. LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention ​
Author: Ziqi Yin, Jianyang Gao, Peiqi Yin, Jiangneng Li, Gao Cong
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.11976v3 Announce Type: replace Abstract: Indexer-TopK, the operation to compute the scores and select the top-k candidates, is widely used by sparse attention algorithms in large language models and vector retrieval in recommendation systems and vector databases. However, existing GPU-bas...
410. Implementations of Quantum and Classical Topology-Aligned Architectures for Molecular Property Prediction ​
Author: James T. Pegg, Hubert Okadome Valencia, Ronin Wu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.13737v2 Announce Type: replace Abstract: For low-data and resource-constrained regimes typical of quantum chemistry, parameter-efficient learning is a key objective. Here, we propose a topology-aligned inductive bias in which the model architecture mirrors the molecular bond graph: atoms ...
411. LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget ​
Author: Changhai Zhou, Kieran Liu, Yuhua Zhou, Qian Qiao, Jun Gao, Harry Zhang, Irvine Lu, Nolan Ho, Lucian Li, Andrew Lei, Cleon Cheng, Steven Chiang, Yihang Zeng, Di Zhang, Rio Yang, Kaijie Chen, Andrew Chen, Pony Ma, Weizhong Zhang, Cheng Jin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.DC
arXiv:2607.14952v3 Announce Type: replace Abstract: Long-context RL post-training is constrained by the lifetime of state and gradients, not attention cost alone. In GRPO, one multi-million-token prompt must serve old-policy and reference scoring plus multiple policy responses, while conventional au...
412. Knowledge-Guided Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Label-Conditioned Contrastive Alignment ​
Author: Xinran Liu, Yuwen Li, Hongxiang Gao, Heyang Xu, Jianqing Li, Zongmin Wang, Chengyu Liu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.15928v2 Announce Type: replace Abstract: Adult and pediatric electrocardiogram (ECG) interpretation relies on age-sensitive criteria, and models pretrained mainly on adult ECGs often transfer poorly to pediatric populations when pediatric labels are scarce. Existing multimodal ECG--text m...
413. JAGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models ​
Author: Ruiyi Ding, Jie Li, He Kang, Ziyan Liu, Chengru Song, Yuan cheng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV, cs.SY, eess.SY
arXiv:2607.17572v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) is a powerful reinforcement learning algorithm for aligning generative models with human preferences. While successful in large language models~\cite{shao2024deepseekmathpushinglimitsmathematical}, its exte...
414. The Label Complexity of Class-Conditional Coverage under Distribution Shift ​
Author: Weijia Han, Lisha Qu
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CV
arXiv:2607.18088v2 Announce Type: replace Abstract: Conformal prediction certifies that a classifier's prediction sets cover the truth, and that certificate is marginal. Many recognition benchmarks build distribution shift into evaluation, placing disjoint conditions in the training and test splits....
415. SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions ​
Author: Hoang-Thang Ta
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.18290v2 Announce Type: replace Abstract: In recent years, Kolmogorov-Arnold Networks (KANs) have attracted increasing attention due to their effectiveness in machine learning and scientific computing, offering a new paradigm for neural network design. In this paper, we present SechKAN, a ...
416. Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary ​
Author: Jan Kirin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.18553v3 Announce Type: replace Abstract: Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes? We test both questions in a frozen 2.6B looped transformer, Ouro-RLTT. On GSM8K, a strict pre-answer probe ex...
417. A Self-Evolving Default Action for Cooperative Tasks with Continuous Action Space ​
Author: Shuangyao Huang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.MA
arXiv:2607.18597v3 Announce Type: replace Abstract: Counterfactual credit assignment has proven effective in multi-agent reinforcement learning (MARL) for discrete action spaces, yet its extension to continuous-action cooperative tasks remains challenging. Existing methods that approximate the count...
418. H$^2$SD: Hybrid Hindsight Self-Distillation ​
Author: Qiye Cai, Yichuan Ma, Peiji Li, Yongkang Chen, Qipeng Guo, Yicheng Zou, Linyang Li, Xiaocheng Feng, Bing Qin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.18955v4 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level guidance. Existing self-distillation methods add a privileged teache...
419. Generating Bearing Vibration Signals at User-Specified Fault Probabilities Using PR-GAN and Counterfactual Methods ​
Author: Seyed Mohammadreza Alavi, Ardeshir Shojaeinasab, Reza Jalayer, Masoud Jalayer, Behnam Bahrak
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.19455v2 Announce Type: replace Abstract: In bearing vibration datasets, most samples receive predicted fault probabilities close to 0 or 1, while samples with intermediate (gray-zone) probabilities are rare. Such borderline samples are important because they reflect conditions in which ma...
420. PhantomFill: When the Form Demands an Answer, Language Models Invent One ​
Author: Rana Muhammad Usman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.20492v2 Announce Type: replace Abstract: Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself causes hallucination. We ask thirteen models the same question about the same input and change on...
421. Adaptive Multi-Horizon Reinforcement Learning ​
Author: Manoosh Samiei, Doina Precup, Paul Masset
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI
arXiv:2607.20656v2 Announce Type: replace Abstract: Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL), this trade-off is typically controlled through a fixed discount factor, which imposes a single ...
422. The Weight of Silence: A Causal Case for Weights Over the Scratchpad in Latent Chess Reasoning ​
Author: Ishan S. Kshirsagar
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.CL
arXiv:2607.20952v2 Announce Type: replace Abstract: Latent, or silent, reasoning lets language models carry out intermediate computation in continuous vector space instead of words, and is widely assumed to function as an internal scratchpad the model consults during inference. Whether that assumpti...
423. Error Certificates for KV-Cache Eviction via Randomized Design ​
Author: Peng Xie
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL
arXiv:2607.21475v2 Announce Type: replace Abstract: Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest. We prove that this design cannot know what it destroyed: evicted values can be altered so that everything the serving system retains is unchang...
424. On the Identifiability of Controlled World Models ​
Author: Xiangteng Zhang, Yang Guan, Bo Zhang, Hongyang Li, Ya-Qin Zhang, Shengbo Eben Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22430v2 Announce Type: replace Abstract: World model serves as a promising tool to infer environment dynamics under high-dimensional observations and candidate actions. Recently, LeCun's JEPA provides a compelling framework for learning such models in representation space. Its action-cond...
425. Susceptible Reservoir Architectures for Regime-Conditional Volatility Forecasting ​
Author: Aliaksei Kaliutau
Published: 7/28/2026, 4:00:00 AM
Categories: cs.LG
arXiv:2607.22491v2 Announce Type: replace Abstract: Volatility forecasting is dominated by persistence and measurement noise, leaving limited residual structure for nonlinear models to exploit. We introduce Susceptible Architectures (SUSA), a reservoir-design principle for volatility forecasting, an...
426. Weighted Low-Rank Matrix Approximation: Acceleration and Applications ​
Author: Elena Tuzhilina, Trevor Hastie
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.ME
arXiv:2109.11057v2 Announce Type: replace-cross Abstract: Weighted low-rank matrix approximation (WLRMA) generalizes classical low-rank approximation and matrix completion by allowing arbitrary elementwise weights. Such formulations arise naturally in a broad class of statistical models, including g...
427. Transfer learning for conflict and duplicate detection in software requirement pairs ​
Author: Garima Malik, Savas Yildirim, Mucahit Cevik
Published: 7/28/2026, 4:00:00 AM
Categories: cs.SE, cs.LG
arXiv:2301.03709v3 Announce Type: replace-cross Abstract: Consistent and holistic expression of software requirements is important for the success of software projects. In this study, we aim to enhance the efficiency of the software development processes by automatically identifying conflicting and ...
428. Evaluation of Blood Vessel Segmentation Methods on Hard-to-Detect Vascular Structures ​
Author: Jo~ao Pedro Parella, Matheus Viana da Silva, Cesar Henrique Comin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2406.13128v2 Announce Type: replace-cross Abstract: Due to the intricate structure of vascular trees, minor segmentation errors can significantly alter connectivity patterns and increase variability in extracted morphological properties. Global metrics such as the Dice coefficient, precision, ...
429. Procedural Content Generation via Generative Artificial Intelligence ​
Author: Xinyu Mao, Wanli Yu, Kazunori D Yamada, Michael R. Zielewski
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2407.09013v2 Announce Type: replace-cross Abstract: The attempt to utilize machine learning in PCG has been made in the past. In this survey paper, we investigate how generative artificial intelligence (AI), which saw a significant increase in interest in the mid-2010s, is being used for PCG. ...
430. No Optimal Language Set Exists for Multilingual Instruction Tuning: Insights from a Linguistically-Informed Study ​
Author: G"urkan Soykan, G"ozde G"ul \c{S}ahin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.LG
arXiv:2410.07809v2 Announce Type: replace-cross Abstract: Multilingual instruction tuning (MIT) is challenged by the curse of multilinguality, data scarcity, and high computational cost. A natural hypothesis is that carefully selecting a linguistically diverse set of languages yields universally bet...
431. LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models ​
Author: Nam V. Nguyen, Thong T. Doan, Luong Tran, Van Nguyen, Quang Pham
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2411.00918v5 Announce Type: replace-cross Abstract: Mixture of experts (MoE) architectures have become a cornerstone for scaling up and are a key component in most large language models such as GPT-OSS, DeepSeek-V3, Llama-4, and Gemini-2.5. However, systematic research on MoE remains severely ...
432. ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins ​
Author: Safayat Bin Hakim, Muhammad Adil, Alvaro Velasquez, Houbing Herbert Song
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.HC, cs.LG, cs.SC
arXiv:2501.08561v5 Announce Type: replace-cross Abstract: Digital twins are increasingly used to monitor and optimize industrial systems, yet many existing frameworks remain difficult to interpret, slow to adapt, and limited in their ability to incorporate explicit domain knowledge. This paper prese...
433. Shift-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment ​
Author: Song-Lin Lv, Yu-Yang Chen, Zhi Zhou, Lan-Zhe Guo
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2501.19060v4 Announce Type: replace-cross Abstract: Vision-language models (VLMs), such as CLIP, adapt effectively to downstream tasks through prompt tuning, but fine-tuning can misalign predictive confidence and accuracy, particularly on unseen classes. Existing VLM-specific calibration metho...
434. Dementia classification from spontaneous speech using wrapper-based feature selection ​
Author: Marko Niemel"a, Mikaela von Bonsdorff, Sami "Ayr"am"o, Tommi K"arkk"ainen
Published: 7/28/2026, 4:00:00 AM
Categories: eess.AS, cs.LG, cs.SD
arXiv:2502.03484v3 Announce Type: replace-cross Abstract: Dementia encompasses a group of syndromes that impair cognitive functions such as memory, reasoning, and the ability to perform daily activities. As populations globally age, nearly 10 million new dementia cases occur annually. Clinical diagn...
435. Robustness and Cybersecurity in the EU Artificial Intelligence Act ​
Author: Henrik Nolte, Miriam Rateike, Mich`ele Finck
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CR, cs.CY, cs.LG
arXiv:2502.16184v3 Announce Type: replace-cross Abstract: The EU Artificial Intelligence Act (AIA) establishes different legal principles for different types of AI systems. While prior work has sought to clarify some of these principles, little attention has been paid to robustness and cybersecurity...
436. Hybrid AI-Physical Modeling for Penetration Bias Correction in X-band InSAR DEMs: A Greenland Case Study ​
Author: Islam Mansour, Georg Fischer, Ronny Haensch, Irena Hajnsek
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG
arXiv:2504.08909v2 Announce Type: replace-cross Abstract: Digital elevation models derived from Interferometric Synthetic Aperture Radar (InSAR) data over glacial and snow-covered regions often exhibit systematic elevation errors, commonly termed "penetration bias." We leverage existing physics-base...
437. PD$^3$: A Project Duplication Detection Framework via Adapted Multi-Agent Debate ​
Author: Dezheng Bao, Yueci Yang, Chutian Yu, Xin Chen, Zeguo Fei, Xiang Yuan, Lijun Zhang, Jiangqian Huang, Zhengxuan Jiang, Daoze Zhang, Junru Chen, Yang Yang
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG
arXiv:2505.17492v2 Announce Type: replace-cross Abstract: Project duplication detection is critical for project quality assessment because it helps avoid investment in repeated proposals. Existing methods usually cast it as ranking and rely on surface matching or direct large language models judging...
438. Farm-LightSeek: An Edge-centric Multimodal Agricultural IoT Data Analytics Framework with Lightweight LLMs ​
Author: Dawen Jiang, Zhishu Shen, Qiushi Zheng, Tiehua Zhang, Wei Xiang, Jiong Jin
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2506.03168v2 Announce Type: replace-cross Abstract: Amid the challenges posed by global population growth and climate change, traditional agricultural Internet of Things (IoT) systems is currently undergoing a significant digital transformation to facilitate efficient big data processing. Whil...
439. TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law ​
Author: Zheng Hui, Yijiang River Dong, Ehsan Shareghi, Nigel Collier
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.CY, cs.LG
arXiv:2507.21134v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their domain-specific safety and compliance becomes critical. While prior work has largely focused on...
440. Realizing Scaling Laws in Recommender Systems: A Foundation-Expert Paradigm for Hyperscale Model Deployment ​
Author: Dai Li, Kevin Course, Wei Li, Hongwei Li, Jie Hua, Yiqi Chen, Zhao Zhu, Rui Jian, Xuan Cao, Bi Xue, Yu Shi, Jing Qian, Kai Ren, Matt Ma, Qunshu Zhang, Rui Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.IR, cs.AI, cs.LG
arXiv:2508.02929v3 Announce Type: replace-cross Abstract: Scaling laws have been established for recommender systems, yet efficiently deploying foundation model (FM) across multiple recommendation surfaces remains a major unsolved challenge. Existing methods for transfer learning face fundamental li...
441. Machine Learning for Cloud Detection in IASI Measurements: A Data-Driven SVM Approach with Physical Constraints ​
Author: Chiara Zugarini, Cristina Sgattoni, Luca Sgheri
Published: 7/28/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.LG
arXiv:2508.10120v2 Announce Type: replace-cross Abstract: Cloud detection is fundamental for the interpretation and operational exploitation of hyperspectral infrared sounders, yet the capability of infrared radiances alone to provide reliable cloud information remains insufficiently assessed. We in...
442. Using Reinforcement Learning to Optimize the Global and Local Crossing Number ​
Author: Timo Brand, Henry F"orster, Stephen Kobourov, Daniel Kohrt, Robin Schukrafft, Markus Wallinger, Johannes Zink
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CG, cs.LG
arXiv:2509.06108v3 Announce Type: replace-cross Abstract: Graph drawing concerns the algorithmic visualization of graphs. A good drawing of a graph is easy to read and facilitates solving tasks on the graph. Several properties have been identified to occur in good drawings of graphs. Such properties...
443. MaskAttn-SDXL: Controllable Region-Level Text-To-Image Generation ​
Author: Yu Chang, Jiahao Chen, Anzhe Cheng, Paul Bogdan
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2509.15357v3 Announce Type: replace-cross Abstract: Diffusion models have achieved strong results in text-to-image generation, but important limitations remain as prompts become more structured and multi-object. On the architecture side, U-Net backbones are efficient and stable, yet their loca...
444. VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference ​
Author: Jiaming Tang, Yufei Sun, Yilong Zhao, Shang Yang, Yujun Lin, Zhuoyang Zhang, James Hou, Yao Lu, Zhijian Liu, Song Han
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2512.01031v2 Announce Type: replace-cross Abstract: Vision-Language-Action models (VLAs) are becoming increasingly capable across diverse robotic tasks. However, these models are typically deployed under synchronous inference, where the robot waits for model inference to complete before acting...
445. Self-Motivated Growing Neural Network for Adaptive Architecture via Local Structural Plasticity ​
Author: Yiyang Jia, Chengxu Zhou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.NE, cs.LG
arXiv:2512.12713v2 Announce Type: replace-cross Abstract: Control policies are often implemented with fixed-capacity multilayer perceptrons trained by backpropagation, which require architecture selection in advance and cannot adapt their capacity during learning. This paper introduces the Self-Moti...
446. INSIGHT: Spatially resolved survival modelling from routine histology crosslinked with molecular profiling reveals prognostic epithelial-immune axes in stage II/III colorectal cancer ​
Author: Piotr Keller, Mark Eastwood, Zedong Hu, Aim'ee Selten, Ruqayya Awan, Gertjan Rasschaert, Sara Verbandt, Vlad Popovici, Hubert Piessevaux, Hayley T Morris, Petros Tsantoulis, Thomas Alexander McKee, Andr'e D'Hoore, C'edric Schraepen, Xavier Sagaert, Gert De Hertogh, Sabine Tejpar, Fayyaz Minhas
Published: 7/28/2026, 4:00:00 AM
Categories: q-bio.QM, cs.LG
arXiv:2512.22262v2 Announce Type: replace-cross Abstract: Routine histology contains rich prognostic information in stage II/III colorectal cancer, much of which is embedded in complex spatial tissue organisation. We present INSIGHT, a graph neural network that predicts survival directly from routin...
447. The Optimal Sample Complexity of Linear Contracts ​
Author: Mikael M{\o}ller H{\o}gsgaard
Published: 7/28/2026, 4:00:00 AM
Categories: cs.GT, cs.AI, cs.LG
arXiv:2601.01496v3 Announce Type: replace-cross Abstract: In this paper, we settle the problem of learning optimal linear contracts from data in the offline setting, where agent types are drawn from an unknown distribution and the principal's goal is to design a contract that maximizes her expected ...
448. Operator learning for models of tear film breakup ​
Author: Qinying Chen, Arnab Roy, Tobin A. Driscoll
Published: 7/28/2026, 4:00:00 AM
Categories: math.NA, cs.CV, cs.LG, cs.NA
arXiv:2601.08001v2 Announce Type: replace-cross Abstract: Tear film (TF) breakup is a key driver of understanding dry eye disease, yet estimating TF thickness and osmolarity from fluorescence (FL) imaging typically requires solving computationally expensive inverse problems. We propose an operator l...
449. Risk reversal for least squares estimators under nested convex constraints ​
Author: Omar Al-Ghattas
Published: 7/28/2026, 4:00:00 AM
Categories: math.ST, cs.LG, math.OC, stat.TH
arXiv:2601.16041v2 Announce Type: replace-cross Abstract: In constrained stochastic optimization, one expects that restricting the feasible set, provided it still contains the true parameter, should not increase the statistical risk of the corresponding projection estimator. We show that this intuit...
450. Event Driven Clustering Algorithm ​
Author: David El-Chai Ben-Ezra, Adar Tal, Daniel Brisk
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2602.00115v2 Announce Type: replace-cross Abstract: This paper introduces a novel asynchronous, event-driven algorithm for real-time detection of small event clusters in event camera data. Similar to hierarchical agglomerative clustering methods, the proposed algorithm detects clusters based o...
451. Does Faithfulness-Guided Alignment Hurt Accuracy? Unlocking Accurate and Faithful Post-Retrieval Reasoning ​
Author: Yu Liu, Wenxiao Zhang, Diandian Guo, Cong Cao, Fangfang Yuan, Qiang Sun, Yanbing Liu, Jin B. Hong, Zhiyuan Ma
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.LG
arXiv:2602.01348v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) can achieve strong answer accuracy on multi-hop questions, but outcome-level rewards often leave reasoning traces weakly grounded and difficult to audit. Under noisy retrieval, models may exhibit right-ans...
452. Aortic Valve Disease Screening from PPG via Physiology-Guided Self-Supervised Learning ​
Author: Jiaze Wang, Qinghao Zhao, Zizheng Chen, Zhejun Sun, Deyun Zhang, Yuxi Zhou, Shenda Hong
Published: 7/28/2026, 4:00:00 AM
Categories: eess.SP, cs.LG
arXiv:2602.04266v2 Announce Type: replace-cross Abstract: Aortic valve disease (AVD) represents a major public health burden, while its diagnosis relies on echocardiography, which is limited by cost and specialist expertise, restricting scalable screening and risk stratification. Existing portable s...
453. Parallel Swin Transformer-Enhanced 3D MRI-to-CT Synthesis for MRI-Only Radiotherapy Planning ​
Author: Zolnamar Dorjsembe, Hung-Yi Chen, Furen Xiao, Hsing-Kuo Pao
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2602.05387v2 Announce Type: replace-cross Abstract: MRI provides superior soft tissue contrast without ionizing radiation; however, the absence of electron density information limits its direct use for dose calculation. As a result, current radiotherapy workflows rely on combined MRI and CT ac...
454. Multi-Task GRPO: Reliable LLM Reasoning Across Tasks ​
Author: Shyam Sundhar Ramesh, Xiaotong Ji, Matthieu Zimmer, Sangwoong Yoon, Zhiyong Wang, Haitham Bou Ammar, Aurelien Lucchi, Ilija Bogunovic
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2602.05547v2 Announce Type: replace-cross Abstract: RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deployment requires reliable performance across diverse tasks. A straightforward multi-task adaptation of GRPO...
455. Gradient Networks for Universal Magnetic Modeling of Synchronous Machines ​
Author: Junyi Li, Tim Foissner, Floran Martin, Antti Piippo, Marko Hinkkanen
Published: 7/28/2026, 4:00:00 AM
Categories: eess.SY, cs.LG, cs.SY
arXiv:2602.14947v3 Announce Type: replace-cross Abstract: This paper presents a physics-constrained neural network framework for dynamic modeling of saturable synchronous machines, including spatial harmonics. The proposed architecture embeds gradient networks directly into the fundamental machine e...
456. Self-Distillation of Hidden Layers for Self-Supervised Representation Learning ​
Author: Scott C. Lowe, Anthony Fuller, Sageev Oore, Evan Shelhamer, Graham W. Taylor
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2603.15553v2 Announce Type: replace-cross Abstract: The landscape of self-supervised learning (SSL) is currently dominated by generative approaches (e.g. MAE) that reconstruct raw low-level data, and predictive approaches (e.g. I-JEPA) that predict high-level abstract embeddings. While generat...
457. LanteRn: Latent Visual Structured Reasoning ​
Author: Andr'e G. Viveiros, Nuno Gon\c{c}alves, Matthias Lindemann, Andr'e Martins
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2603.25629v2 Announce Type: replace-cross Abstract: While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs). As a result, most LMMs default to verbalizing perceptual content into text, a strong limitation for tasks re...
458. PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms ​
Author: Wei Wang, Tianyu Shi, Shuai Zhang, Boyang Xia, Zequn Xie, Chenyu Zeng, Qi Zhang, Lynn Ai, Yaqi Yu, Kaiming Zhang, Feiyue Tang, Lei Ding
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG
arXiv:2603.27476v2 Announce Type: replace-cross Abstract: AI-powered people search platforms are increasingly used in recruiting, sales prospecting, and professional networking, yet no widely accepted benchmark exists for evaluating their performance. We introduce PeopleSearchBench, an open-source b...
459. AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models ​
Author: Mozhgan Pourkeshavarz, Tianran Liu, Nicholas Rhinehart
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.LG
arXiv:2603.28963v2 Announce Type: replace-cross Abstract: Simulation with realistic traffic agents is essential for validating autonomous driving systems. Existing data-driven simulators learn agent behavior from higher-level abstractions such as 3D bounding boxes and polylines, inferred by upstream...
460. Transferable FB-GNN-MBE Framework for Potential Energy Surfaces: Data-Adaptive Transfer Learning in Deep Learned Many-Body Expansion Theory ​
Author: Siqi Chen, Zhiqiang Wang, Yili Shen, Xianqi Deng, Xi Cheng, Cheng-Wei Ju, Jun Yi, Guo Ling, Dieaa Alhmoud, Hui Guan, Zhou Lin
Published: 7/28/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.LG
arXiv:2604.09320v3 Announce Type: replace-cross Abstract: Mechanistic understanding and rational design of complex chemical systems depend on fast and accurate predictions of electronic structures beyond individual building blocks. However, if the system exceeds hundreds of atoms, first-principles q...
461. Unraveling the Mechanism of Drug Binding to SARS-CoV-2 RNA Pseudoknot with Thermodynamics-Driven Machine Learning ​
Author: Mariia Ivonina, Jakub Rydzewski
Published: 7/28/2026, 4:00:00 AM
Categories: physics.bio-ph, cs.LG
arXiv:2604.14906v4 Announce Type: replace-cross Abstract: The pseudoknot secondary structure in SARS-CoV-2 RNA is essential for regulating protein synthesis through $-$1 programmed ribosomal frameshifting ($-1$ PRF), a mechanism that allows the virus to generate both structural and non-structural pr...
462. MOCA: A Transformer-based Modular Causal Inference Framework with One-way Cross-attention and Cutting Feedback ​
Author: Lei Wang, Debashis Ghosh
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.ME
arXiv:2604.23107v2 Announce Type: replace-cross Abstract: Causal effect estimation from observational data requires careful adjustment for confounding. Classical estimators such as inverse probability weighting and augmented inverse probability weighting can perform well under favorable model specif...
463. Variational Neural Belief Parameterizations for Robust Dexterous Grasping under Multimodal Uncertainty ​
Author: Clinton Enwerem, Shreya Kalyanaraman, John S. Baras, Calin Belta
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.LG, cs.SY, eess.SY
arXiv:2604.25897v2 Announce Type: replace-cross Abstract: Contact variability, sensing uncertainty, and external disturbances make grasp execution stochastic. Expected-quality objectives ignore tail outcomes and often select grasps that fail under adverse contact realizations. Risk-sensitive POMDPs ...
464. Principles and Guidelines for Randomized Controlled Trials in AI Evaluation ​
Author: Christopher Kelly, Angelica Chowdhury, Alexandra Campili, Bimpe Ayoola, Devin Barbour, Thomas Chen Dawson, Ze Shen Chin, Rokas Gipi\v{s}kis
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC, cs.LG
arXiv:2605.02050v2 Announce Type: replace-cross Abstract: This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established practices from disciplines with established RCT traditions, including software engineering, economics, clin...
465. Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs ​
Author: Sohan Venkatesh
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.LG
arXiv:2605.09239v2 Announce Type: replace-cross Abstract: Large language models fail at counting how many times a word repeats in a list, even though they perform well on far harder reasoning tasks. These failures are commonly attributed to limitations in internal count tracking. We show this attrib...
466. A Cascaded Edge-Cloud Architecture for Automated Diabetic Retinopathy Screening ​
Author: Nishi Doshi, Shrey Shah
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG
arXiv:2605.14108v2 Announce Type: replace-cross Abstract: Diabetic Retinopathy (DR) is one of the leading causes of preventable blindness, and automated screening can help extend specialist capacity in resource-constrained clinical workflows. Cloud-based deep learning systems can provide strong grad...
467. Beyond Adoption Intention How Trust in Augmented Analytics Relates to Perceived Decision Quality Among Non-Technical BI Users ​
Author: Thuy Pham Thi Phuong, Hieu Vu Le Trung, Ha Nguyen Manh, Ngan Nguyen Thi Thuy, Lan Hoang Thi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.HC, cs.CY, cs.LG
arXiv:2605.20198v2 Announce Type: replace-cross Abstract: Augmented analytics has transformed how Business Intelligence (BI) systems support decision-making, shifting non-technical managers from manual analysis toward dependence on automated insights. Current BI research often overlooks the cognitiv...
468. A2QTGN: Adaptive Amplitude Quantum-Integrated Temporal Graph Network for Dynamic Link Prediction ​
Author: Nouhaila Innan, M. Murali Karthick, Simeon Kandan Sonar, Vivek Chaturvedi, Muhammad Shafique
Published: 7/28/2026, 4:00:00 AM
Categories: quant-ph, cs.LG
arXiv:2605.21916v2 Announce Type: replace-cross Abstract: Dynamic link prediction is important for modeling evolving interactions in social, communication, financial, and transportation networks. Classical temporal graph models capture changes over time, but they may struggle to represent rapidly ev...
469. On the Detection of Commutative Factors in Factor Graphs: Necessary and Sufficient Conditions ​
Author: Malte Luttermann, Ralf M"oller, Marcel Gehrke
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.DS, cs.LG
arXiv:2605.26908v2 Announce Type: replace-cross Abstract: Exploiting the indistinguishability of objects in a probabilistic graphical model such as a factor graph is key to lifted probabilistic inference algorithms and allows for tractable probabilistic inference problems with respect to domain size...
470. Causal Density Functions ​
Author: Sridhar Mahadevan
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ME, cs.AI, cs.LG
arXiv:2606.00754v2 Announce Type: replace-cross Abstract: We study the full density ratio between a specified intervention regime $P_a$ and an observational regime $P_0$, $\rho_a=dP_a/dP_0$, under the prerequisite $P_a\ll P_0$. We call the regime-indexed ratio a causal density function when $P_a$ is...
471. It does what it says on the tin: safe synthetic data from coarsened margins ​
Author: Gillian M Raab
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.AP
arXiv:2606.02101v3 Announce Type: replace-cross Abstract: This paper proposes a method of creating synthetic data (SD) that will have two important advantages for the user compared to other methods currently available. The first is transparency; unlike other methods, the person in receipt of the SD ...
472. Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery ​
Author: Syed Rifat Raiyan, Mohsinul Kabir, Hasan Mahmud, Md Kamrul Hasan, Sophia Ananiadou
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.CV, cs.LG
arXiv:2606.08728v4 Announce Type: replace-cross Abstract: Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it has moved from a niche problem within NLP to one of the most consequential AI frontiers. This survey provides a unified account of th...
473. Market Design for AI: Beyond the Copyright Binary ​
Author: Yan Dai, Maryam Farboodi, Negin Golrezaei, Sepehr Shahshahani
Published: 7/28/2026, 4:00:00 AM
Categories: econ.TH, cs.AI, cs.GT, cs.LG, stat.ML
arXiv:2606.12260v2 Announce Type: replace-cross Abstract: How can we design a market of human-generated content for use in training AI models that both enables technological progress and preserves individual incentives for high-quality content creation? Existing approaches take polar positions: a "f...
474. The Hitchhiker's Guide to Agentic AI: From Foundations to Systems ​
Author: Haggai Roitman
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, cs.LG
arXiv:2606.24937v2 Announce Type: replace-cross Abstract: The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. The book covers the full stack from first principles to production deployment, organized around a central thesis: building gr...
475. EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures ​
Author: Bu\u{g}ra Alperen Ulu{\i}rmak, Rifat Kurban
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG, cs.SE
arXiv:2606.30219v4 Announce Type: replace-cross Abstract: This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capa...
476. DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning ​
Author: Hengyu Fu, Tianyu Guo, Zixuan Wang, Hanlin Zhu, Jason D. Lee, Jiantao Jiao, Stuart Russell, Song Mei
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.00341v2 Announce Type: replace-cross Abstract: Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought (CoT). However, many questions require the model to internalize the multi-step reasoning within a sing...
477. TSP with Predictions: Heatmap to Tour with Provable Guarantees ​
Author: Marek Eli'a\v{s}, Fabrizio Grandoni, Adam Polak, Eleonora Vercesi
Published: 7/28/2026, 4:00:00 AM
Categories: cs.DS, cs.LG
arXiv:2607.03791v2 Announce Type: replace-cross Abstract: The Traveling Salesperson Problem (TSP) has long served as a benchmark for evaluating the strength of optimization techniques in the classical theory of algorithms. In recent efforts to apply ML to algorithmic problems, TSP has also become a ...
478. XS-VLA: Coupling Coarse-grained Spatial Distillation with Latent Flow Matching for Lightweight Robotic Control ​
Author: Iok Tong Lei, Qingchen Xie, Wei Huang, Ying Jie Yap, Yujie Zhang, Qianzhi Li, Xiaolong Liu, Zhidong Deng
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.LG
arXiv:2607.04171v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) have shown strong multimodal understanding and spatial grounding, but their computational cost limits real-time robotic control. In contrast, lightweight models are suitable for edge deployment but often s...
479. What You See Is What You Get: Observation-Aligned Supervision for Chart-to-Code Generation ​
Author: Tianhao Niu, Qingfu Zhu, Wanxiang Che
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.LG
arXiv:2607.04726v2 Announce Type: replace-cross Abstract: Chart-to-code generation is commonly trained with supervised fine-tuning on reference plotting scripts, implicitly treating the gold code as a fully observable target. We argue that this assumption is often invalid: many chart programs contai...
480. Integrating GNSS-Derived Zenith Wet Delay into a Weather Foundation Model Improves Precipitation Forecasting ​
Author: Leonardo Trentini, Fanny Lehmann, Laura Crocetti, Benedikt Soja
Published: 7/28/2026, 4:00:00 AM
Categories: physics.ao-ph, cs.LG, physics.geo-ph
arXiv:2607.05658v2 Announce Type: replace-cross Abstract: Global Navigation Satellite Systems (GNSS), best known for positioning, also serve weather science, as atmospheric water vapour delays their signals. This delay, the Zenith Wet Delay (ZWD), is a direct, all-weather measure of column moisture....
481. Enhanced Seam Segmentation for Automated Welding Robot in Construction Through Transfer Learning: Addressing Limitations of Bilateral Segmentation Network ​
Author: Keonvin Park, Yong Ann Voeurn, Hyeokjun Kweon, Doyun Lee
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CV, cs.LG
arXiv:2607.06150v2 Announce Type: replace-cross Abstract: Reliable seam segmentation is essential for autonomous robotic welding in construction, where harsh illumination, specular reflections, and thin weld geometries often degrade segmentation performance. This study proposes a reflection-robust s...
482. Finding a stationary point of a stochastic convex problem ​
Author: Felipe Areces, John Duchi, Malo Sommers
Published: 7/28/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.OC
arXiv:2607.06883v3 Announce Type: replace-cross Abstract: We consider the problem of finding stationary points for stochastic convex optimization problems. Rather than surrogates to stationarity, such as a proximity-to-stationarity guarantee or small gradient of the Moreau envelope, we ask for a str...
483. Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models ​
Author: Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim, Suyoun Kim, Bo-Ru Lu, Qingming Tang, Ankur Gandhe, Hung-yi Lee, Chieh-Chi Kao, Chao Wang
Published: 7/28/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.CL, cs.LG, cs.SD
arXiv:2607.13408v2 Announce Type: replace-cross Abstract: Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global simila...
484. Interpreting Quantum Learning Models via Stochastic Processes ​
Author: Johannes Fankhauser, Lukas J. Fiderer, Hans J. Briegel
Published: 7/28/2026, 4:00:00 AM
Categories: quant-ph, cs.LG
arXiv:2607.17327v2 Announce Type: replace-cross Abstract: Quantum machine learning models define probabilistic input--output maps through coherent quantum evolution and measurement. While such models can exhibit computational advantages, their internal functioning and decision making generally resis...
485. SLPO: Scaling Latent Reasoning via a Surrogate Policy ​
Author: Runyang You, Zhiyuan Liu, Yongqi Li, Wenjie Li
Published: 7/28/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG
arXiv:2607.19691v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must ...
486. Hard Guarantees at a Measured Price: Entropy-Stable Learned Finite Volumes for Compressible Flow ​
Author: Denis Gueyffier (ONERA -- Institut Polytechnique de Paris)
Published: 7/28/2026, 4:00:00 AM
Categories: physics.flu-dyn, cs.LG, cs.NA, math.NA
arXiv:2607.20171v2 Announce Type: replace-cross Abstract: Learned solvers for compressible flow are usually compared to classical methods at equal mesh resolution rather than at equal computational cost, and they typically offer no guarantee that their solutions remain physically admissible. We pres...
487. Emergent Compositional Skills in Mixture-of-Experts VLAs ​
Author: Shlok Shah, Rhiaan Jhaveri, Tharun Kumar Tiruppali Kalidoss, Chirayu Nimonkar, Ishaan Javali, Dhruv Shah
Published: 7/28/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG
arXiv:2607.20771v2 Announce Type: replace-cross Abstract: We consider the problem of learning compositional robot policies end-to-end from expert demonstrations, without any pre-specified notion of task decomposition or hierarchy. We ask whether a VLA trained with a simplified Mixture-of-Experts (Mo...
488. TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI ​
Author: Ritik Raj, Souvik Kundu, Sarbartha Banerjee, Dheemanth Joshi, Ishita Vohra, Tushar Krishna
Published: 7/28/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, cs.MA
arXiv:2607.22465v2 Announce Type: replace-cross Abstract: Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, age...