Skip to content

arXiv cs.LG - 2026-07-20 ​

206 items collected.


1. Structure of the Circular-Dyadic Convolution Error ​

Author: Ben Fauber, Alireza Moradzadeh
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.NA, math.NA

arXiv:2607.15293v1 Announce Type: new Abstract: Dyadic and circular convolution can both be computed in $O(N\log N)$ time using the Hadamard transform and the FFT-computed discrete Fourier transform (DFT), respectively. The Hadamard transform is preferable for its real-valued sign flips, yet its sub...

📖 Read original article


2. Position: Quantum Program Generation Must Prioritize Validity Over Probabilistic Scaling ​

Author: Junhao Song, Yu Zhou, William Knottenbelt, Yudong Cao
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15313v1 Announce Type: new Abstract: The scaling hypothesis assumes that increasing model parameters yields emergent reasoning capabilities. This position paper argues that applying this probabilistic paradigm to generic quantum circuit synthesis is a directional error. Unlike natural lan...

📖 Read original article


3. A Transportable Threshold-Based Framework for Interpretable Classification of Medical Data ​

Author: Antony Garcia, Adrian Noriega, Gabrielle Britton, Xinming Huang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15394v1 Announce Type: new Abstract: Black-box models limit the adoption of artificial intelligence in medicine due to their lack of interpretability and reproducibility. We introduce a statistically grounded framework that provides fully interpretable, rule-based clinical classification ...

📖 Read original article


4. Regularity-Aware Stochastic MGDA with Adaptive Conflict-Avoidant Update Direction Control ​

Author: Chentong Huang, Lisha Chen
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, math.OC

arXiv:2607.15412v1 Announce Type: new Abstract: Multi-objective learning (MOL) aims to optimize multiple objectives simultaneously. The multi-gradient descent algorithm (MGDA) is a workhorse that iteratively updates along a common descent or conflict-avoidant (CA) direction across objectives. In sto...

📖 Read original article


5. AI Trading: Evaluating Large Language Models for Technical Market Analysis ​

Author: Geofrey Ntale
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, q-fin.CP

arXiv:2607.15414v1 Announce Type: new Abstract: Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial markets. This paper presents a systematic, comparative evaluation of five prominent LLMs: GPT-4 Turbo, Claude 3 Op...

📖 Read original article


6. qZACH-ViT: Quantization-Aware Intrinsic Explanations with Recursive Attribution-Stabilized Optimization ​

Author: Athanasios Angelakis
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2607.15421v1 Announce Type: new Abstract: Compact medical-image classifiers need efficiency and interpretable evidence, yet these goals are often addressed separately. We introduce qZACH-ViT, a quantization-aware extension of the zero-token (CLS-token-free), position-free ZACH-ViT backbone wit...

📖 Read original article


7. From hyperplanes to hyperellipsoids: characterizing the inherent interpretability of linear and single-qubit mixed-state binary classification models ​

Author: Kaitlin Gili
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, quant-ph

arXiv:2607.15433v1 Announce Type: new Abstract: We characterize and compare the inherent interpretability offerings of a standard linear model with a single qubit mixed state model for the task of supervised binary classification. A side by side comparison reveals that a single qubit mixed state mod...

📖 Read original article


8. Stochastic Reset Pathfinding: Path-Level Regret for Cascading Bandits over Graph Paths ​

Author: Guni Sharon, Wei Zhang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15440v1 Announce Type: new Abstract: We introduce Stochastic Reset Pathfinding (SRP), an episodic learning problem on a known directed graph with unknown stationary edge success probabilities. In each episode, the agent commits to a source-to-goal path, and any edge failure during executi...

📖 Read original article


9. Who Became Financially Vulnerable After COVID-19? A Population-Level Machine Learning Analysis Using MEPS Data ​

Author: Alexey Kresin, Zien Cheng, Ammar Ahad, Ebiyomare Kelvin, Manish Sivaratri, Prabhjeet Singh, Omar Aljawfi, Olabisi Ojo, Nawar Shara
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15446v1 Announce Type: new Abstract: The cost of healthcare remains a concern in the United States and may have been influenced by disruptions associated with the COVID-19 pandemic. This study examines healthcare financial vulnerability before and after the pandemic using Medical Expendit...

📖 Read original article


10. LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models ​

Author: Jingteng Li, Alexander Capstick, Louise Rigny, Iona Biggart, Neil J Sebire, Payam Barnaghi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15447v1 Announce Type: new Abstract: Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models to foundation models, utilising modern representation learning methods. Here, foundation models are p...

📖 Read original article


11. Relevant and Irrelevant: A Renormalization Group Analysis of Transformer Attention ​

Author: Parviz Haggi-Mani, Irina Rish
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15449v1 Announce Type: new Abstract: Using the language of Wilsonian renormalization group theory (RG), we treat the Transformer's attention mechanism as a perturbation of the trained MLP residual-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operato...

📖 Read original article


12. Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers ​

Author: James O' Neill, Fergal Reid
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2607.15456v1 Announce Type: new Abstract: Looped, weight-tied Transformers reduce parameters by reusing a block, but decoding still stores a separate K/V cache for every recurrence step. We show that this loop-indexed cache is highly structured. For a fixed token, layer and head, K/V vectors t...

📖 Read original article


13. Robust Peak-cost Constrained Reinforcement Learning ​

Author: Shilpa Mukhopadhyay, Sourav Ganguly, Santosh Mohan Rajkumar, Honghao Wei, Debdipta Goswami, Arnob Ghosh
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15457v1 Announce Type: new Abstract: We study robust peak-cost constrained reinforcement learning (RP-CRL), where the objective is to maximize expected reward while controlling the maximum cost encountered along a trajectory. This setting is motivated by safety-critical applications in wh...

📖 Read original article


14. ADS-C: Antidistillation Sampling for Classification ​

Author: Khawaja Abaid Ullah, Mohammad Javad Khojasteh
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.CR

arXiv:2607.15467v1 Announce Type: new Abstract: Knowledge distillation enables an adversary to replicate a proprietary classifier by querying its prediction interface and training a surrogate on the returned probability vectors. Antidistillation sampling, proposed for large language models, counters...

📖 Read original article


15. Deep Learning Approaches for Sleep Apnea Classification from Polysomnographic EEG Signals ​

Author: Shashank Manjunath, Mukesh Cheemakurthi, Aarti Sathyanarayana
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15477v1 Announce Type: new Abstract: Sleep apnea diagnosis via polysomnography remains resource intensive and relies on time consuming manual data analysis and scoring. Recent work has demonstrated that central nervous system effects of sleep apnea events can be detected through electroen...

📖 Read original article


16. Inpainting Insights: Elevating Visual XAI with Photorealistic Perturbations ​

Author: Josef Lindl, Mariana Chaves, Damien Garreau
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15482v1 Announce Type: new Abstract: The increasing complexity of state-of-the-art machine learning models has made their behavior progressively harder to interpret, spurring rapid advancements in the field of eXplainable Artificial Intelligence (XAI). Among many methods proposed, perturb...

📖 Read original article


17. Diffusion models recover accurate mixture weights despite score function insensitivity ​

Author: Andrew Dennehy, Ramchandran Muthukumar, Rebecca Willett, Nisha Chandramoorthy
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, math.PR, stat.ML

arXiv:2607.15485v1 Announce Type: new Abstract: Score-based generative models exhibit a puzzling behavior: they often appear to cover all modes of a target multimodal distribution and yet may fail to learn the correct relative mode amplitudes, which can be interpreted as mixture weights. We resolve ...

📖 Read original article


18. An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism ​

Author: Mobina Kashaniyan, Mehrdad Ashtiani, Amirhossein Ghassemi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DC, cs.PF

arXiv:2607.15511v1 Announce Type: new Abstract: Serverless computing provides automatic resource management and pay-per-use execution, but effective autoscaling remains challenging because of dynamic workloads, cold-start latency, and dependencies among functions. We present a dependency-aware autos...

📖 Read original article


19. Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching ​

Author: Yan Song
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15516v1 Announce Type: new Abstract: Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent). The compression literature has standardized on query-aware methods that produce ...

📖 Read original article


20. Recursive Harness Self-Improvement ​

Author: Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Matei Zaharia, Yujin Tang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15524v1 Announce Type: new Abstract: Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both im...

📖 Read original article


21. Kolmogorov--Arnold Networks for Small Language Models ​

Author: Felippe Alves, Renato Vicente
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15525v1 Announce Type: new Abstract: Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks. We test these claims separ...

📖 Read original article


22. Publicly-Verifiable Certificates for Statistical Algorithms ​

Author: Michael Ngo, Michael P. Kim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.CR, cs.DS

arXiv:2607.15528v1 Announce Type: new Abstract: Following Goldwasser, Rothblum, Shafer, and Yehudayoff, who defined a framework for interactive proofs of learning [ITCS'21], we initiate the study of non-interactive proofs of learning. We define and study a new notion: Publicly-Verifiable Certificate...

📖 Read original article


23. From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation ​

Author: Himel Dev, Tanmoy Sen, Madhusudan Basak, Bashima Islam
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15552v1 Announce Type: new Abstract: Generating personalized trip itineraries is a complex planning task and involves a tension between hard combinatorial feasibility and soft latent desirability. Classical optimization enforces constraints but fails to capture subjective traveler prefere...

📖 Read original article


24. Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation ​

Author: Himel Dev, Madhusudan Basak, Tanmoy Sen, Paromita Shome, Bashima Islam
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15562v1 Announce Type: new Abstract: Packing for air travel is recurring and error-prone: the checklist must be personal and context-aware, yet feasible under safety rules, item dependencies, and luggage limits. Existing packing assistants are template-driven and generic, or recommendatio...

📖 Read original article


25. Information-Directed Sampling for Causal Bandits ​

Author: Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15577v1 Announce Type: new Abstract: Causal bandits exploit structural relationships among variables to share information across interventions and accelerate the identification of high-reward decisions. In many applications, however, some variables cannot be directly manipulated, even tho...

📖 Read original article


26. Rethinking Transfer in Continual Learning: A Replay-Based Realisation ​

Author: Yang Meng, Zhenya Liu, Zhuokai Zhao, Yuxin Chen
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15587v1 Announce Type: new Abstract: Continual learning studies how deployed language models can continually acquire new tasks without expensive retraining from scratch. Existing methods, whether rehearsal-based (replaying stored past data) or rehearsal-free (regularising or isolating par...

📖 Read original article


27. Field-Aware RankMixer with Dual-Stream Bilinear Fusion for the Tencent UNI-REC Challenge ​

Author: Yufeng Zhang, Zhengqi Xu, Jiajun Cui
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15590v1 Announce Type: new Abstract: This paper presents our solution to the KDD Cup 2026 Tencent UNIREC Challenge. The task requires joint modeling of multi-domain user behavior sequences and non-sequential multi-field features for target-ad pCVR prediction. We develop a Field-Aware Rank...

📖 Read original article


28. Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data ​

Author: Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2607.15606v1 Announce Type: new Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing, yet a generator can reproduce every marginal and every foreign-key relationship while emitting timestamps that run backwards or repeat, and while sending entit...

📖 Read original article


29. ASK-NN: An Asymmetric Nearest-Neighbor Test that detects Distribution Drifts in Natural Language ​

Author: Sergey Zakharov, Rodion Oblovatny, Alexey Zaytsev
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2607.15607v1 Announce Type: new Abstract: Hallucinations and artificial text in LLM-generated outputs often appear as distributional deviations between prompt and response hidden-state distributions. Since prompts or retrieved contexts typically serve as reference samples and responses as quer...

📖 Read original article


30. Neural Non-Equilibrium Hamiltonian Monte Carlo for Corrected Boltzmann Sampling ​

Author: Moxian Qian
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cond-mat.stat-mech, hep-lat

arXiv:2607.15682v1 Announce Type: new Abstract: Sampling from an unnormalized Boltzmann density requires proposals that move probability mass globally while retaining enough path-probability information for statistical correction. We introduce Neural Non-Equilibrium Hamiltonian Monte Carlo (NHMC), a...

📖 Read original article


31. Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework ​

Author: Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than s...

📖 Read original article


32. A Benchmark for Electrical Load Forecasting Across Grid Levels: Time-Series Transformers Outperform Established Methods ​

Author: Matthias Hertel, Sebastian P"utz, Jonathan Kolar, Benjamin Sch"afer, Ralf Mikut, Veit Hagenmeyer
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15705v1 Announce Type: new Abstract: Accurate load forecasting at multiple grid levels is essential for future smart grids, ranging from aggregated control area forecasts for balancing supply and demand to forecasts of individual end-consumer loads for demand-side management and energy ma...

📖 Read original article


33. CardioMeta: Calibrated Multi-Task Prediction of Diabetes, Hypertension, and Cardiovascular Disease Across Population and EHR Data ​

Author: S M Asif Hossain, Ruksat Khan Shayoni, M. F. Mridha, Jungpil Shin
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15721v1 Announce Type: new Abstract: Cardiometabolic diseases remain among the most persistent drivers of preventable morbidity because diabetes, hypertension, and cardiovascular disease frequently co-occur and share metabolic, vascular, demographic, and behavioral determinants. Existing ...

📖 Read original article


34. Learning Faster without Deeper Networks: A*-Inspired Batch Selection for Efficient CNN Training ​

Author: Anxhelo Shehu, Enes Stastoli, Arben Cela
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15745v1 Announce Type: new Abstract: Common practice when training Convolutional Neural Networks (CNNs) is to use randomly shuffled mini-batches. This creates two limitations: slower convergence, and a diminishing learning signal, since many samples are quickly classified as easy during t...

📖 Read original article


35. Trainable Spline Representations for Physics-Informed Learning ​

Author: Giovanni Canali, Nicola Demo, Gianluigi Rozza
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.NA

arXiv:2607.15751v1 Announce Type: new Abstract: This work introduces Physics-Informed Splines (PI-Splines), a structured spline-based architecture for physics-informed learning. Instead of representing the solution of a differential equation with a neural network, PI-Splines directly parametrize the...

📖 Read original article


36. CoG-Guided Weight Correction for Fault-Tolerant Deep Neural Networks ​

Author: Bahram Parchekani, Samira Nazari, Ali Azarpeyvand, Mohammad Hasan Ahmadilivani, Tara Ghasempouri, Jaan Raik
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AR, cs.NA, math.NA

arXiv:2607.15753v1 Announce Type: new Abstract: Deep Neural Networks (DNNs) used in safety-critical applications are vulnerable to hardware and memory faults that corrupt network weights and degrade reliability. In this paper, we propose a Center of Gravity (CoG) guided weight correction method that...

📖 Read original article


37. From Diffusion to Reaction-Diffusion: A Dynamical-Systems View of Oversmoothing in Hypergraph Neural Networks ​

Author: Zhiheng Zhou, Mengyao Zhou, Yancheng Chen, Dengyi Zhao, Xingqin Qi, Guiying Yan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15773v1 Announce Type: new Abstract: Higher-order couplings enhance the expressive power of hypergraph neural networks (HGNNs), but they also intensify representation collapse in deep propagation due to strong multi-way feature mixing. This work investigates hypergraph oversmoothing from ...

📖 Read original article


38. Scaling Time Series Classification via XAI-Driven Data Reduction ​

Author: Davide Italo Serramazza, Thach Le Nguyen, Georgiana Ifrim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15774v1 Announce Type: new Abstract: Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI, a novel methodology t...

📖 Read original article


39. AquaAugmentor: A Novel Feature Augmentation Algorithm for Water Potability Prediction ​

Author: Muntasir Tabasum, Al Zadid Sultan Bin Habib, Tanpia Tasnim, Md. Ekramul Islam, Md Younus Ahamed, Md Asif Bin Syed
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE

arXiv:2607.15775v1 Announce Type: new Abstract: Access to potable water is crucial for health, economic development, and sustainability. However, accurately classifying water quality remains a significant challenge due to the complexity and variability of water source data. This paper addresses the ...

📖 Read original article


40. Knowledge-Assisted Multi-Graph Dependency Learning for Multivariate Time Series Anomaly Detection in Multi-Stage Industrial Processes ​

Author: Jaeyeong Lee, Taeseong Yoon, Wonmo Koo, Heeyoung Kim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15799v1 Announce Type: new Abstract: Industrial processes often generate complex, interdependent time-series data from multiple sensors across multiple stages, forming complex dependencies among variables and process stages. Effective monitoring and timely anomaly detection of these time ...

📖 Read original article


41. QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides ​

Author: Zhengyang Zhuge, Hao Yu, Xin Wang, Zheng Li, Yizhong Cao, Dayiheng Liu, Jianwei Zhang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15810v1 Announce Type: new Abstract: Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP8. As an emerging low-precision format, NVFP4 combines fine-grained scal...

📖 Read original article


42. Graph Coloring Approach to Solving Sudoku with Oscillatory Neural Networks ​

Author: Filip Sabo, Aida Todri-Sanial
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15814v1 Announce Type: new Abstract: Oscillatory Neural Networks (ONNs) present an attractive physics-based computing paradigm rooted in the dynamics of a network of typically fully coupled oscillators aiming to minimize an underlying energy function. In this paper, we propose an ONN-base...

📖 Read original article


43. In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention ​

Author: Katsuyuki Hagiwara
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15819v1 Announce Type: new Abstract: In-context learning is a remarkable property of transformers and has recently received a lot of interest. In many studies of in-context learning, it has been shown that transformers are capable of implementing solver for linear and non-linear regressio...

📖 Read original article


44. Data-Native Global Optimization for Big Data K-means Clustering ​

Author: Ravil Mussabayev, Rustam Mussabayev, Zukhra Yerdaliyeva, Kuldeyev Nursultan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15835v1 Announce Type: new Abstract: Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids. We target arbitrarily tall da...

📖 Read original article


45. ContinuityBench: A Benchmark and Systems Study of Stateful Failover in Multi-Provider LLM Routing ​

Author: Vishal Pandey, Gopal Singh
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15899v1 Announce Type: new Abstract: In production large language model (LLM) deployments, high API availability guarantees do not equate to conversational continuity. When a primary provider experiences an outage or strict rate-limiting, naive stateless failover mechanisms successfully m...

📖 Read original article


46. A Semiparametric Framework for Stochastic Fundamental Diagram Modeling ​

Author: Pengnan Chi, Xiaoliang Ma, Magnus Jansson, Magnus Nordenvaad
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15907v1 Announce Type: new Abstract: The stochastic fundamental diagram (SFD) provides a probabilistic description of the relationship between traffic density and flow or speed, enabling uncertainty-aware traffic modeling. However, existing stochastic models frequently struggle to accommo...

📖 Read original article


47. (MPO)$^2$: Multivariate Polynomial Optimization based on Matrix Product Operators ​

Author: Niccol`o Ciolli, Anders Vestergaard N{\o}rskov, Michael Kastoryano, Petr Taborsky, Morten M{\o}rup
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15916v1 Announce Type: new Abstract: Central to machine learning and signal processing is the ability to perform universal function approximation and learn complex input-output relationships from limited numbers of observations. Multivariate polynomial models offer a natural way to expres...

📖 Read original article


48. On the Failure of Boundary-Seeking Distillation in Bottlenecked Generative Architectures ​

Author: Mohamed Amine Kina
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.15919v1 Announce Type: new Abstract: Data-free knowledge distillation transfers the knowledge encoded in a teacher model to a student model without access to the original training data. Prior work such as Contrastive Abductive Knowledge Extraction (CAKE) achieves this for classifiers by s...

📖 Read original article


49. Knowledge-Guided Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Label-Conditioned Contrastive Alignment ​

Author: Xinran Liu, Yuwen Li, Hongxiang Gao, Heyang Xu, Jianqing Li, Zongmin Wang, Chengyu Liu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15928v1 Announce Type: new Abstract: Adult and pediatric electrocardiogram (ECG) interpretation relies on age-sensitive criteria, and models pretrained mainly on adult ECGs often transfer poorly to pediatric populations when pediatric labels are scarce. Existing multimodal ECG--text metho...

📖 Read original article


50. An Exploratory Study of Single Channel Surface Electromyography for Hand Gesture Classification ​

Author: Daanish Hindustani
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.15972v1 Announce Type: new Abstract: Accurate hand gesture recognition using surface electromyography (sEMG) typically relies on multichannel sensor arrays and computationally intensive models. This limits practical deployment in low-power and embedded systems. This study investigates the...

📖 Read original article


51. DebrisTracer: Reliable Tracking in Hypervelocity Impact Fast Imaging ​

Author: Th'eophane Loloum, Fabien Vivodtzev, David H'ebert, Baptiste Reynier, Michel Arrigoni, Julien Tierny
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.CV, cs.GR, eess.IV

arXiv:2607.15986v1 Announce Type: new Abstract: This application paper presents DebrisTracer, a framework for the reliable tracking of debris in hypervelocity impact fast imaging. These noisy and highly specific datasets capture the ejection of a large number of debris fragments after the impact of ...

📖 Read original article


52. Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation ​

Author: Zhaoyang Jiang, Zhizhong Fu, Zicheng Li, Yunsoo Kim, Jiacong Mi, Xuanqi Peng, Fei Teng, Honghan Wu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.16019v1 Announce Type: new Abstract: AI systems increasingly retrieve from records that revise themselves: issue threads, encyclopedic histories, policy logs, and long conversations. The challenge is not only finding relevant evidence, but deciding which claims remain in force, which were...

📖 Read original article


53. Constrained Hebbian Learning Supports Efficient Representational Allocation under Structural Constraints ​

Author: Patrick Inoue, Florian R"ohrbein, Andreas Knoblauch
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.NE

arXiv:2607.16027v1 Announce Type: new Abstract: Introduction: Biological systems face anatomical and metabolic constraints, including costly synaptic maintenance and limited connectivity. These constraints favor neural codes that compress behaviorally relevant information into low-redundancy pattern...

📖 Read original article


54. CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward Approach ​

Author: Andrei Neagu, Eeham Khan, Leila Kosseim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.16028v1 Announce Type: new Abstract: This paper presents our system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires daily long, flat, or short trading decisions for Bitcoin (BTC) and Tesla (TSLA) using news and historical market data. We formulate the problem as a discrete-actio...

📖 Read original article


55. Revisiting data-driven dynamic security assessment with a tabular foundation model ​

Author: Olayiwola Arowolo, Maosheng Yang, Jochen Cremer
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.16031v1 Announce Type: new Abstract: Data-driven pre-fault dynamic security assessment (DSA) rapidly evaluates the dynamic risk of credible contingencies on a power system using machine learning. Existing approaches face two limitations. First, they require a large labelled database for t...

📖 Read original article


56. DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model Embeddings ​

Author: Yuya Kawakami, Daniel Cayan, Dongyu Liu, Kwan-Liu Ma, Tom Corringham
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.16050v1 Announce Type: new Abstract: Pluvial (rainfall-driven) flooding accounts for 45% of National Flood Insurance Program (NFIP) claims in the United States and is harder to predict than its riverine and coastal counterparts, with existing approaches limited to coarse resolution, regio...

📖 Read original article


57. When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis ​

Author: S. Aaron McClendon
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.16062v1 Announce Type: new Abstract: Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is essentially never tested against the baseline it claims to replace: methods merge independently released agents prec...

📖 Read original article


58. Physics-Based Deep Spatiotemporal Hyperlocal Radar Nowcasting with a Multi-Variable U-Net for High-Resolution Precipitation Forecasting ​

Author: Akshay Sunil, Muhammed Rashid, Raja Sekhar Sivaraju, Sushma Nair, Subimal Ghosh
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, eess.IV

arXiv:2607.16080v1 Announce Type: new Abstract: Precipitation nowcasting over the immediate 10-90 min period is important for flood management and real-time decision-making in urban regions. Conventional short-range forecasting with high-resolution numerical weather prediction requires frequent data...

📖 Read original article


59. Neural spectroscopy of AlphaFold2 reveals encoded protein conformational landscapes ​

Author: Kaustav Mehta
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, q-bio.BM

arXiv:2607.16087v1 Announce Type: new Abstract: AlphaFold2's 93 million parameters, shaped by the evolutionary record of protein structure encoded in the Protein Data Bank and in sequence alignments, are conventionally treated only as machinery for converting sequence to structure. We propose they a...

📖 Read original article


60. DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning ​

Author: Hanyang Chen, Anirudh Satheesh, Longchao Da, Hua Wei
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.16090v1 Announce Type: new Abstract: Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains. In this paper, we consider the setting of online dynamics adaptation, where policies are trained...

📖 Read original article


61. Understanding Reasoning from Pretraining to Post-Training ​

Author: Jingyan Shen, Ang Li, Salman Rahman, Yifan Sun, Micah Goldblum, Matus Telgarsky, Pavel Izmailov
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2607.16097v1 Announce Type: new Abstract: Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain ope...

📖 Read original article


62. The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic Infrastructure ​

Author: Jun He, Deying Yu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.DC, cs.MA

arXiv:2607.16109v1 Announce Type: new Abstract: State machine replication (SMR) and Byzantine fault-tolerant (BFT) consensus guarantee agreement despite a bounded number of arbitrary, colluding faulty participants. However, these guarantees rely on participants outside this set correctly executing t...

📖 Read original article


63. When Do Multi-Agent Systems Help? An Information Bottleneck Perspective ​

Author: Wendi Yu, Lianhao Zhou, Xiangjue Dong, Sai Sudarshan Barath, Declan Staunton, Byung-Jun Yoon, Xiaoning Qian, James Caverlee, Shuiwang Ji
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.16133v1 Announce Type: new Abstract: LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with performance varying inconsistently across settings. Here, we provide an informa...

📖 Read original article


64. Improving Improved Kernel PLS ​

Author: Ole-Christian Galbo Engstr{\o}m
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.DS

arXiv:2607.16138v1 Announce Type: new Abstract: Improved Kernel Partial Least Squares (IKPLS) algorithms 1 and 2 are among the fastest PLS calibration algorithms. This article focuses on two shared steps, the computation of the $\mathbf{X}$ rotations, $\mathbf{R}$, and the $\mathbf{Y}$ loadings, $\m...

📖 Read original article


65. PRISA: Proactive Infrastructure LiDAR Framework for Intersection Safety Assessment ​

Author: Tam Bang, Hussam Abubakr, Emiliano de la Garza Villarreal, Truc Phuong Nguyen, Austin Harris, Toru Hirano, Mina Sartipi, Yunfei Xu, Hoang H. Nguyen
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.SY, eess.SY

arXiv:2607.16156v1 Announce Type: new Abstract: Urban intersections are among the most hazardous locations in road networks, posing significant risks to vehicles and vulnerable road users (VRUs) such as pedestrians and cyclists. The complexity of multi-agent interactions demands continuous, real-tim...

📖 Read original article


66. Behaviour-Conditioned Neural Processes for Adaptive Residential Short-Term Load Forecasting ​

Author: Ramin Soleimani, Andrea Visentin, Dirk Pesch
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.16168v1 Announce Type: new Abstract: Residential short-term load forecasting (STLF) is challenging because household demand is heterogeneous, temporally variable, and shaped by diverse behavioural routines. This work investigates whether inferred behavioural structure can be embedded with...

📖 Read original article


67. When Does Muon Help Agentic Reinforcement Learning? ​

Author: Kai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.16169v1 Announce Type: new Abstract: Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW on ALFWorld us...

📖 Read original article


68. Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems ​

Author: Matteo Tomasetto, Nicol`o Botteghi, Gabriele Bruni, Andrea Manzoni
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, math.OC

arXiv:2607.16177v1 Announce Type: new Abstract: Reinforcement learning (RL) has recently emerged as a promising feedback control strategy for nonlinear and complex dynamical systems. However, RL algorithms are sample inefficient and require a large number of interaction with the environment to synth...

📖 Read original article


69. A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing ​

Author: Owen Lockwood, J'er'emy B'ejanin, Joost Bus, Christopher Chamberland, Patrick Huembeli, Frank Sch"afer, Guillaume Verdon
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.ET, physics.app-ph

arXiv:2607.16183v1 Announce Type: new Abstract: To address the escalating energy and latency demands of machine-learning workloads, we introduce a blueprint for an energy-efficient and fast thermodynamic computing stack that leverages stochastic analog processes in physical hardware. In this work, w...

📖 Read original article


70. PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization ​

Author: Yuchen Yang, Yifan Zhao, Anisha Dasgupta, Sasa Misailovic
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.16184v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) is a popular class of large language models (LLMs), offering high efficiency and accuracy. However, in KV-cache-intensive serving scenarios, MoEs often exhibit a tension between the GPU memory requirements of the model weights ...

📖 Read original article


71. Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal ​

Author: Kia-J"ung Yang, Dominik Meier, Jiachen Zhao, Terry Ruas, Bela Gipp
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2605.26772v1 Announce Type: cross Abstract: Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may complicate control mechanisms such as refusal. Unlike instruction-tuned LLMs, where refusal is mediated...

📖 Read original article


72. Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior ​

Author: Ali khalil, Aly M. Kassem, Mohamed Abdelrazek, Santu Rana, Negar Rostamzadeh, Golnoosh Farnadi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.CL, cs.LG

arXiv:2607.15286v1 Announce Type: cross Abstract: We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using an emergent-misalignment organism and a refusal-ablated jailbroken org...

📖 Read original article


73. An Empirical Study of Handcrafted Feature Learning and Convolutional Neural Networks for Facial Expression Recognition ​

Author: Chethiya Galkaduwa
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.15288v1 Announce Type: cross Abstract: Facial expression recognition is an important computer vision task with applications in human--computer interaction, mental health monitoring, driver alert systems, and behavioral analysis. While convolutional neural networks (CNNs) dominate modern f...

📖 Read original article


74. A Physics-Informed Neural Network with a Modified Lorentzian Activation for Nonlocal Gradient-Flow Equations in Dynamic Density Functional Theory ​

Author: Dimitrios Gourzoulidis, Soumaya Elkantassi, Serafim Kalliadasis
Published: 7/20/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA

arXiv:2607.15291v1 Announce Type: cross Abstract: We develop a physics-informed neural network (PINN) framework for nonlocal partial differential equations arising in dynamic density functional theory (DDFT). Such equations are challenging for standard PINN methods because they involve nonlinearitie...

📖 Read original article


75. AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning ​

Author: Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin
Published: 7/20/2026, 4:00:00 AM
Categories: cs.MM, cs.AI, cs.LG, cs.SD

arXiv:2607.15295v1 Announce Type: cross Abstract: We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of global and per-modality ...

📖 Read original article


76. MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation ​

Author: Zhiyuan Zhao, Bin Wang, Linke Ouyang, Yiqi Lin, Pan Zhang, Xiaoyi Dong, Jiaqi Wang, Conghui He
Published: 7/20/2026, 4:00:00 AM
Categories: cs.MM, cs.CV, cs.LG

arXiv:2607.15299v1 Announce Type: cross Abstract: In this paper, we propose MLLM-DataEngine, a novel closed-loop system that bridges data generation, model training, and evaluation. Within each loop iteration, the MLLM-DataEngine first analyzes the weakness of the model based on the evaluation resul...

📖 Read original article


77. DyneTrion: A Spatio-temporally Coherent Generative Emulator for Protein Dynamics Across Timescales ​

Author: Kaihui Cheng, Zhiqiang Cai, Peng Tu, Yisong Yao, Limei Han, Libo Wu, Siyu Zhu, Tzuhsiung Yang, Yuan Qi
Published: 7/20/2026, 4:00:00 AM
Categories: q-bio.QM, cs.LG

arXiv:2607.15309v1 Announce Type: cross Abstract: Proteins function through coordinated motion across multiple spatial and temporal scales, underpinning processes such as ligand binding, allostery, and catalysis. However, accessing long-timescale conformational change through molecular dynamics (MD)...

📖 Read original article


78. On the Impact of Entropy-based Features ​

Author: Iuri Mundstock, Abreu Quevedo, J'eferson Campos Nobre, Roben C. Lunardi, Thiago L. T. da Silveira, Bruno L. Dalmazo
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.LG

arXiv:2607.15379v1 Announce Type: cross Abstract: Network anomaly detection is increasingly challenging due to the growing diversity and variability of traffic patterns, which are not always well captured by traditional statistical features. In this work, we explore the use of entropy as an addition...

📖 Read original article


79. Improving Network Anomaly Detection via Choquet-Integral-Based Feature Aggregation ​

Author: Abreu Quevedo, Roger Immich, Giancarlo Lucca, Gra\c{c}aliz Dimuro, Bruno L. Dalmazo
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.LG

arXiv:2607.15389v1 Announce Type: cross Abstract: This work investigates a generalized Choquet-integral-based feature aggregation framework to improve anomaly detection in high-dimensional network traffic data. The approach combines adaptive weighting with incremental feature selection to address fe...

📖 Read original article


80. Unsupervised Keypoints for Real-Time Fall Detection: Comparative Analysis Under Real-world Conditions with Predictive Bandwidth Reduction ​

Author: Tasmiah Haque, Jacob Kosinski, Sumit Mohan, Srinjoy Das, Mohammad Abdullah Al-Mamun
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.15400v1 Announce Type: cross Abstract: Falls among older adults are a major safety challenge, but continuous monitoring is difficult to sustain. Video captures fall-related posture and motion, yet deployment is limited by privacy, computation, and bandwidth. Supervised pose estimation is ...

📖 Read original article


81. Closed-Loop Bayesian Bandit Encoder with GRAND Receiver for a Bursty Interference Channel ​

Author: Bhaskar Krishnamachari
Published: 7/20/2026, 4:00:00 AM
Categories: cs.IT, cs.LG, eess.SP, math.IT

arXiv:2607.15404v1 Announce Type: cross Abstract: Interleaving mitigates burst errors but introduces decoding delay and removes temporal error structure that a channel-aware decoder could exploit. We consider packet-level selection between a random linear code and the same code used with cross-codew...

📖 Read original article


82. Proactive Inpatient Bed Requests for Emergency Department Admissions ​

Author: QIan Cheng, Nilay Tanik Argon, Aniruddhan Ganesaraman, Serhan Ziya
Published: 7/20/2026, 4:00:00 AM
Categories: stat.AP, cs.LG, math.OC, stat.ML, stat.OT

arXiv:2607.15432v1 Announce Type: cross Abstract: Emergency department (ED) boarding occurs when admitted patients remain in the ED while awaiting inpatient beds. Boarding is a major driver of ED crowding and has been associated with poor patient outcomes. We propose a framework to help EDs reduce b...

📖 Read original article


83. Prediction-Only Distillation in Linear and Logistic Regression ​

Author: Hien Dang, Pratik Patil, Alessandro Rinaldo
Published: 7/20/2026, 4:00:00 AM
Categories: math.ST, cs.LG, stat.ML, stat.TH

arXiv:2607.15450v1 Announce Type: cross Abstract: Self-distillation (SD) is typically studied when the student is retrained on the teacher's original training inputs. In many practical deployments, however, the labeled training data are no longer available, and one has access only to the trained pre...

📖 Read original article


84. Design-Based Supervised Learning with Noisy Human Labels ​

Author: Robert Chew, Matthew R. Williams
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.AI, cs.LG, stat.AP

arXiv:2607.15455v1 Announce Type: cross Abstract: Researchers increasingly use automated classifiers to label unstructured data for statistical analysis. Existing rectification methods can correct errors in these automated labels using a probability-sampled audit set, but they usually treat the audi...

📖 Read original article


85. FLINT: Fingerprinting Federated Learning Architectures from 5G PHY-Layer Side Channels ​

Author: Md Nahid Hasan Shuvo, Mahmudul Hassan Ashik, Moinul Hossain
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2607.15469v1 Announce Type: cross Abstract: Federated Learning (FL) over 5G cellular networks protects raw data but remains vulnerable to side-channel leakage. Prior fingerprinting attacks assume packet-level network visibility, an assumption that does not hold at the 5G Physical (PHY) layer, ...

📖 Read original article


86. Ptolemy's Equant Equates to a Universal Dynamical Clock via Machine Learning ​

Author: Jingdong Zhang, Luan Yang, Murilo S. Baptista, Zefeng Zhang, Qunxi Zhu, Wei Lin, Celso Grebogi
Published: 7/20/2026, 4:00:00 AM
Categories: math.DS, cs.LG, physics.bio-ph

arXiv:2607.15472v1 Announce Type: cross Abstract: Oscillatory dynamics arise ubiquitously in nonlinear systems, yet identifying a physically interpretable phase and phase dynamics in nonlinear, high-dimensional oscillations remains a central unresolved problem. Here we establish the principle of a u...

📖 Read original article


87. Verbalizable Representations Form a Global Workspace in Language Models ​

Author: Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, Jack Lindsey
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.15495v1 Announce Type: cross Abstract: Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous function...

📖 Read original article


88. VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs ​

Author: Shahrzad Esmat, Dhawal Shah, Ali Jannesari
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2607.15498v1 Announce Type: cross Abstract: The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score importance from an observati...

📖 Read original article


89. Fast and Scalable Caputo Fractional Gradient Descent via Perturbation-Preserving Memory Compression ​

Author: Hwanseo Lee, Junseo Lee, Hyunju Kim
Published: 7/20/2026, 4:00:00 AM
Categories: math.OC, cs.LG, cs.NA, cs.NE, math.NA

arXiv:2607.15505v1 Announce Type: cross Abstract: Fractional gradient descent (FGD) incorporates long-range memory through Caputo-type operators and has been shown to improve stability in ill-conditioned and nonconvex optimization problems. Despite these advantages, its practical use remains limited...

📖 Read original article


90. LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4 ​

Author: Mobina Kashaniyan, Amirhossein Ghassemi, Nasser Mozayani
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.15509v1 Announce Type: cross Abstract: We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition. Each large language model independently ge...

📖 Read original article


91. Intentional Electromagnetic Interference Attacks on Facial Recognition ​

Author: Tyler Fitzsimmons, Adam Czajka
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.CR, cs.LG

arXiv:2607.15512v1 Announce Type: cross Abstract: Attacks on general computer vision algorithms are often relegated to the digital domain, with the optimization performed purely in the digital world and then translated to physical mediums for implementation. In the field of biometrics, including fac...

📖 Read original article


92. E3DGS: Unified Geometric-Photometric Equivariance for 3D Gaussian Splatting via Color-as-Geometry Embedding ​

Author: Chankyo Kim, Maani Ghaffari
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.15536v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) captures scenes by coupling explicit geometry (position, covariance) with view-dependent photometry (Spherical Harmonics). However, building $\mathrm{SE}(3)$-equivariant architectures on these primitives presents a fundam...

📖 Read original article


93. Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models ​

Author: Rakshanda Hassan Abhinandan, John Galeotti, Deva Ramanan, Gautam Rajendrakumar Gare
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV

arXiv:2607.15565v1 Announce Type: cross Abstract: Where should the question go in a vision-language model (VLM) prompt: before the image or after it? Intuition says before: knowing what is asked should tell the model where to look. Yet across visual question answering benchmarks, question-first prom...

📖 Read original article


94. Process Reward Informed Tree Rollout for Effective Multi-Turn RL ​

Author: Xintong Li, Sha Li, Yuwei Zhang, Changlong Yu, Rongmei Lin, Hongye Jin, Shuyi Guan, Xin Liu, Linwei Li, Qingyu Yin, Jingbo Shang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.15610v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple independently sampled complete trajectories for advantage estimation. In long-horizon agentic tasks, such a uniform ...

📖 Read original article


95. Retraining Seeks Stable Signals ​

Author: Moritz Hardt
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2607.15623v1 Announce Type: cross Abstract: Predictive models deployed at scale influence future data, a phenomenon called performativity. And there is always one way to cope: Train the model on new data, deploy it again, and repeat. This process, called retraining or repeated risk minimizatio...

📖 Read original article


96. Testing Distributions Against Bounded Distinguishers ​

Author: Mark Bun, Rathin Desai, Renato Ferreira Pinto Jr
Published: 7/20/2026, 4:00:00 AM
Categories: cs.DS, cs.CC, cs.LG

arXiv:2607.15645v1 Announce Type: cross Abstract: Motivated by the challenge of testing distributions over high-dimensional or continuous domains, we study distribution testing with respect to bounded classes of distinguishers. A representative task is to use samples from an unknown distribution $P$...

📖 Read original article


97. Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models ​

Author: Yingqian Cui, Wei Deng, Lantao Mei, Hang Li, Charu C. Aggarwal, Hui Liu, Yue Xing
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2607.15655v1 Announce Type: cross Abstract: Masked diffusion language models (DLMs) enable parallel text generation by iteratively refining masked tokens, offering a promising alternative to autoregressive decoding. Recent lookahead-based decoding methods improve the accuracy--efficiency trade...

📖 Read original article


98. Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents ​

Author: Halima Bouzidi, Mboutidem Ekemini Mkpong, Mohammad Abdullah Al Faruque
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.CV, cs.LG

arXiv:2607.15657v1 Announce Type: cross Abstract: Multimodal AI agents increasingly rely on persistent long-term memory to ground generation in past visual and textual episodes. We show that unconditional trust in visual data creates a critical vulnerability. We propose Lucid, a black-box adversaria...

📖 Read original article


99. SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models ​

Author: Jinwen Xin, Xixiang Lv
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SD, cs.CR, cs.LG

arXiv:2607.15697v1 Announce Type: cross Abstract: Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data. In security-sensitive applications such as voice interaction fo...

📖 Read original article


100. Hierarchical Specialised Ensembles for Classification of Zebrafish Phenotypes Using the Selected Image Recognition Methods ​

Author: Piotr S. Maci\k{a}g, Monika Maci\k{a}g, Magdalena Majdan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.15698v1 Announce Type: cross Abstract: We propose and evaluate three hierarchical ensemble setups for zebrafish phenotype classification from embryo images. In all setups, stage 1 uses a single four-class classifier to assign images to one of the exclusive phenotypes: Normal, Chorion, Dea...

📖 Read original article


101. A Statistical Formulation Gap for Nonlinear Multiscale Physics-Informed Learning ​

Author: Ronald Katende
Published: 7/20/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA, math.AP

arXiv:2607.15702v1 Announce Type: cross Abstract: We prove a finite-sample formulation gap for physics-informed learning of nonlinear multiscale elliptic equations. For a uniformly monotone divergence-form class with coefficients oscillating at scale $\epsilon$, we derive a finite-width, finite-samp...

📖 Read original article


102. Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization ​

Author: Yong Chu, Xun Zhou, Zenglin Xu, Hui Wang, Yue Yu
Published: 7/20/2026, 4:00:00 AM
Categories: eess.SP, cs.AI, cs.LG

arXiv:2607.15713v1 Announce Type: cross Abstract: Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing. Despite its importance, achieving precise localization across diverse environmen...

📖 Read original article


103. Natural Backdoor Attacks on Speech Recognition Models ​

Author: Jinwen Xin, Xixiang Lyu, Jing Ma
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.LG, cs.SD

arXiv:2607.15724v1 Announce Type: cross Abstract: With the rapid development of deep learning, its vulnerability has gradually emerged in recent years. This work focuses on backdoor attacks on speech recognition systems. We adopt sounds that are ordinary in nature or in our daily life as triggers fo...

📖 Read original article


104. Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling ​

Author: Bo-An Chang, Yu-Chih Chen
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.MM

arXiv:2607.15740v1 Announce Type: cross Abstract: As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI. Existing T2I evaluation metrics and multimodal judges often rely on ...

📖 Read original article


105. Aggregation of Statistical Evidence under Exchangeability ​

Author: Antonin Schrab, Rajen Shah, Arthur Gretton, Ilmun Kim
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ME, cs.LG, math.ST, stat.ML, stat.TH

arXiv:2607.15823v1 Announce Type: cross Abstract: We study aggregation of statistical evidence under unknown and potentially complex dependence using group-invariance. Building on permutation-based constructions that treat transformed datasets as exchangeable units, we aggregate evidence across stat...

📖 Read original article


106. Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment ​

Author: Haowei Hua
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2607.15829v1 Announce Type: cross Abstract: Automated essay scoring (AES) enables scalable assessment and timely feedback but remains challenged by transformer input-length limitations, which can cause information loss when processing long essays. This study proposes a generative AI-assisted s...

📖 Read original article


107. RTL-Sequencer: Towards Scalable RTL Timing Prediction with the Sequence-based Paradigm ​

Author: Ziyan Guo, Wenji Fang, Wenkai Li, Yuchao Wu, Shang Liu, Zhiyao Xie
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AR, cs.AI, cs.LG

arXiv:2607.15830v1 Announce Type: cross Abstract: Accurate timing prediction at the register-transfer level (RTL) is a longstanding challenge in design automation. Existing graph-based methods struggle with limited receptive fields, high complexity, and a lack of signal directionality. We present RT...

📖 Read original article


108. A zero-one law for one-shot system identification ​

Author: Nicolas Boull'e, Diana Halikias, Samuel E. Otto, Alex Townsend
Published: 7/20/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA, math.DS

arXiv:2607.15832v1 Announce Type: cross Abstract: Can a model be identified from one experiment? We study analytic systems that are linearly parameterized by a combination of prescribed dictionary terms, such as partial differential operators and dynamical systems. For a single input-response pair, ...

📖 Read original article


109. Von Mises-Fisher Mixture Model with Dynamic Shrinkage for Realistic Test-Time Transduction ​

Author: Jiazhen Huang, Zhiming Liu, Changhu Wang, Wei Ju, Ziyue Qiao, Xiao Luo
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.15851v1 Announce Type: cross Abstract: A range of methods aim to enhance the performance of vision-language models (VLMs) at test time. Among them, transduction has emerged as a promising paradigm due to its strong compatibility and efficiency. However, realistic evaluations often involve...

📖 Read original article


110. Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation ​

Author: Zhenduo Shang, Xiyao Liu, Bohan Li, Xudong Wang, Teng Ren, Lianqing Liu, Zhi Han
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.LG

arXiv:2607.15880v1 Announce Type: cross Abstract: Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization. The existing methods predominantly focus on imitation from in-domain tasks and consequ...

📖 Read original article


111. Which Hyperparameters Matter? A Game-Theoretic Framework for Interpretable Hyperparameter Sensitivity Analysis ​

Author: Nyi Nyi Aung, Heepeom Shin, Abigail Lawlor, Adrian Stein
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.CO

arXiv:2607.15884v1 Announce Type: cross Abstract: This work presents a game-theoretic framework for interpretable hyperparameter-objective interaction analysis rather than proposing a new optimization algorithm. In the proposed framework, Shapley Effects are employed for global sensitivity analysis,...

📖 Read original article


112. Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models ​

Author: Andy Catruna, Emilian Radoi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.15893v1 Announce Type: cross Abstract: While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is known about diffusion language models (DLMs), an emerging alternative that generates text by iterative denoising. In this work, we study how...

📖 Read original article


113. Orbis 2: A Hierarchical World Model for Driving ​

Author: Sudhanshu Mittal, Arian Mousakhan, Silvio Galesso, Karim Farid, Jonannes Dienert, Rajat Sahay, Thomas Brox
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, cs.RO

arXiv:2607.15898v1 Announce Type: cross Abstract: Current world models operate at a single level of abstraction, with most prioritizing perceptual fidelity while lacking the spatial reasoning and semantic understanding required for real-world downstream tasks. We present a hierarchical driving world...

📖 Read original article


114. Learning Reach-Avoid Task with Reinforcement Learning: Vectorized Simulation and Benchmark ​

Author: Jonas Weihing, Shahram Eivazi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.LG

arXiv:2607.15935v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) has a longstanding tradition in addressing the reach-avoid task problem, especially for controlling robotic arms. While this task serves as a baseline environment within the research community, the ability of DRL to ...

📖 Read original article


115. More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe ​

Author: Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.15942v1 Announce Type: cross Abstract: Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, of...

📖 Read original article


116. Code-Poisoning Property Inference Attacks ​

Author: Xukun Luan, Yuhui Gong, Gang Zhang, Zixuan Huang, Yuanguo Bi, Xuesong Li, Jinyan Liu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.LG

arXiv:2607.15970v1 Announce Type: cross Abstract: The flourishing code hosting platforms and coding agents enable even beginners with private data to build tailored Machine Learning (ML) models using available code quickly. The training data for ML models, often regarded as private property (e.g., c...

📖 Read original article


117. CanonicalPhys: Pose-Robust Remote Photoplethysmography via Canonical-Space Priors ​

Author: Hui Wei, Seyedata Jodeiri Seyedian, Xiaobai Li, Guoying Zhao
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2607.15995v1 Announce Type: cross Abstract: Deep remote photoplethysmography (rPPG) attains sub-bpm heart-rate error on frontal, stationary faces yet degrades sharply under head pose: on MMPD, the state-of-the-art FactorizePhys backbone's MAE grows $1.60\times$ from frontal ($|\text{yaw}|{<}15...

📖 Read original article


118. Rethinking Quantum Continual Learning with Quantum Fisher Information ​

Author: Yu-Chao Hsu, Yu-Cheng Lin, Tai-Yue Li, Nan-Yow Chen, En-Jui Kuo
Published: 7/20/2026, 4:00:00 AM
Categories: quant-ph, cs.AI, cs.LG

arXiv:2607.16030v1 Announce Type: cross Abstract: Quantum continual learning aims to train quantum models on sequential tasks without losing previously learned knowledge. However, variational quantum classifiers (VQCs) are prone to catastrophic forgetting under nonstationary task distributions. We p...

📖 Read original article


119. Deep and Probabilistic Models for Gene Regulatory Network Inference ​

Author: Claudia Skok Gibbs
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.AP, stat.ME

arXiv:2607.16053v1 Announce Type: cross Abstract: Gene regulatory networks (GRNs) link transcription factor (TF) proteins to their target genes, yet reconstructing these networks from genome-wide data remains challenging under practical and methodological constraints. Many methods couple modeling as...

📖 Read original article


120. Pick-to-Learn Calibration of an MPC Policy for an Origin-to-Destination Flight Problem ​

Author: Marco C. Campi, Simone Garatti
Published: 7/20/2026, 4:00:00 AM
Categories: eess.SY, cs.LG, cs.SY, math.OC

arXiv:2607.16084v1 Announce Type: cross Abstract: This paper illustrates the Pick-to-Learn methodology applied to the calibration of a Model Predictive Control policy. While developed around a specific example, the presentation is meant to highlight a methodology of broad applicability. The example ...

📖 Read original article


121. CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data ​

Author: Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru, MohammadHossein Rezaei, Aakash Sabharwal, Yunzhong He
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2607.16122v1 Announce Type: cross Abstract: Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak examples, topics, or...

📖 Read original article


122. Learning Standard Model structure from LHC data with Riemannian flow matching ​

Author: Midori Kato, Kevin A. Urqu'ia-Calder'on, Inar Timiryasov, Oleg Ruchayskiy
Published: 7/20/2026, 4:00:00 AM
Categories: hep-ph, cs.LG, hep-ex

arXiv:2607.16144v1 Announce Type: cross Abstract: In this work we demonstrate that a single transformer-based generative model can capture Standard Model structure spanning five decades of invariant mass, from the sub-GeV regime to the TeV continuum, a range that no single Monte Carlo sample covers....

📖 Read original article


123. An Exam for Active Observers ​

Author: Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.CL, cs.LG

arXiv:2607.16165v1 Announce Type: cross Abstract: Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. W...

📖 Read original article


124. Cluster-Aware Matching via Laplacian Optimal Transport ​

Author: Gabriel Samberg, YoonHaeng Hur, Yuehaw Khoo, Nir Sharon
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, cs.NA, math.NA, stat.ME

arXiv:2607.16178v1 Announce Type: cross Abstract: In many applications of matching, the point clouds to be matched are not merely unstructured sets of points but rather samples from distributions with an intrinsic cluster structure. In such cases, as individual points are often interchangeable withi...

📖 Read original article


125. Perception-Aligned AI Outputs: End-to-End Visual Prediction for Uncertainty Communication in Clinical Decision-Making ​

Author: Mohammad Eslami, Solale Tabarestani, Saber Kazeminasab, Ehsan Adeli, Glyn Elwyn, Tobias Elze, Mengyu Wang, Nazlee Zebardast, Lucia Sobrin, Nassir Navab, Daniel Shu Wei Ting, Malek Adjouadi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2205.04599v2 Announce Type: replace Abstract: Explainable Artificial Intelligence (XAI) is essential for trustworthy AI in healthcare, yet many existing methods rely on technical explanations that are difficult for clinicians and patients to interpret. We introduce Visualized Learning for Mach...

📖 Read original article


126. AutoSpec: Automated Generation of Neural Network Specifications ​

Author: Shuowei Jin, Taobo Liao, Anuj Kalia, Xenofon Foukas, Huan Zhang, Cheng Tan, Z. Morley Mao, Francis Y. Yan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.SE

arXiv:2409.10897v3 Announce Type: replace Abstract: The increasing adoption of neural networks in learning-augmented systems highlights the growing need for model safety and robustness, especially in safety-critical domains. While recent advances in neural network verification offer formal guarantee...

📖 Read original article


127. Derivation of effective gradient flow equations and dynamical truncation of training data in Deep Learning ​

Author: Thomas Chen
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.AP, math.OC, stat.ML

arXiv:2501.07400v3 Announce Type: replace Abstract: We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient descent for the Euclidean loss in the input layer, and under the assumption that the weights are, in a precis...

📖 Read original article


128. CTC: The Composite Task Challenge for Cooperative Multi-Agent Reinforcement Learning ​

Author: Yurui Li, Yuxuan Chen, Xiaoli Yang, Shijian Li, Gang Pan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.MA

arXiv:2502.00345v2 Announce Type: replace Abstract: The critical role of division of labor (DOL) in enhancing cooperation is well-recognized in real-world applications. Consequently, many cooperative multi-agent reinforcement learning (MARL) methods have incorporated DOL mechanisms to improve cooper...

📖 Read original article


129. Rethinking the Global Knowledge of CLIP in Training-Free Open-Vocabulary Semantic Segmentation ​

Author: Jingyun Wang, Cilin Yan, Guoliang Kang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2502.06818v4 Announce Type: replace Abstract: Recent works modify CLIP to perform open-vocabulary semantic segmentation in a training-free manner (TF-OVSS). In vanilla CLIP, patch-wise image representations mainly encode homogeneous image-level properties, which hinders the application of CLIP...

📖 Read original article


130. MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation ​

Author: Haonan Yu, Junhao Liu, Xin Zhang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2502.11068v3 Announce Type: replace Abstract: Anchors is a popular local model-agnostic explanation technique whose applicability is limited by its computational inefficiency. To address this limitation, we propose a memorization-based framework that accelerates Anchors while preserving explan...

📖 Read original article


131. AuditVotes: Elevating Provable Defense for GNNs with Efficient Augmentation and Conditional Smoothing ​

Author: Yuni Lai, Yulin Zhu, Yixuan Sun, Yulun Wu, Bin Xiao, Gaolei Li, Jianhua Li, Qi Xie, Kai Zhou
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2503.22998v2 Announce Type: replace Abstract: Despite advancements in Graph Neural Networks (GNNs), adaptive attacks continue to challenge their robustness. Certified robustness via randomized smoothing offers provable guarantees but suffers from a severe accuracy-robustness trade-off, limitin...

📖 Read original article


132. Honesty in Causal Forests: When It Helps and When It Hurts ​

Author: Yanfang Hou, Carlos Fern'andez-Lor'ia
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2506.13107v5 Announce Type: replace Abstract: Causal forests estimate how treatment effects vary across individuals, guiding personalized interventions in areas like marketing, operations, and public policy. A standard practice is honest estimation: dividing the data into two samples, one to d...

📖 Read original article


133. A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model ​

Author: Andris Ambainis, Joao F. Doriguello, Debbie Lim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC, quant-ph, stat.ML

arXiv:2507.22854v3 Announce Type: replace Abstract: We propose novel classical and quantum online algorithms for learning finite- and infinite-horizon Markov Decision Processes (MDPs). Our algorithms are based on a hybrid online-offline reinforcement learning model wherein the agent can, from time t...

📖 Read original article


134. Discovering Generalizable Governing Equations for Graph Dynamical Systems with Interpretable Neural Networks ​

Author: Riccardo Cappi, Paolo Frazzetto, Nicol`o Navarin, Alessandro Sperduti
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2508.18173v2 Announce Type: replace Abstract: The discovery of symbolic governing equations is a central goal in science; yet, it remains challenging particularly for graph dynamical systems, where the network topology further shapes the system behavior. While artificial intelligence offers po...

📖 Read original article


135. Multi-marginal temporal Schr\"odinger Bridge Matching from unpaired data ​

Author: Thomas Gravier, Thomas Boyer, Auguste Genovesio
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2510.01894v3 Announce Type: replace Abstract: Many natural dynamic processes -- such as in vivo cellular differentiation or disease progression -- can only be observed through the lens of static sample snapshots. While challenging, reconstructing their temporal evolution to decipher underlying...

📖 Read original article


136. Are Heterogeneous Graph Neural Networks Truly Effective for Node Classification? A Causal Perspective ​

Author: Xiao Yang, Xuejiao Zhao, Zhiqi Shen
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2510.05750v2 Announce Type: replace Abstract: Graph neural networks (GNNs) have achieved remarkable success in node classification. Building on this progress, heterogeneous graph neural networks (HGNNs) integrate relation types and node and edge semantics to leverage heterogeneous information....

📖 Read original article


137. Mixing Configurations for Downstream Prediction ​

Author: Juntang Wang, Hao Wu, Yihan Wang, Dongmian Zou, Shixin Xu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2510.19248v2 Announce Type: replace Abstract: Clustering-based features are widely used in machine learning, but most methods must choose a resolution -- a choice that is global, fixed, and ad hoc. Recent work shows that varying the resolution parameter produces only a finite set of structural...

📖 Read original article


138. Analysis of Semi-Supervised Learning on Hypergraphs ​

Author: Adrien Weihs, Andrea L. Bertozzi, Matthew Thorpe
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, math.ST, stat.TH

arXiv:2510.25354v3 Announce Type: replace Abstract: Hypergraphs provide a natural framework for modeling multiway interactions. We analyze a class of variational semi-supervised learning problems posed on random geometric hypergraphs and establish asymptotic consistency in the large-data limit. In p...

📖 Read original article


139. DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone ​

Author: Vaibhav Singh, Oleksiy Ostapenko, Pierre-Andr'e No"el, Eugene Belilovsky, Torsten Scholak
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2511.15927v4 Announce Type: replace Abstract: Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency due to quadratic attention or KV-cache overhead. We introduce DiffuM...

📖 Read original article


140. SloMo-Fast: Slow-Momentum and Fast-Adaptive Teachers for Source-Free Continual Test-Time Adaptation ​

Author: Md Akil Raihan Iftee, Mir Sazzat Hossain, Rakibul Hasan Rajib, Tariq Iqbal, Md Mofijul Islam, M Ashraful Amin, Amin Ahsan Ali, AKM Mahbubur Rahman
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2511.18468v2 Announce Type: replace Abstract: Continual Test-Time Adaptation (CTTA) is crucial for deploying models in real-world applications with unseen, evolving target domains. Existing CTTA methods, however, often rely on source data or prototypes, limiting their applicability in privacy-...

📖 Read original article


141. Energy-Efficient Federated Learning via Adaptive Encoder Freezing for MRI-to-CT Conversion: A Green AI-Guided Research ​

Author: Ciro Benito Raggio, Lucia Migliorelli, Nils Skupien, Mathias Krohmer Zabaleta, Oliver Blanck, Francesco Cicone, Giuseppe Lucio Cascini, Paolo Zaffino, Maria Francesca Spadea
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV, cs.DC, physics.med-ph

arXiv:2512.03054v3 Announce Type: replace Abstract: Federated Learning (FL) holds the potential to advance equality in health by enabling diverse institutions to collaboratively train deep learning (DL) models, even with limited data. However, the significant resource requirements of FL often exclud...

📖 Read original article


142. Time-varying Mixing Matrix Design for Energy-efficient Decentralized Federated Learning ​

Author: Xusheng Zhang, Tuan Nguyen, Ting He
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.DC, math.OC

arXiv:2512.24069v2 Announce Type: replace Abstract: We consider the design of mixing matrices to minimize the operation cost for decentralized federated learning (DFL) in wireless networks, with focus on minimizing the maximum per-node energy consumption. As a critical hyperparameter for DFL, the mi...

📖 Read original article


143. Dichotomous Diffusion Policy Optimization ​

Author: Ruiming Liang, Yinan Zheng, Kexin Zheng, Tianyi Tan, Jianxiong Li, Liyuan Mao, Zhihao Wang, Guang Chen, Hangjun Ye, Jingjing Liu, Jinqiao Wang, Xianyuan Zhan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.RO

arXiv:2601.00898v3 Announce Type: replace Abstract: Diffusion-based policies have gained growing popularity in solving a wide range of decision-making tasks due to their superior expressiveness and controllable generation during inference. However, effectively training large diffusion policies using...

📖 Read original article


144. PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning ​

Author: Zhiyan Hou, Haiyun Guo, Haokai Ma, Yandu Sun, Yonghui Yang, Jinqiao Wang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2601.13020v2 Announce Type: replace Abstract: Continual instruction tuning (CIT) requires multimodal large language models (MLLMs) to adapt to a stream of tasks without forgetting prior capabilities. A common strategy is to isolate updates by routing inputs to different LoRA experts. However, ...

📖 Read original article


145. SC-JEPA: Stabilizing Latent Predictive Learning for Time-Series Anomaly Prediction ​

Author: Yanan He, Yunshi Wen, Xin Wang, Tengfei Ma
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2602.04643v2 Announce Type: replace Abstract: Time-series anomaly prediction aims to forecast future system failures before they fully emerge, making latent predictive models such as JEPA a promising framework for capturing precursor dynamics. However, directly applying continuous self-distill...

📖 Read original article


146. An Embarrassingly Simple Way to Optimize Orthogonal Matrices at Scale ​

Author: Adri'an Javaloy, Antonio Vergari
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, math.DG, math.OC

arXiv:2602.14656v2 Announce Type: replace Abstract: Orthogonality constraints are ubiquitous in robust and probabilistic machine learning. Unfortunately, current optimizers are computationally expensive and do not scale to problems with hundreds or thousands of constraints. One notable exception is ...

📖 Read original article


147. Coverage Guarantees for Pseudo-Calibrated Conformal Prediction under Distribution Shift ​

Author: Farbod Siahkali, Ashwin Verma, Vijay Gupta
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, eess.IV

arXiv:2602.14913v3 Announce Type: replace Abstract: Conformal prediction (CP) offers distribution-free marginal coverage guarantees under an exchangeability assumption, but these guarantees can fail if the data distribution shifts. We analyze the use of pseudo-calibration as a tool to counter this p...

📖 Read original article


148. SODA: Semi On-Policy Black-Box Distillation for Large Language Models ​

Author: Xiwen Chen, Jingjing Wang, Wenhui Zhu, Peijie Qiu, Xuanzhao Dong, Yueyue Deng, Hejian Sang, Zhipeng Wang, Alborz Geramifard, Feng Luo
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2604.03873v4 Announce Type: replace Abstract: Black-box knowledge distillation for large language models presents a strict trade-off. Simple off-policy methods (e.g., sequence-level knowledge distillation) struggle to correct the student's inherent errors. Fully on-policy methods (e.g., Genera...

📖 Read original article


149. Soft $Q(\lambda)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces ​

Author: Pranav Mahajan, Ben Seymour
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.13780v2 Announce Type: replace Abstract: Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a penalty on the divergence from a reference policy. Despite its success, the multi-step extensions o...

📖 Read original article


150. What Is the Minimum Architecture for Prolepsis? Early Irrevocable Commitment Across Tasks in Small Transformers ​

Author: 'Eric Jacopin
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CL

arXiv:2604.15010v2 Announce Type: replace Abstract: When do transformers commit to a decision, and what prevents them from correcting it? We introduce prolepsis: a transformer commits early, task-specific attention heads sustain the commitment, and no layer corrects it. Replicating Lindsey et al.'s ...

📖 Read original article


151. Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics ​

Author: Dominik Dahlem, Diego Maniloff, Mac Misiura
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.CL, stat.ML

arXiv:2605.04893v3 Announce Type: replace Abstract: Every attention head defines a degree-normalized transport operator, and a growing family of diagnostics reads model behavior (hallucination among them) from its spectrum. We ask what such diagnostics can and cannot infer. The operator splits ortho...

📖 Read original article


152. How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation ​

Author: Shai Feldman, Yaniv Romano
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.06605v3 Announce Type: replace Abstract: Evaluating and predicting the performance of large language models (LLMs) in multi-turn conversational settings is critical yet computationally expensive; key events -- e.g., jailbreaks or successful task completion by an agent -- often emerge only...

📖 Read original article


153. DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts ​

Author: Jiading Gai, Shuai Zhang, Xiang Song, Bernie Wang, George Karypis
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.15422v3 Announce Type: replace Abstract: Modern RL post-training methods such as GRPO and DAPO train on N response sequences of R tokens sampled from a shared prompt of P tokens, but standard FlashAttention replicates all P prompt tokens N times across both forward and backward passes -- ...

📖 Read original article


154. Variational Inference for Evidential Deep Learning ​

Author: Jiawei Tang, Xinyan Du, Hui Liu, Junhui Hou, Yuheng Jia
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.26477v3 Announce Type: replace Abstract: While Deep Neural Networks (DNNs) achieve remarkable performance, their tendency to produce overconfident predictions. Evidential Deep Learning (EDL) mitigates this by formulating predictions as a Dirichlet distribution over class probabilities to ...

📖 Read original article


155. Lightweight CNN-Based Anomaly Detection for High Voltage Converter Modulators in the Spallation Neutron Source ​

Author: Alberto D. Cencillo, Leonardo Concepci'on, Juli'an Luengo, Isaac Triguero
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.31259v2 Announce Type: replace Abstract: Unscheduled trips of high-power pulsed converters are a leading source of downtime at large accelerator facilities. At the Spallation Neutron Source (SNS), the High Voltage Converter Modulators (HVCMs) are consistently the second-largest contributo...

📖 Read original article


156. The Terminal Representation in Reinforcement Learning ​

Author: Amir Esterhuysen, Anders Jonsson
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2605.31289v2 Announce Type: replace Abstract: Representation learning is a powerful tool for spatio-temporal abstraction within reinforcement learning (RL). Two well established approaches are through the successor representation (SR) and the default representation (DR). The SR encodes states ...

📖 Read original article


157. Stop the Sampler! Classifier-Based Adaptive Stopping for Sampling Kernels ​

Author: Kirill Korolev, Nikita Morozov, Stepan Pavlenko, Esmeralda S. Whitammer, Sergey Samsonov
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2606.16073v2 Announce Type: replace Abstract: Sampling from complex, unnormalized probability densities is a fundamental challenge in Bayesian inference and probabilistic modeling. While Markov chain Monte Carlo (MCMC) methods provide asymptotic guarantees, they often suffer from slow mixing a...

📖 Read original article


158. Factorized Neural Operators Decompose Dynamic and Persistent Responses ​

Author: Hao Tang, Yuechen Duan, Jiongyu Zhu, Zimeng Feng, Hao Li, Chao Li
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2606.16900v2 Announce Type: replace Abstract: Physical systems often exhibit heterogeneous mechanisms, where rapidly evolving dynamics coexist with persistent structures. Capturing such multiscale physical behavior remains challenging for existing neural operators, which typically rely on sing...

📖 Read original article


159. GeoRouteNet: A Geometry-Aware Non-Autoregressive Neural Solver for the Euclidean Traveling Salesman Problem ​

Author: Xiang Li
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.22776v2 Announce Type: replace Abstract: Non-autoregressive neural solvers amortize computation across traveling salesman problem (TSP) instances, but models trained on random Euclidean instances can degrade when the number or spatial distribution of nodes changes. We study whether explic...

📖 Read original article


160. Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors ​

Author: Alexander H"agele, Alejandro Hern'andez-Cano, Atli Kosson, Martin Jaggi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2606.25971v2 Announce Type: replace Abstract: Modern neural network training relies on optimizers such as Adam and Muon which act on each weight matrix as a single object. Yet every weight matrix carries two distinct quantities -- a \emph{magnitude} and a \emph{direction} -- and all optimizers...

📖 Read original article


161. Anti-Collapse Dynamics and the Emergence of Multi-Time-Scale Learning in Recurrent Neural Networks ​

Author: Lorenzo Livi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, physics.data-an

arXiv:2606.29519v2 Announce Type: replace Abstract: Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data. This fade is c...

📖 Read original article


162. LLM-Guided Transportation Hub Capacity Planning with Textual Business Inputs ​

Author: Xiaoyue Liu, Zheng Dong
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, math.OC

arXiv:2607.03651v2 Announce Type: replace Abstract: While traditional hub capacity planning models optimize effectively for quantitative inputs, they often fail to digest qualitative business context. We propose a novel framework where a large language model (LLM) agent iteratively proposes hub capa...

📖 Read original article


163. Workload-Preserving Differentially Private Synthetic Data for Causal Inference via Maximum-Entropy Calibration ​

Author: Amir Asiaee, Kaveh Aryan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.08122v2 Announce Type: replace Abstract: Workload-based differentially private (DP) synthetic data methods privately measure aggregate queries and post-process the noisy answers into synthetic records. Generic workloads can achieve strong distributional fidelity, but causal estimands such...

📖 Read original article


164. Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph ​

Author: Duen Horng Chau, Donghao Ren, Fred Hohman, Dominik Moritz
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.DS, cs.HC

arXiv:2607.08746v2 Announce Type: replace Abstract: While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally. This graph encodes the data mani...

📖 Read original article


165. Energy-guided Recursive Model ​

Author: Yifei Zhao, Ying Tang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2607.10128v2 Announce Type: replace Abstract: Recursive reasoning models address structured problems by repeatedly updating latent states of small neural networks. However, their test-time scaling lacks a principled inference mechanism: increasing depth or stochastic breadth generates more tra...

📖 Read original article


166. A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs ​

Author: Rahul Krishnan, Volker Schulz
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.CL, math.OC

arXiv:2607.12550v2 Announce Type: replace Abstract: The key-value (KV) cache has become the dominant memory cost of transformer inference: it grows with batch size, context length, and depth, and at long context it, rather than the model weights, sets the throughput ceiling. Existing reductions fall...

📖 Read original article


167. Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems ​

Author: Dhruv Shivkant, Saket Mohanty, Somya Rai, Utkarsh Wadhwa
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.13735v2 Announce Type: replace Abstract: The rapid deployment of machine learning systems across cloud, edge, and enterprise environments has brought model optimization to the forefront of systems-engineering. Despite a rich literature spanning quantization, pruning, knowledge distillatio...

📖 Read original article


168. MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model ​

Author: Charilaos Papaioannou, Ioannis Tsantilas, Dimitris Giannakakos, Vasilis Michalakopoulos, Sotiris Pelekis, Vangelis Marinakis, Arsam Aryandoust, Antonello Monti, Ricardo J. Bessa, Perdo P. Vergara, Jochen Cremer, Elissaios Sarmas
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13763v2 Announce Type: replace Abstract: Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the lowest in-distribution error degrade the most under topology shift. We term this topology overfitting: the t...

📖 Read original article


169. Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values ​

Author: Jan Betley, Johannes Treutlein, Jan Dubi'nski, Harry Mayne, Karol Ga{\l}\k{a}zka, Niels Warncke, Anna Sztyber-Betley, Owain Evans
Published: 7/20/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CR

arXiv:2607.14345v2 Announce Type: replace Abstract: People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to th...

📖 Read original article


170. Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory ​

Author: Guangyue Peng, Tao Ge, Wen Luo, Wei Li, Houfeng Wang
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2303.01421v2 Announce Type: replace-cross Abstract: Semiparametric language models (LMs) have shown promise in various Natural Language Processing (NLP) tasks. However, they utilize non-parametric memory as static storage, which lacks learning capability and remains disconnected from the inter...

📖 Read original article


171. Instability in Complex Oscillator Networks: Limitations and Potentials of Network Measures and Machine Learning ​

Author: Christian Nauck, Michael Lindner, Nora Molkenthin, J"urgen Kurths, Eckehard Sch"oll, J"org Raisch, Frank Hellmann
Published: 7/20/2026, 4:00:00 AM
Categories: nlin.AO, cs.LG, cs.SI

arXiv:2402.17500v2 Announce Type: replace-cross Abstract: A central question of network science is how functional properties of systems emerge from their structure. For networked dynamical systems, structure is typically captured through network measures. We investigate the relationship between thes...

📖 Read original article


172. Pretrained Event Classification Model for High Energy Physics Analysis ​

Author: Joshua Ho, Benjamin Ryan Roberts, Shuo Han, Haichen Wang
Published: 7/20/2026, 4:00:00 AM
Categories: hep-ph, cs.LG

arXiv:2412.10665v3 Announce Type: replace-cross Abstract: We introduce a foundation model for event classification in high-energy physics, built on a Graph Neural Network architecture and trained on 120 million simulated proton-proton collision events spanning 12 distinct physics processes. The mode...

📖 Read original article


173. SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training ​

Author: Jiaheng Hu, Peter Stone, Roberto Mart'in-Mart'in
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG

arXiv:2506.04147v5 Announce Type: replace-cross Abstract: Building capable household and industrial robots requires mastering the control of versatile, high-degree-of-freedom (DoF) systems such as mobile manipulators. While reinforcement learning (RL) holds promise for autonomously acquiring robot c...

📖 Read original article


174. Minimax and Bayes Optimal Best-Arm Identification ​

Author: Masahiro Kato
Published: 7/20/2026, 4:00:00 AM
Categories: econ.EM, cs.LG, math.ST, stat.ME, stat.ML, stat.TH

arXiv:2506.24007v5 Announce Type: replace-cross Abstract: This study investigates minimax and Bayes optimal strategies for fixed-budget best-arm identification. We consider an adaptive procedure consisting of a sampling phase followed by a recommendation phase, and we design an adaptive experiment w...

📖 Read original article


175. Interpretable Role-Based Clustering in Multi-Layer Financial Networks ​

Author: Christian Franssen, Thao Le, Iman van Lelyveld, Bernd Heidergott
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SI, cs.LG

arXiv:2507.00600v3 Announce Type: replace-cross Abstract: Understanding the functional roles of financial institutions within interconnected markets is critical for effective supervision, systemic risk assessment, and resolution planning. We propose an interpretable role-based clustering approach fo...

📖 Read original article


176. Boosted Enhanced Quantile Regression Neural Networks with Spatiotemporal Permutation Entropy for Complex System Prognostics ​

Author: David J Poland
Published: 7/20/2026, 4:00:00 AM
Categories: eess.SP, cs.LG, cs.SY, eess.SY

arXiv:2507.14194v3 Announce Type: replace-cross Abstract: This paper presents an integrative prognostic framework that combines Spatiotemporal Permutation Entropy (STPE), Boosted Enhanced Quantile Regression Neural Networks (B-EQRNNs), Gated Temporal Attention, a Spiking Neural Network (SNN) refinem...

📖 Read original article


177. A Standardized Benchmark for Skeleton-Based Rehabilitation Assessment Using Deep Learning ​

Author: Ali Ismail-Fawaz, Maxime Devanne, Stefano Berretti, Jonathan Weber, Germain Forestier
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2507.21018v2 Announce Type: replace-cross Abstract: Automated assessment of human motion plays a vital role in rehabilitation, enabling objective evaluation of patient performance and progress. Unlike general human activity recognition, rehabilitation motion assessment focuses on analyzing the...

📖 Read original article


178. Label-Free Concept Drift Assessment for Reliable AI in Emerging Wireless Applications ​

Author: Athanasios Tziouvaras, Carolina Fortuna, George Floros, Kostas Kolomvatsos, Panagiotis Sarigiannidis, Marko Grobelnik, Bla\v{z} Bertalani\v{c}
Published: 7/20/2026, 4:00:00 AM
Categories: cs.NI, cs.LG

arXiv:2508.00042v2 Announce Type: replace-cross Abstract: Machine learning models deployed in non-stationary environments degrade silently, since as the input distribution drifts their accuracy decays without an error signal and without labels to reveal it. Sustaining reliable AI therefore requires ...

📖 Read original article


179. Comparative Field Deployment of Reinforcement Learning and Model Predictive Control for Residential HVAC ​

Author: Ozan Baris Mulayim, Elias N. Pergantis, Levi D. Reyes Premer, Bingqing Chen, Guannan Qu, Kevin J. Kircher, Mario Berg'es
Published: 7/20/2026, 4:00:00 AM
Categories: eess.SY, cs.LG, cs.SY

arXiv:2510.01475v2 Announce Type: replace-cross Abstract: Model Predictive Control (MPC) has demonstrated significant performance improvements over today's control methods for residential Heating, Ventilation, and Air Conditioning (HVAC), but deploying MPC often requires substantial engineering effo...

📖 Read original article


180. Manifold Dimension Estimation via Local Graph Structure ​

Author: Zelong Bi, Pierre Lafaye de Micheaux
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.AP

arXiv:2510.15141v5 Announce Type: replace-cross Abstract: Most existing manifold dimension estimators rely on the assumption that the underlying manifold is locally flat within the neighborhoods under consideration. More recently, curvature-adjusted principal component analysis (CA-PCA) has emerged ...

📖 Read original article


181. RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs ​

Author: Ruike Hu, Shulei Wu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2512.00319v3 Announce Type: replace-cross Abstract: The Structure Gap between probabilistic LLM generation and deterministic schema requirements hinders automated workflows. We propose RL-Struct, a lightweight framework using Gradient Regularized Policy Optimization (GRPO) with a hierarchical ...

📖 Read original article


182. Bifocal Attention: Harmonizing Geometric and Spectral Positional Embeddings for Algorithmic Generalization ​

Author: Kanishk Awadhiya
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.FL, cs.LG

arXiv:2601.22402v2 Announce Type: replace-cross Abstract: Rotary Positional Embeddings (RoPE) have become the standard for Large Language Models (LLMs) due to their ability to encode relative positions through geometric rotation. However, we identify a significant limitation we term ''Spectral Rigid...

📖 Read original article


183. Density-Informed Pseudo-Counts for Calibrated Evidential Deep Learning ​

Author: Pietro Carlotti, Nevena Gligi'c, Arya Farahi
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2602.01477v3 Announce Type: replace-cross Abstract: Evidential Deep Learning (EDL) is a popular framework for uncertainty-aware classification that models predictive uncertainty via Dirichlet distributions parameterized by neural networks. Despite its popularity, its theoretical foundations an...

📖 Read original article


184. Improving Backward Conformal Prediction via Non-Conformity Score Transformation ​

Author: Junxian Liu, Hao Zeng, Hongxin Wei
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2602.01733v3 Announce Type: replace-cross Abstract: Conformal Prediction (CP) provides a statistical framework for uncertainty quantification that constructs prediction sets with coverage guarantees. While CP yields uncontrolled prediction set sizes, Backward Conformal Prediction (BCP) inverts...

📖 Read original article


185. Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking ​

Author: Zhicheng Fang, Jingjie Zheng, Chenxu Fu, Wei Xu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.CL, cs.LG

arXiv:2602.24009v4 Announce Type: replace-cross Abstract: Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datasets, harnesses, and judging protocols. We introduce JAILBREAK F...

📖 Read original article


186. KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models ​

Author: Songming Zhang, Xue Zhang, Tong Zhang, Bojie Hu, Yufeng Chen, Jinan Xu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2603.01875v3 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is an essential technique to compress large language models (LLMs) into smaller ones. However, despite the distinct roles of the student model and the teacher model in KD, most existing frameworks still use a homog...

📖 Read original article


187. When Bigger is Worse: A Practitioner's Guide to Model Selection Under Data Scarcity ​

Author: Kwame Mbobda-Kuate, Gabriel Kasmi
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2603.02142v2 Announce Type: replace-cross Abstract: Scaling laws assume larger models trained on more data consistently outperform smaller ones -- an assumption that drives model selection in computer vision but remains untested in resource-constrained Earth observation (EO). We conduct a syst...

📖 Read original article


188. Conformal Graph Prediction with Z-Gromov-Wasserstein Distances ​

Author: Gabriel Melo, Thibaut de Saivre, Anna Calissano, Florence d'Alch'e-Buc
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2603.02460v5 Announce Type: replace-cross Abstract: Supervised graph prediction addresses regression problems where the outputs are structured graphs. Although several approaches exist for graph-valued prediction, principled uncertainty quantification remains limited. We propose a conformal pr...

📖 Read original article


189. LVSum: A Benchmark for Timestamp-Aware Long Video Summarization ​

Author: Alkesh Patel, Melis Ozyildirim, Ying-Chang Cheng, Ganesh Nagarajan
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2604.10024v2 Announce Type: replace-cross Abstract: Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally gro...

📖 Read original article


190. Robust Explanations for User Trust in Enterprise NLP Systems ​

Author: Guilin Zhang, Kai Zhao, Jeffrey Friedman, Xu Chu, Amine Anoun, Jerry Ting
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2604.12069v4 Announce Type: replace-cross Abstract: Robust explanations are increasingly required for user trust in enterprise NLP, yet pre-deployment validation is difficult in the common case of black-box deployment (API-only access) where representation-based explainers are infeasible and e...

📖 Read original article


191. VeriX-Anon: A Multi-Layered Framework for Mathematically Verifiable Outsourced Target-Driven Data Anonymization ​

Author: Miit Daga, Swarna Priya Ramu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CR, cs.DB, cs.LG

arXiv:2604.12431v2 Announce Type: replace-cross Abstract: Organisations increasingly outsource privacy-sensitive data transformations to cloud providers, yet no practical mechanism lets the data owner verify that the contracted algorithm was faithfully executed. VeriX-Anon is a multi-layered verific...

📖 Read original article


192. FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment ​

Author: Sophie Chiang, Tom Brennan, Fethiye Irmak Dogan, Jiaee Cheong, Hatice Gunes
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2604.23786v2 Announce Type: replace-cross Abstract: In recent years, the integration of multimodal machine learning in wellbeing assessment has offered transformative potential for monitoring mental health. However, with the rapid advancement of Vision-Language Models (VLMs), their deployment ...

📖 Read original article


193. Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory ​

Author: Bole Ma, Jan Eitzinger, Harald Koestler, Gerhard Wellein
Published: 7/20/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.LG

arXiv:2605.20982v2 Announce Type: replace-cross Abstract: AlltoAll dispatch is the dominant bottleneck of MoE expert parallelism, and the interconnect community has responded with four families of mitigations: predictive sample placement, adaptive expert relayout, hierarchical collectives, and EP-aw...

📖 Read original article


194. RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching ​

Author: Jinhyeok Yang, Hyeongju Kim, Yechan Yu, Joon Byun, Frederik Bous, Juheon Lee
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SD, cs.LG, eess.AS

arXiv:2605.22083v2 Announce Type: replace-cross Abstract: While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repeat errors from imperfect alignment. We propose RobustSpeechFlo...

📖 Read original article


195. RhinoVLA Technical Report ​

Author: Huixi Technology, :, Chen Zhang, Chenyang Zhou, Guanglei Ding, Guanghui He, Haibin Gao, Jiajia Chen, Jianyong Zhang, Lianyi Yu, Ningyi Xu, Ping Xu, Qingchen Li, Yingjun Hu, Yijia Zhang, Yuxi Liu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.LG

arXiv:2606.07383v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but real-time deployment on edge hardware remains challenging. In this work, we identify VLM visual and context tokens as a major source of deployment l...

📖 Read original article


196. AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models ​

Author: Andreas Einwiller, Max Klabunde, Florian Lemmerich
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2606.16127v2 Announce Type: replace-cross Abstract: The worldwide rise of authoritarianism and the growing role of Large Language Models (LLMs) in users' everyday lives raise the question of whether specific models exhibit or promote authoritarian attitudes. We introduce AuAu, a comprehensive ...

📖 Read original article


197. An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars ​

Author: Vinoth Nandakumar, Qiang Qu, Pramod Thebe, Sakshi Khachariya, Tongliang Liu
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2606.17522v2 Announce Type: replace-cross Abstract: Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract and compositional features across layers. In language modeling, ...

📖 Read original article


198. Where Will They Go? Modelling Multimodal Pedestrian Manoeuvres from Ego-centric Videos ​

Author: Yuxuan Xie, Nicolas Pugeault, Chongfeng Wei, Hubert P. H. Shum, Edmond S. L. Ho
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2606.18824v2 Announce Type: replace-cross Abstract: Pedestrian trajectory prediction from an on-board ego-centric camera is challenging since it depends on complex interactions with vehicles and scene context, as well as the intention of the pedestrian. The task becomes even more challenging s...

📖 Read original article


199. Dirac-Frenkel dynamics with inertia for nonlinearly parametrized solutions of evolution problems ​

Author: Matteo Raviola, Benjamin Peherstorfer
Published: 7/20/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA

arXiv:2606.24769v2 Announce Type: replace-cross Abstract: Even when Dirac-Frenkel dynamics determine a well-defined evolution in function space, the corresponding parameter dynamics can be non-unique or ill-conditioned for redundant nonlinear parametrizations, such as typical neural networks or mixt...

📖 Read original article


200. Missing Data Imputation under Manifold Hypothesis ​

Author: Zelong Bi, Amuchechukwu Ibenegbu
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2607.03641v2 Announce Type: replace-cross Abstract: The manifold hypothesis posits that high-dimensional data are concentrated near a low-dimensional embedded manifold. Recent advances in mixture variational autoencoders (VAEs) provide a powerful tool for extracting such underlying structure i...

📖 Read original article


201. Length Penalties Make Chain-of-Thought Less Monitorable ​

Author: Bryce Little
Published: 7/20/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.LG

arXiv:2607.09786v2 Announce Type: replace-cross Abstract: Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In our experiments, training with length penalties does not stop misleading hints from steering models, e...

📖 Read original article


202. NeuralActuator: Neural Actuation Modeling for Robot Dynamics and External Force Perception ​

Author: Zhiyang Dou, John U. Onyemelukwe, Hangxing Zhang, Heng Zhang, Minghao Guo, Yunsheng Tian, Michal Piotr Lipiec, Joshua Jacob, Chao Liu, Peter Yichen Chen, Yuri Ivanov, Wojciech Matusik
Published: 7/20/2026, 4:00:00 AM
Categories: cs.RO, cs.CV, cs.GR, cs.LG

arXiv:2607.11734v2 Announce Type: replace-cross Abstract: Differentiable simulators have advanced policy learning and model-based control across robotic tasks. Yet actuator dynamics remain underexplored and can be a major source of sim-to-real error, particularly on low-cost platforms, where the lin...

📖 Read original article


203. LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition ​

Author: Donghwan Kim
Published: 7/20/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2607.13347v2 Announce Type: replace-cross Abstract: LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a controlled tes...

📖 Read original article


204. Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code ​

Author: Niels M"undler-Sasahara, Hristo Venev, Dawn Song, Martin Vechev, Jingxuan He
Published: 7/20/2026, 4:00:00 AM
Categories: cs.PL, cs.AI, cs.LG

arXiv:2607.13921v2 Announce Type: replace-cross Abstract: Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generation more difficult. Off-the-shelf compilers can provide useful feedback post-generation, but does not gui...

📖 Read original article


205. NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs ​

Author: Jiarong Zhao, Zhikai Lei, Zhiheng Xi, Rui Zheng, Hang Yan, Jie Zhou, Qin Chen, Liang He
Published: 7/20/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.LG

arXiv:2607.14186v2 Announce Type: replace-cross Abstract: Synthesizing training data to scale agent capabilities in LLM post-training is bottlenecked by substrate-bound task synthesis: tasks are generated from fixed tools, repositories, or skill graphs, so expanding coverage requires manual substrat...

📖 Read original article


206. cGAP: Generalized Association Plots with HOMALS-Guided Heatmaps for Visualization of High-Dimensional Categorical Data ​

Author: Chun-houh Chen, Shun-Chuan Chang, Chiun-How Kao, Yi-Ju Lee, Shang-Ying Shiu, Yin-Jing Tien, ShengLi Tzeng, Han-Ming Wu
Published: 7/20/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, stat.CO, stat.ME

arXiv:2607.15018v2 Announce Type: replace-cross Abstract: High-dimensional categorical data arise in genetics, biomedicine, and the social sciences, yet visualization tools for such data remain far less developed than those for continuous variables. Existing methods either scale poorly, rely heavily...

📖 Read original article