Skip to content

arXiv cs.LG - 2026-08-10 ​

237 items collected.


1. Latent Fact-Checking: Detecting Misinformation through Activation Engineering ​

Author: Pedro Barcelos, Ot'avio Parraga, Marcelo M. Mussi, Lucas M. Fraga, Lucas S. Kupssinsk"u, Rodrigo C. Barros
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.06417v1 Announce Type: new Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level linguistic features or external knowledge retrieval, we examine truthfulness as a geometric property of a ...

📖 Read original article


2. Risk-Aware Decision Policies for Agents Under Noisy Perception ​

Author: David Szczecina
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06420v1 Announce Type: new Abstract: Perception in biological systems is inherently noisy, requiring organisms to make decisions under uncertainty where misclassification can be costly or fatal. We present an Artificial Life predator-prey model of foraging under noisy perception, and comp...

📖 Read original article


3. Sharding Prevents LLM Oversight Failures and Adversarial Exploitation ​

Author: Victor Akinwande, J. Zico Kolter, Aran Nayebi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06422v1 Announce Type: new Abstract: Giving an LLM judge more compute does not necessarily make it check more requirements. When one call must return many verdicts, some decisions become weakly grounded in the evidence, even when that call receives the same token or tool budget as a panel...

📖 Read original article


4. Adversarial Causal Intervention Falsification ​

Author: Mojtaba Eslami
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.GT, econ.EM, stat.ME

arXiv:2608.06427v1 Announce Type: new Abstract: Generative models can reproduce an observational distribution while encoding an incorrect causal structure. We study a sequential game in which a structural causal generator proposes observational and interventional distributions, while an adversarial ...

📖 Read original article


5. Fixed and Adaptive Topological DeepONets: Functional Measurements on Hausdorff Locally Convex Spaces ​

Author: Khemraj Shukla, George Em Karniadakis
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, math-ph, math.GN, math.MP

arXiv:2608.06428v1 Announce Type: new Abstract: Deep Operator Networks (DeepONets; arXiv:1910.03193) typically encode an input function through point values on a fixed discretization. Building on the Topological DeepONet framework of Ismailov (arXiv:2603.11972), we replace point samples by continuou...

📖 Read original article


6. MiGHT-EHR: A Multi-task Graph Transformer for Heterogeneous Temporal Electronic Health Records ​

Author: Anirudh Rayas, Yuan Wang, Pavan Turaga
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, q-bio.QM

arXiv:2608.06430v1 Announce Type: new Abstract: Learning from Electronic Health Records (EHRs) has gained significant attention due to its potential to improve clinical prediction. However, effective learning remains challenging because EHRs encode heterogeneous, temporally ordered clinical interact...

📖 Read original article


7. SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction ​

Author: Guofan Yu, Sitian Chen, Zhenheng Tang, Xiaowen Chu, Amelie Chi Zhou
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.DC

arXiv:2608.06441v1 Announce Type: new Abstract: Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges. We present SNI-GNN, a SmartNIC-assisted full-graph training system that reduces communication while preser...

📖 Read original article


8. ED-CSP: Crystal Structure Prediction from Electron Diffraction ​

Author: Germain Poloudenny, Ya"el Fr'egier, Arnaud Demorti`ere
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06448v1 Announce Type: new Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. Existing ED-based learning methods mainly predict crystallographic labels, reconstruct structures from...

📖 Read original article


9. Beyond Attention: Signed Integrated Gradients Attribution in a BiomeGPT-Style Microbiome Transformer ​

Author: Oren Nelson
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06486v1 Announce Type: new Abstract: In a feature-tokenized transformer (arXiv:2106.11959) such as BiomeGPT (doi:10.64898/2026.01.05.697599), each input token is built by fusing a fixed identity with a sample-specific measurement: a fixed species and a variable abundance, T = S + A. To in...

📖 Read original article


10. Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability ​

Author: Guanghui Min, Liang Wu, Mayank Darbari, Chen Chen, Liangjie Hong
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06503v1 Announce Type: new Abstract: Recurrent context compression controls context growth in long-horizon agents, but its behavioral effects remain poorly understood. In this preliminary empirical study, we show that compression can weaken the influence of recent interactions, increasing...

📖 Read original article


11. Unmasking Removal-Budget Confounding: A Matched Operating-Point Evaluation Framework for Adaptive Data Cleaning ​

Author: Wei-Hsiang Chen, Pin-Hsuan Yu, Chen-Hsuan Fang, Jung-Hua Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06511v1 Announce Type: new Abstract: Adaptive data-cleaning methods replace manual filtering thresholds with data-driven partitions. However, changing the partition granularity, the number of groups used to segment samples by estimated corruption risk, can implicitly shift the decision bo...

📖 Read original article


12. Target-Weighted Neyman Allocation: Experimental Design for Heterogeneous Treatment Effects under Population Shift ​

Author: Hoang Dang, Luan Pham, Minh Nguyen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, stat.ME

arXiv:2608.06512v1 Announce Type: new Abstract: Randomized experiments are often run in one population to guide decisions in another. Allocating by experimental proportions wastes budget on groups that rarely appear in deployment, whereas allocating by deployment proportions under-samples groups tha...

📖 Read original article


13. CertBind from Multimodal Connectivity to Certifiable Retrieval Decisions ​

Author: Shuheng Cao, Zhenhao Zhang, Ruiqi Chen, Renjie Cao, Weijia Zhang, Siyu Zhang, Jiaxin Liu, Xiangyu Zeng, Haotian Geng, Fan Gu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06516v1 Announce Type: new Abstract: Lightweight connectors make frozen multimodal encoders composable at the representation level. Deployment exposes a second problem at the level of task decisions. A connected route can expand cross-modal reach while changing an established native retri...

📖 Read original article


14. Online Security Learning in Cooperative Multi-Agent Systems under Hidden Byzantine Attacks ​

Author: Ximing Sun, Yue Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06520v1 Announce Type: new Abstract: We study online cooperative control of a multi-agent system under Byzantine attacks. Namely, an unknown, fixed subset of agents are Byzantine comprised and can stealthily overwrite its own coordinates of the team's planned joint action after observing ...

📖 Read original article


15. Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions ​

Author: Yuepeng Yang, Yuxin Chen, Yuejie Chi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, math.OC, stat.ML

arXiv:2608.06545v1 Announce Type: new Abstract: Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty. We study how many samples are necessary and sufficient to learn an $\varepsilon$-optimal robust policy under the av...

📖 Read original article


16. Newton-Schulz Retraction-Based Inference Enables Hidden Quantum Markov Models to Outperform Classical HMMs ​

Author: Ning Ning
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, quant-ph, stat.ME

arXiv:2608.06554v1 Announce Type: new Abstract: Hidden Markov models (HMMs) are widely used probabilistic models for discrete sequential data but can be limited when hidden dynamics are complex. Hidden quantum Markov models (HQMMs) generalize HMMs by replacing probability vectors with density matric...

📖 Read original article


17. Bootstrap-Conditioned Action Selection with Tabular Foundation Models ​

Author: Devansh Gupta, Shiv Tavker, Dmitry Efimov, Suchitra Sathyanarayana, Gitanjali Bhutani, Boris N. Oreshkin
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06559v1 Announce Type: new Abstract: Contextual bandits offer a natural framework for sample-efficient personalization, but practical deployment remains difficult under sparse, biased interaction data, unreliable uncertainty estimates, and severe cold starts. We study whether pre-trained ...

📖 Read original article


18. Theoretical Foundations of Communication-Efficient, Robust, and Practical Distributed and Federated Optimization ​

Author: Grigory Malinovsky
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, math.OC

arXiv:2608.06563v1 Announce Type: new Abstract: Machine learning and optimization have advanced together, with practical demands motivating new theory and theoretical breakthroughs enabling new applications. Modern large-scale training relies on classical optimization principles, but the constraints...

📖 Read original article


19. Quantization Damage Is Multiplicative, Not Additive ​

Author: Zekun Wu, Swati Dhiman, Adriano Koshiyama
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.06564v1 Announce Type: new Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt. What nobody can say is which of the model's decisions will change at a given bit-width. The damage is silent: a compressed agent stops calling its...

📖 Read original article


20. CrystalGRPO: Target-Aligned and Coverage-Preserving Reinforcement Learning for Flow-Based Crystal Structure Prediction ​

Author: Kaixiang Su, Hongfei Xue, Qiang Zhu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06582v1 Announce Type: new Abstract: Flow-based generative models can efficiently produce candidate structures for crystal structure prediction (CSP), but their pretrained objectives do not directly optimize downstream target recovery. Reinforcement-learning post-training offers a flexibl...

📖 Read original article


21. Flowing Through States: Neural ODE Regularization for Reinforcement Learning ​

Author: Mohamed Ghanem, Bernd Finkbeiner
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06595v1 Announce Type: new Abstract: Neural networks applied to sequential decision-making tasks typically rely on latent representations of environment states. While environment dynamics dictate how semantic states evolve, the corresponding latent transitions are usually left implicit, c...

📖 Read original article


22. Retrofitting Linear Attention into Diffusion Language Models ​

Author: Jinha Kim, Younghun Roh, Jaeyeon Kim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06628v1 Announce Type: new Abstract: Diffusion language models (dLLMs) offer a promising alternative to autoregressive models by accelerating inference through parallel decoding. Recent dLLMs commonly use blockwise semi-autoregressive decoding, generating blocks autoregressively while den...

📖 Read original article


23. The Sparsity Whisperer ​

Author: Linghao Kong, Inimai Subramanian, Micah Adler, Dan Alistarh, Dan Gutfreund, Nir Shavit
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that this overlooks a key computation performed by particularly sparsity-sensitive neurons in...

📖 Read original article


24. Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation ​

Author: Chunhui Shi, Xinwen Fu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06631v1 Announce Type: new Abstract: Cryptanalytic extraction has been demonstrated for ReLU networks, for networks using componentwise activations such as GELU or SiLU, and for a Transformer's final projection matrix. These methods do not recover the bias-free Gated Linear Unit (GLU) fee...

📖 Read original article


25. Bypassing Krum: Selection-Aware Backdoor Attacks in Federated Learning ​

Author: Srinivasan Subramanian, Md. Abdullah Al Hafiz Khan, Kazi Aminul Islam
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06637v1 Announce Type: new Abstract: Robust aggregation methods are widely used in federated learning to mitigate the impact of adversarial client behavior. Distance-based aggregation rules, such as Krum and Multi-Krum, select updates that are closest to the majority under the assumption ...

📖 Read original article


26. Dirichlet Follow-the-Leader Closes the Gap in Simultaneous Multiclass U-Calibration ​

Author: Pahan Dewasurendra
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06656v1 Announce Type: new Abstract: Can one forecaster attain the optimal regret rate for every bounded proper loss and also adapt to every smooth proper loss? Recent work answered this up to a dimension gap. Its self-concordant perturbation gives roughly $K^{5/4}\sqrt{T}$ worst-case reg...

📖 Read original article


27. EpiFlow: A framework for improving the utility of wastewater signals for disease forecasting ​

Author: Aniruddha Adiga, Jingyuan Chou, Gursharn Kaur, Andrew Warren, Srinivasan Venkatramanan, Baltazar Espinoza, Bryan Lewis, Justin Crow, Alexandra Lorentz, Rekha Singh, Madhav Marathe
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06671v1 Announce Type: new Abstract: Wastewater-based surveillance is an effective tool for disease monitoring and can provide early warning of outbreaks. Although wastewater viral loads (WVL) correlate with disease burden, their utility for improving real-time forecasting remains under i...

📖 Read original article


28. A Transferable Autologistic Model for Predicting Rare Failures in Heterogeneous Equipment ​

Author: Islam Benamirouche, Djemel Ziou, Feriel Fass
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06695v1 Announce Type: new Abstract: Predicting failures before they occur remains a major challenge in predictive maintenance, particularly when failures are rare, when equipment of the same family differ in sensor configurations, and when the goal is anticipation rather than diagnosis o...

📖 Read original article


29. Dueling World Models: Advantage-Style Action Channels for Common-Mode Distractor Rejection ​

Author: Jiazhuo Li, Yiming Fei, Zhiruo Zhou, Heikichi Hayashi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06706v1 Announce Type: new Abstract: Latent world models plan by predicting future states from an action, but when a scene contains motion the agent does not control, they quietly go action-blind: predictions for different actions become indistinguishable even as the training loss keeps i...

📖 Read original article


30. Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors ​

Author: Saeid Shokoufa, Mohammad Erfan Sadeghi, Mehdi Kamal, Massoud Pedram
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06723v1 Announce Type: new Abstract: The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making accurate estimation essential for sustainable artificial intelligence deployment and hardware-aware desi...

📖 Read original article


31. Solver-Guided Reasoning for Mixed-Equilibrium Strategies ​

Author: Han Wang, Philippe Beardsell, Boning Li, Aaron Sasmita, Shuai Li, Hongyuan Zha, Baoxiang Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.GT

arXiv:2608.06741v1 Announce Type: new Abstract: Reasoning in large language models (LLMs) is often grounded in human text, human demonstrations, and human-generated rationales. For equilibrium reasoning in complex games, however, relying on human data can be suboptimal. In fact, human play is often ...

📖 Read original article


32. KReF: Training-Free Retrieval for Long-Term Time-Series Forecasting and Predictive Uncertainty ​

Author: Yang Zhang, Rui Su
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06748v1 Announce Type: new Abstract: Probabilistic long-term time-series forecasting commonly relies on trained models. Training-free conformal methods typically construct intervals around a pre-existing point forecaster and do not natively represent a complete predictive distribution; se...

📖 Read original article


33. Sub-Quadratic Bisimulation Metrics via Approximate Nearest Neighbors: Coverage-Augmented Guarantees and Computable Two-Sided Certificates ​

Author: Ibne Farabi Shihab, Joyanta Jyoti Mondal
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06762v1 Announce Type: new Abstract: Bisimulation metrics quantify behavioral similarity in Markov decision processes, but their Wasserstein fixed-point operator updates every state pair and incurs quadratic pairwise work. We give a certificate-carrying sub-quadratic method for MDPs with ...

📖 Read original article


34. CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights ​

Author: Xuetian Gao
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.DC

arXiv:2608.06763v1 Announce Type: new Abstract: Weight quantization for large-language-model inference must balance adaptive reconstruction levels with representations regular enough for efficient GPU execution. Uniform integers constrain each group to a linear grid. Low-bit floating-point formats u...

📖 Read original article


35. Hidden Gauge Controls Feature Specialization in ReLU Networks ​

Author: Tongxi Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06766v1 Announce Type: new Abstract: Training changes a network's predictions while allocating task-relevant structure across its internal units. In an overparameterized ReLU network, several neurons can begin with exactly the same functional role, yet one may acquire a teacher feature wh...

📖 Read original article


36. ArchEGraph: A Large-Scale Graph Dataset for Geometry-Topology-Physics Aligned Building Energy Modeling ​

Author: Yihui Li, Yihui Chen, Kaidi Zha, Xiaoyue Yan, Zhexuan Yu, Shiqi Dai, Jun Xiao, Jun Yin, Ramon Elias Weber, Borong Lin
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06772v1 Announce Type: new Abstract: Accurate estimation of building energy use is essential for achieving carbon neutral and sustainable buildings. To better understand the influence of design decisions on building energy use and calibrate machine learning models that can give architects...

📖 Read original article


37. Faster Query-Key Learning Sharpens Attention in Self-Attention Models ​

Author: Rahul Vashisht, Harish G. Ramaswamy
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06776v1 Announce Type: new Abstract: A standard self-attention layer consists of two interacting circuits: the query-key circuit that governs attention allocation, and the output-value circuit that maps attended representations to predictions. Collapsed and factorized parameterizations of...

📖 Read original article


38. Understanding Differentiable Embeddings Through Differential and Integral Geometry ​

Author: Xinyu Zhang, Klaus Mueller
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06809v1 Announce Type: new Abstract: How can an analyst decide whether a nonlinear dimensionality reduction embedding can be trusted? Existing diagnostics provide only partial answers: projection glyphs characterize local sensitivity, map-continuity scores measure local conditioning, and ...

📖 Read original article


39. Multiscale Reward Hedging from Correct Demonstrations ​

Author: Pahan Dewasurendra
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06825v1 Announce Type: new Abstract: Learning from correct demonstrations is harder than supervised learning when many answers are correct: after predicting, the learner sees one valid answer but not whether its own answer was valid, nor any reward. Existing reward-hedging guarantees cons...

📖 Read original article


40. Graph Machine: Exploring Edge Mechanisms as an Inductive Bias ​

Author: Lintai Hou
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06834v1 Announce Type: new Abstract: Transformers provide a powerful architecture for global content-based matching, but reasoning problems may benefit from a stronger inductive bias toward iterative traversal of latent relations. We introduce Graph Machine, an architecture with two expli...

📖 Read original article


41. Mathematical Principles and Experimental Discoveries of the Emergence of Symbolic Patterns in Artificial Neural Networks ​

Author: Quanshi Zhang, Qihan Ren, Siyu Lou
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06839v1 Announce Type: new Abstract: Artificial Neural networks (ANNs) are often treated as black-box models, making explainability a central challenge in deep learning. Many engineering methods have been proposed to approximately explain the ANN from various perspectives, such as feature...

📖 Read original article


42. Bridging the Gap Between Hyperdimensional Computing and Kernel Methods via the Nystr\"om Method ​

Author: Quanling Zhao, Anthony Hitchcock Thomas, Ari Brin, Xiaofan Yu, Tajana Rosing
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06860v1 Announce Type: new Abstract: Hyperdimensional computing (HDC) is an approach from the cognitive science literature for solving information processing tasks using data represented as high-dimensional random vectors. The technique has a rigorous mathematical backing, and is easy to ...

📖 Read original article


43. SkillAligner: Treating Retrieved Skills as Adaptable Drafts at Execution Time ​

Author: Qinfeng Li, Dalin He, Yuntai Bao, Ying Yang, Ruoxi Chen, Xinyan Yu, Lizhou Liang, Ge Su, Wenqi Zhang, Xuhong Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06880v1 Announce Type: new Abstract: General-purpose skills promise reusable procedural knowledge for language agents, yet semantic relevance does not guarantee execution utility: a retrieved skill may encode assumptions that conflict with the current task, execution environment, or other...

📖 Read original article


44. PRISM: Principled Reference Identification for Schrodinger Bridge Model ​

Author: Forouzan Fallah, Yezhou Yang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06893v1 Announce Type: new Abstract: Schr"odinger bridge models restore a clean signal from a degraded observation by following the conditional bridges of a reference process, yet this reference is chosen heuristically, typically white noise with a hand-tuned schedule. We develop PRISM, ...

📖 Read original article


45. Recent advances in weakly supervised learning: New supervision paradigms, assumption relaxations, and practical solutions ​

Author: Wei Wang, Gang Niu, Masashi Sugiyama
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06896v1 Announce Type: new Abstract: Deep learning has achieved great success in recent years thanks to the availability of high-quality, well-annotated training data. However, this requirement is often not met in real-world applications. Weakly supervised learning aims to train an accura...

📖 Read original article


46. MiCoPro: End-to-End Mixed Precision HW/SW Co-design with HW-aware Proxy Model ​

Author: Zijun Jiang, Yangdi Lyu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06916v1 Announce Type: new Abstract: Quantized Neural Networks~(QNN) with low-bitwidth data have proven promising in efficient storage and computation on edge devices. To mitigate accuracy degradation while maximizing speedup, layer-wise mixed-precision quantization~(MPQ) becomes a popula...

📖 Read original article


47. Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning ​

Author: Moloud Damandeh, Meead Saberi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2608.06934v1 Announce Type: new Abstract: Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and preferences. Existing studies, however, often reduce these diverse judgements to aggregated scores, implicitl...

📖 Read original article


Author: Woojin Cho, Junghwan Park, Sangcheol Sim, Steve Andreas Immanuel, Junhyuk Heo, Darongsae Kwon
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2608.06942v1 Announce Type: new Abstract: The acquisition of multispectral imagery via small satellites (e.g., CubeSats) presents significant data downlink challenges due to high data volumes and restricted communication windows. While onboard image compression is critical to address this bott...

📖 Read original article


49. A Rate Separation for Agnostic Direct Sums ​

Author: Mihir More, Aritra Das, Debayan Gupta
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.06951v1 Announce Type: new Abstract: Hanneke, Moran, and Waknine \cite{HannekeMoranWaknine2024} asked how the agnostic PAC learning curve of the direct sum $C^r$ depends on the single-instance learning curve $\epsagn(n\mid C)$ and on $r$. We show that the single-instance learning rate doe...

📖 Read original article


50. How Molecular Generative Models Organize Molecular Identity ​

Author: Raul Ortega-Ochoa, Tejs Vegge, Jens S. Bakander, Luis Mantilla Calderon, Alan Aspuru-Guzik, Tonio Buonassisi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, physics.chem-ph

arXiv:2608.06956v1 Announce Type: new Abstract: Generative models for matter are often evaluated as samplers over output representations, and their latent spaces are commonly used as proxies for navigating chemical space. Much less is known about how these models internally arrange discrete chemical...

📖 Read original article


51. Density-aware Hierarchical Clustering Based on Element-Categorized Connection Subgraphs ​

Author: Yuning Yu, Jos'e Rodr'iguez-Pi~neiro, Xuefeng Yin, Bin Feng
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06990v1 Announce Type: new Abstract: Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning. Among various clustering methods, hierarchical clustering, density-based clustering, and graph clustering stand out as representative approaches. F...

📖 Read original article


52. Beyond Foundation Models: Dimension-Aware Neural Architecture Search with Small-Data Representation Models for Cryocooler Lifetime Prediction ​

Author: Gregor Molan (Comtrade 360 d.o.o., Letali\v{s}ka cesta 29b, Ljubljana, 1000, Slovenia), Grafika Jati (Comtrade 360 d.o.o., Letali\v{s}ka cesta 29b, Ljubljana, 1000, Slovenia), Francesco Barchi (Alma Mater Studiorum - Universita di Bologna, Department of Electrical, Electronic, and Information Engineering), Andrea Acquaviva (Alma Mater Studiorum - Universita di Bologna, Department of Electrical, Electronic, and Information Engineering), Alja\v{z} Osterman (LE-Tehnika d.o.o., \v{S}uceva 27, Kranj, 4000, Slovenia), Martin Molan (Comtrade AI GmbH, Grafenauweg 8, Zug, 6300, Switzerland)
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06993v1 Announce Type: new Abstract: Large-scale pretrained time-series models achieve strong results through large-scale pretraining and task-agnostic representation learning, but they rely on abundant, diverse data that industrial and scientific domains often lack. We therefore propose ...

📖 Read original article


53. Every Cache Entry Earns Its Place: Global Allocation of Resolution and Coverage for KV Cache Compression ​

Author: Haolin Tian, Yuzhe Liu, Tonghan Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07001v1 Announce Type: new Abstract: As large language models (LLMs) process increasingly long contexts, KV cache storage and repeated access have become a major bottleneck. Existing KV cache compression methods rely on predefined, fixed compression rules and are typically developed aroun...

📖 Read original article


Author: Robert Jankowski, Maksim Kitsak, Dorota Celi'nska-Kopczy'nska
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.SI, physics.soc-ph

arXiv:2608.07029v1 Announce Type: new Abstract: Hyperbolic embeddings provide compact geometric representations of complex networks in hyperbolic spaces, but systematic comparisons of methods developed in machine learning, network science, and algorithmics remain rare. We benchmark 13 unsupervised h...

📖 Read original article


55. Accounting Graph Transformer for Short-History Multi-KPI Forecasting in Small Businesses ​

Author: Shrutendra Harsola, Vignesh Subrahmaniam
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07037v1 Announce Type: new Abstract: Small businesses often have only 12-24 months of accounting history, yet planning and risk workflows require coordinated forecasts across financial statements. We study joint 12-month forecasting of 13 income-statement, balance-sheet, cash-flow, and wo...

📖 Read original article


56. Beyond Isolation: Unlocking Reinforcement Learning Component Synergy for Sample-Efficient Continuous Control ​

Author: Qi Zhao, Guozheng Ma, Yilun Kong, Lu Li, Haoyu Wang, Zilin Wang, Tiantian Zhang, Yuxing Wang, Jian Sha, Yongzhe Chang, Xueqian Wang, Dacheng Tao
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07086v1 Announce Type: new Abstract: Reinforcement learning systems are significantly more complex than other machine learning paradigms due to inherent properties, causing RL system design to jointly account for many tightly coupled factors. Despite advances in individual algorithmic com...

📖 Read original article


57. Synthetic LiDAR Data Generation and Deterministic Downsampling for Point Cloud Classification on the Edge ​

Author: Niclas Meyer, Stefan Reitmann
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.RO

arXiv:2608.07106v1 Announce Type: new Abstract: Deploying three-dimensional deep learning frameworks to low-power embedded processors is bottlenecked by the unstructured nature of spatial data and the resource-intensive distance sorting algorithms often used before neural network inference. To addre...

📖 Read original article


58. Modular TTT: Rethinking Test-Time Training as Composable Modules ​

Author: Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu, Ya Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.07110v1 Announce Type: new Abstract: Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard-code each variant separately, ...

📖 Read original article


59. Online Conformal Prediction Beyond Feedback ​

Author: Joar Skalse, Edoardo Pona, Osvaldo Simeone, Nicola Paoletti
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07139v1 Announce Type: new Abstract: Uncertainty quantification is essential when deploying machine learning models in safety-critical applications. Online conformal prediction (OCP) provides theoretically principled uncertainty quantification for arbitrary black-box classifiers and non-i...

📖 Read original article


60. Interpretable reinforcement learning with decision-tree pruning ​

Author: Mark Leon Ringer, Michel Tokic
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07151v1 Announce Type: new Abstract: Reinforcement learning policies are difficult to inspect, but interpreting them is a prerequisite for trustworthiness. Converting a trained policy into explicit decision-tree rules improves transparency and the resulting artifacts often remain too comp...

📖 Read original article


61. Machine Learning-Based Inter-Crystal Scatter Recovery for Ultra-High Resolution PET Imaging ​

Author: Alexandre Bernier, Roger Lecomte, Jean-Baptiste Michaud
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, physics.med-ph

arXiv:2608.07155v1 Announce Type: new Abstract: Inter-crystal scatter (ICS) events pose a significant challenge in ultrahigh- resolution positron emission tomography (UHR-PET), especially as detector crystals become smaller and their readouts increasingly segmented. Current approaches either reject ...

📖 Read original article


62. Capacity Confounds and Coverage Guarantees in Adaptive Sub-model Federated Learning ​

Author: Alireza Moayedikia, Alicia Troncoso Lora
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.DC

arXiv:2608.07157v1 Announce Type: new Abstract: Sub-model federated learning lets resource-constrained clients train width-reduced versions of a global model, but existing methods allocate capacity by device resources alone. A natural next step, allocating capacity by each client's data heterogeneit...

📖 Read original article


63. Edge Sparsification via Temporal Forman-Ricci Curvature for Dynamic Graph Learning ​

Author: Poupak Azad, Cuneyt Gurcan Akcora, Kiarash Shamsi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07158v1 Announce Type: new Abstract: Temporal graph learning has become essential for analyzing real-world systems whose interactions continuously evolve over time, including financial transaction networks, communication systems, and online social platforms. However, learning from large-s...

📖 Read original article


64. Fluid-DiT: Graph-Free Diffusion Transformers for Fluid Flow Simulations Learning ​

Author: Shentong Mo, Guolin Ke
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE

arXiv:2608.07161v1 Announce Type: new Abstract: Simulating complex fluid flows requires capturing full equilibrium distributions rather than just mean trajectories, yet high-fidelity solvers remain computationally prohibitive. Recent advances, such as Diffusion Graph Networks (DGNs), have combined d...

📖 Read original article


65. Momba: Network Modernization Improves Multi-Objective Reinforcement Learning ​

Author: Adam \v{S}tafa, Santeri Heiskanen, Petr Novotn'y, Joni Pajarinen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07180v1 Announce Type: new Abstract: Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in sample efficiency and asymptotic performance without altering the underlying algorithms. In contrast, work on mult...

📖 Read original article


66. Conformal Fusion Under Missing Modalities ​

Author: Alireza Moayedikia
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07183v1 Announce Type: new Abstract: Multimodal fusion architectures typically assume all modalities are available at inference, yet sensor failures, acquisition variability, and cost constraints routinely produce incomplete observations. Existing work treats modality absence as a predict...

📖 Read original article


67. MAUPITI: On-Device Prototype-Based Learning on a Smart Infrared Sensor ​

Author: Beatrice Alessandra Motetti, Tanguy Dugas du Villard, Matteo Risso, Alessio Burrello, Francesco Daghero, Enrico Macii, Massimo Poncino, Marco Castellano, Alfio Basile, Daniele Jahier Pagliari
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07192v1 Announce Type: new Abstract: Low-resolution infrared (IR) array sensors represent an interesting solution for privacy-preserving human sensing in embedded systems. In this letter, we describe a smart multi-pixel IR sensor integrating a 16$\times$16 thermal MOSFET (TMOS) array and ...

📖 Read original article


68. An AI4AI Framework for Visual Token Pruning ​

Author: Zhen Liu, Wenli Huang, Wei Song, Yuhan Liu, Zhiqin Yang, Jingwen Fu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2608.07193v1 Announce Type: new Abstract: Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fixed, handcrafted heuristics and costly expert trial and error. As pruning objectives, budgets, and mode...

📖 Read original article


69. Stochastic Autoregressive Learning ​

Author: Ilan Doron-Arad, Idan Mehalel, Elchanan Mossel
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07224v1 Announce Type: new Abstract: Motivated by LLMs, which generate outputs by iteratively sampling from next-token distributions, we introduce a PAC-learning model for binary stochastic autoregressive learning. This generalizes the deterministic autoregressive learning framework of Jo...

📖 Read original article


70. Learning Suffers More Than the Policy Class Under Partial Observability: A Closed-Form Analysis ​

Author: Idil G"ozel (University College London)
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, math.OC

arXiv:2608.07228v1 Announce Type: new Abstract: When a reinforcement learning agent cannot observe the full state, we usually blame its policies: it cannot see enough to represent a good one. We show that in a solvable case the bigger problem lies elsewhere. Even when a good policy is available and ...

📖 Read original article


71. TOFD: Target-Oriented Feature Decoupling against Poisoning Attacks in Split Federated Learning ​

Author: Yuhan Xie, Jingrong Huang, Chen Lyu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07274v1 Announce Type: new Abstract: Split Federated Learning (SFL) facilitates privacy-preserving collaborative training with reduced client-side overhead. However, its split architecture introduces unique attack surfaces, rendering it vulnerable to diverse poisoning attacks. Most existi...

📖 Read original article


72. A foundation-model approach to pediatric headache classification from rs-fMRI ​

Author: Guilherme S. Imai Aldeia, Clara Moon, Julie Shulman, Navil Sethna, Allison Smith, Alyssa Lebel, William G. La Cava, Scott Holmes
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07287v1 Announce Type: new Abstract: Headache is the most common neurological disorder in children and substantially affects quality of life. We investigated whether resting-state functional MRI (rs-fMRI) can support pediatric headache classification using machine learning. We encoded rs-...

📖 Read original article


73. FUSE: Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabular Flow Matching ​

Author: Suman Cha, Seongchan Lee, Dohyun Ko, Hyunjoong Kim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07294v1 Announce Type: new Abstract: Generating mixed-type tabular data requires jointly modeling diverse feature distributions and their complex cross-column dependencies. Variational flow matching handles distinct endpoints via factorized distributions, yet leaves feature-specific proce...

📖 Read original article


74. From Optimal Actions to World Models: Identifiability of Transition Kernels in Discounted MDPs ​

Author: Neal Batra
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07301v1 Announce Type: new Abstract: We study what can be recovered about the transition probabilities of a Markov decision process from optimal actions alone. This is closely related to the inverse problem considered by Letcher et al., who ask when the dynamics can be recovered from nume...

📖 Read original article


75. Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU ​

Author: Yuting Ge, Pengju Yang, Mingkai Nie
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07323v1 Announce Type: new Abstract: We test whether decoder-only language-model FFNs require SwiGLU's open positive tail. We introduce MemGLU as a closed-tail comparator derived from a memristive branch geometry. Across paired 9M and 30M pretraining runs with three seeds, MemGLU remains ...

📖 Read original article


76. When GNNs Fail: Quantifying and Overcoming Temporal Correlation Volatility in Time Series ​

Author: Chen Shao, Yue Wang, Zhenyi Zhu, Zhanbo Huang, Tobias K"afer, Zonghan Wu, Danai Koutra
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07333v1 Announce Type: new Abstract: Modeling multivariate time series by representing them as graphs, where individual series act as nodes and pairwise temporal corre- lations serve as edges, has gained significant traction. Recent advances in Graph Neural Networks (GNNs) have demonstrat...

📖 Read original article


77. Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks ​

Author: Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07335v1 Announce Type: new Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parallelized Q-Network (PQN) algorithm achieves stable off-policy learning without relying on computationally expensive...

📖 Read original article


78. Residual Algebra for Representation-Preserving Learning ​

Author: Yao Wu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07349v1 Announce Type: new Abstract: Learning from heterogeneous representations is usually reduced to feature concatenation, which erases which representation produced an error. We instead algebraize the residual: a representation is a typed object that owns both a coordinate system and ...

📖 Read original article


79. Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning ​

Author: Haoyu Zheng, Yun Zhu, Qing Wang, Wenqiao Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CL

arXiv:2608.07371v1 Announce Type: new Abstract: Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards. However, a completed rollout can yield many such signals, leaving their appropriate allocation across turns unclear. We introduce TRIAL, a trajectory-rela...

📖 Read original article


80. Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations ​

Author: Ioannis Ziogas, Ensieh Khazaei, Bilal Taha, Aamna Al Shehhi, Ahsan H. Khandoker, Leontios J. Hadjileontiadis, Dimitrios Hatzinakos
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, eess.AS, eess.SP, stat.ML

arXiv:2608.07385v1 Announce Type: new Abstract: Learning disentangled representations is a key requirement for developing versatile, general-purpose, and sustainable models in multi-modal wearable computing. However, existing approaches do not operate as full-stack wearable processors, i.e., they do...

📖 Read original article


81. FedDOSE: Federated Learning Framework Decomposing Site Effects for Modeling Brain Dynamic Functional Connectivity ​

Author: Deepank Girish, Yi Hao Chan, Yubin Zheng, Sukrit Gupta, Jagath C. Rajapakse
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, eess.SP, q-bio.NC

arXiv:2608.07393v1 Announce Type: new Abstract: Functional Magnetic Resonance Imaging ( fMRI ) data are often pooled into collaborative multi-site consortia, as deep learning models for analyses require large datasets to generalize well. While Federated Learning (FL) offers a privacy-preserving para...

📖 Read original article


82. Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration ​

Author: Ruochen Jin, Zhanliang Wang, Zongyu Dai, Jiancong Xiao, Bojian Hou
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07419v1 Announce Type: new Abstract: Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Traditional post-hoc temperature scaling is inherently domain-dependent: a temperature fitted on one domain does not generalize across domains. This moti...

📖 Read original article


83. Beyond Myopic World Models: Long-Horizon End-to-End Training for Direct Future Prediction ​

Author: Xinyi Li, Zaishuo Xia, Chenjie Hao, Yubei Chen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.07420v1 Announce Type: new Abstract: World models are expected to support imagination over extended temporal horizons, yet most are still trained through local few-step prediction objectives and deployed by recursively rolling out their own predictions. This creates a fundamental mismatch...

📖 Read original article


84. Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits ​

Author: Elena Dumitrescu, Gert Lek, Lydia Y. Chen, J'er'emie Decouchant
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.07430v1 Announce Type: new Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as targets and as adversaries...

📖 Read original article


85. A proximal subgradient method for nonconvex stochastic optimization under the Kurdyka-{\L}ojasiewicz condition ​

Author: Felipe Atenas, Alejandro Jofr'e, Pedro P'erez-Aros, David Torregrosa-Bel'en
Published: 8/10/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2608.05460v1 Announce Type: cross Abstract: This work introduces a proximal stochastic subgradient method for minimizing the sum of an expected cost, whose integrand is potentially nonsmooth and nonconvex, and a lower semicontinuous, prox-bounded function. We target a broad class of integrands...

📖 Read original article


86. Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes ​

Author: Aoting Zeng, Kai Wang, Jianwei Wang, Yuxiang Sun, Yizhang He, Wenjie Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.06402v1 Announce Type: cross Abstract: Community detection is a fundamental task in graph analytics that aims to identify cohesive groups of entities with similar behaviors or interests. Classic objective-driven methods struggle with complex graph structures, while deep-learning approache...

📖 Read original article


87. UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys ​

Author: Junxiong Zhou, Xuechen Li, Chonghao Qiu, Lang Qiao, Xiaowei Jia, Qi Yang, Chishan Zhang, Leikun Yin, Nanshan You, Vipin Kumar, David Mulla, Ce Yang, Zhenong Jin, Licheng Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.06404v1 Announce Type: cross Abstract: Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and management response. Modern 3D reconstruction methods perform strongly on generic benchmarks, but render...

📖 Read original article


88. Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery ​

Author: Laura Bader, Muhammad Ammar Ahmed, Xiao Xiang Zhu, G"oran Kauermann
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.06406v1 Announce Type: cross Abstract: Accurate estimation of forest height from satellite imagery is essential for applications such as carbon accounting, biodiversity monitoring, and ecosystem management. While recent deep learning approaches provide accurate predictions, they typically...

📖 Read original article


89. Certified Feedforward Tracking for Unknown Nonlinear Systems via Invertible Neural Networks ​

Author: Berk Altiner, Rajasree Sarkar, Arunava Banerjee, Zongxuan Sun, Kenneth Kim
Published: 8/10/2026, 4:00:00 AM
Categories: eess.SY, cs.LG, cs.SY

arXiv:2608.06419v1 Announce Type: cross Abstract: In this paper, we address the certification of datadriven feedforward control for periodic tracking of unknown nonlinear systems under partial state measurements. To this end, we adopt an invertible neural network (INN) as a surrogate for the unknown...

📖 Read original article


90. Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing ​

Author: Iftach Shoham, Tali Dror, Oren Gal, Haim Permuter, Gilad Katz, Eliya Nachmani
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SD, cs.CL, cs.LG

arXiv:2608.06424v1 Announce Type: cross Abstract: Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. Speech inpainting restores missing segments, whereas speech editing replaces spoken content ...

📖 Read original article


91. NTDH: Complex Reasoning for Comprehensive Affective Analysis ​

Author: Tianlei Zhu, Zhiwei Liu, Yuyan Wang, Xiao-Yang Liu, Sophia Ananiadou
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.06425v1 Announce Type: cross Abstract: Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is context-dependent, requiring conflicting cues to be reconciled rather...

📖 Read original article


92. Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models ​

Author: Yong Yang, Roger Newman-Norlund, Xiang Guan, Saeed Ahmadi, Regan Willis, Nadra Salman, Kalil Warren, Sophie Arheix-Parras, Srihari Nelakuditi, Leonardo Bonilha, Christopher Rorden, Rutvik H. Desai, Julius Fridriksson
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.06429v1 Announce Type: cross Abstract: Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior. In earlier work, we lesioned LLMs to produce error profiles in pict...

📖 Read original article


93. Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection ​

Author: Yuewei Sun, Lang Qin, Zechuan Tian, Jingwen Li, Guiqin Wang, Shengzeng Huo, Wenxin Ren, Tao Fang, Xiaochen Zhang, Guanqing Deng, Xiang Wang, Xiaowen Dong, Qinghai Guo, Yuxin Ma
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.LG

arXiv:2608.06434v1 Announce Type: cross Abstract: Embodied intelligence demands both long-horizon reasoning and real-time closed-loop responsiveness. Recent dual-system Vision-Language-Action (VLA) architectures combine fast reactive control with slow deliberative reasoning to balance inference spee...

📖 Read original article


94. Game-Theoretic Inverse Reinforcement Learning for Modeling Competitive Human Driving: A Cut-in Prediction Study ​

Author: Yu Song
Published: 8/10/2026, 4:00:00 AM
Categories: physics.soc-ph, cs.LG

arXiv:2608.06445v1 Announce Type: cross Abstract: Capturing the strategic decision-making inherent in competitive human driving is critical for autonomous vehicle safety and traffic simulation. This study demonstrates that game-theoretic Inverse Reinforcement Learning (IRL) provides a robust framewo...

📖 Read original article


95. FedTransKD-IDS: Robust Federated Transfer Learning with Knowledge Distillation for Intrusion Detection in IoT ​

Author: Mohammad Hosssein Gholamrezazadeh, Ahmadreza MontazerolghaemAhmadreza Montazerolghaem
Published: 8/10/2026, 4:00:00 AM
Categories: cs.NI, cs.LG

arXiv:2608.06447v1 Announce Type: cross Abstract: In modern distributed network environments, particularly in Internet of Things infrastructures and 5G networks, stringent privacy preservation and scalability requirements have created significant challenges for intrusion detection systems. Although ...

📖 Read original article


96. Fairis: Fairness-Aware Aggregation with Provable Influence Containment against Fairness Poisoning Attacks in Collaborative Machine Learning ​

Author: Devharsh Trivedi, Nesrine Kaaniche, Nikos Triandopoulos, Maryline Laurent, Jackson Walters
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.LG

arXiv:2608.06469v1 Announce Type: cross Abstract: Collaborative machine learning among financial institutions must be both group-fair and robust against deliberate adversarial manipulation. Existing fairness-aware aggregation methods remain formally vulnerable to fairness poisoning: a malicious clie...

📖 Read original article


97. LyEvO: Lyapunov-Guided Evolutionary Optimization for Safe and Robust Sim-to-Real Policy Learning ​

Author: Riccardo Curcio, Hongpeng Cao, Marco Caccamo
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.LG, cs.NE

arXiv:2608.06481v1 Announce Type: cross Abstract: Training controllers that are safe and robust in simulation, and systematically assessing their readiness for real-world deployment, remain key challenges in sim-to-real transfer. To address this, we propose LyEvO, a physics-grounded framework that c...

📖 Read original article


98. Density-Functional Excited-State Gradients and Nonadiabatic Couplings on a Consumer GPU from a Contraction-DAG ​

Author: Rub'en Dar'io Guerrero
Published: 8/10/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.AR, cs.LG

arXiv:2608.06536v1 Announce Type: cross Abstract: Nonadiabatic dynamics needs an excited-state gradient and an interstate nonadiabatic coupling matrix element (NACME) at every nuclear geometry, and a double-hybrid functional's accuracy has been unavailable for the coupling. We report the first analy...

📖 Read original article


99. TaskSense: Focusing on What Matters in World Models ​

Author: SM Mazharul Islam, Manfred Huber
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CV, cs.LG

arXiv:2608.06544v1 Announce Type: cross Abstract: World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve information across the entire visual input. However, task-relevant content often occupies only a ...

📖 Read original article


100. Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving ​

Author: Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair, Daniel Berger, Pantea Zardoshti, Rodrigo Fonseca, Esha Choukse
Published: 8/10/2026, 4:00:00 AM
Categories: cs.DC, cs.LG

arXiv:2608.06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests. LLM serving platforms today define response-latency service-level ...

📖 Read original article


101. Cascading Through the Hierarchy: Regularizer-Induced Feature Detection as Phase Transitions in Deep Linear Neural Networks ​

Author: Bj"orn Ladewig, Ibrahim Talha Ersoy, Karoline Wiesner
Published: 8/10/2026, 4:00:00 AM
Categories: cond-mat.stat-mech, cond-mat.dis-nn, cs.LG

arXiv:2608.06597v1 Announce Type: cross Abstract: A scientific theory of deep learning, comprising learning dynamics and statistical properties of learned models, is rapidly gaining attention. One of the corner stones of this development are analytically solvable toy models, allowing for the fully t...

📖 Read original article


102. Pre-Inference Routing for Cost-Efficient Document Field Extraction ​

Author: Sreerekha Rajendran
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.IR, cs.LG

arXiv:2608.06607v1 Announce Type: cross Abstract: Most document-extraction systems use a single model for all documents. This is simple but can be costly for easy cases and less effective for difficult ones. We examine whether we can predict a document's difficulty before extraction using inexpensiv...

📖 Read original article


103. Beyond Co-Movement: Locality by Exposures Enables a Joint Factor-Graph Framework for Portfolio Diversification ​

Author: Sara Chehab, Giorgos Iacovides, Parisa Yazdanparast, Danilo Mandic
Published: 8/10/2026, 4:00:00 AM
Categories: q-fin.PM, cs.LG, q-fin.ST

arXiv:2608.06618v1 Announce Type: cross Abstract: Current portfolio construction methods are either agnostic to the effects of idiosyncratic shocks (standard factor models) or to the latent data structure driving systematic returns (recent graph-based approaches). This presents an opportunity to com...

📖 Read original article


104. Corrupting Attention: Evasion-Based Adversarial Attacks on Encoder Attention in Detection Transformers ​

Author: Ridma Jayasundara, Shaheer Mohamed, Tharindu Fernando, Harshala Gammulle, Basura Fernando, Sanka Rasnayake, A V Subramanyam, Sridha Sridharan, Clinton Fookes
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.06674v1 Announce Type: cross Abstract: Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in many safety-critical systems. Detection transformers have emerged as leading object detectors, ye...

📖 Read original article


105. Optimal Neural Network Approximation via Empirical Least Squares with Deterministic Samples ​

Author: Xinliang Liu, Tong Mao, Jinchao Xu
Published: 8/10/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA

arXiv:2608.06687v1 Announce Type: cross Abstract: We develop a rigorous theory of discrete residual least-squares approximation for elliptic spectral equations $\mathfrak L_\beta u=f$ using linearized ReLU$^k$ neural networks on the sphere, where $\mathfrak L_\beta$ is a positive elliptic spectral m...

📖 Read original article


106. Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models ​

Author: Zhuoheng Huang, Mukesh Singh
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CR, cs.AI, cs.LG

arXiv:2608.06690v1 Announce Type: cross Abstract: Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different systems question: can trusted authorization determine which newly trained parameters are reachable by the forwar...

📖 Read original article


107. Online Monitoring and Corrective Steering of Programming Agents ​

Author: Shuyang Liu, Saman Dehghan, Ji Young Kim, Jatin Ganhotra, Martin Hirzel, Reyhaneh Jabbarvand
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SE, cs.AI, cs.CL, cs.LG

arXiv:2608.06701v1 Announce Type: cross Abstract: Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or the issue description lacks the information needed to localize and repair it. As a result, agents traverse long t...

📖 Read original article


108. MolBioKG: Grounding Out-of-Graph Molecules in Biomedical Knowledge Graphs via Multi-Resolution Structural Anchoring ​

Author: Yiming Zhang, Hikaru Shindo, Shuan Chen, Kaushalya Madhawa, Jun Jin Choong, Yuna Oikawa, Takashi Fujiwara, Keisuke Ozawa
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.06713v1 Announce Type: cross Abstract: Biomedical knowledge graphs (KGs) accelerate drug discovery, but standard pipelines assume query molecules already exist as graph entities, leaving unregistered molecules disconnected. We address this cold-start challenge, termed the out-of-graph mol...

📖 Read original article


109. bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning ​

Author: Koushik Howlader, Tirtho Roy, Md Tauhidul Islam, Wei Le
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.06727v1 Announce Type: cross Abstract: Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. Mixture-of-Recursions (MoR) improves efficiency through adaptive token-choice or expert-choice routing. W...

📖 Read original article


110. Weak Adversarial Neural Pushforward Method for Boltzmann Equation ​

Author: Jenia Fardousi Koly, Andrew Qing He, Wei Cai
Published: 8/10/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA

arXiv:2608.06823v1 Announce Type: cross Abstract: In this paper, we extend a weak adversary neural network pushforward method for solving time dependent Boltzmann equation and a weak formulation of the collision operator is proposed where an invertible neural pushforward mapping is used to generatin...

📖 Read original article


111. FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition ​

Author: Ghani Haider, Majid Kundroo, Boyun Eom, Dong Hwan Park, Chen Chen, Taehong Kim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.DC, cs.LG

arXiv:2608.06876v1 Announce Type: cross Abstract: In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentralized intelligence paradigm for Video Anomaly Recognition (VAR). This task is vital for maintaining high-fidelity ...

📖 Read original article


112. Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models ​

Author: Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.06901v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges for deployment in c...

📖 Read original article


113. Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests ​

Author: Seitaro Ono, Senna Ross, Jun Saiki
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.CY, cs.LG

arXiv:2608.06908v1 Announce Type: cross Abstract: We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT). WEAT is a bias measurement method widely used in both computational social science and AI fairness research. I...

📖 Read original article


114. Stream Learning: Partition-Fair Gossip Learning Without Tokens ​

Author: Fabien Mathieu (NPA), Alexandre Pham (NPA), Maria Gradinariu Potop-Butucaru (NPA), S{'e}bastien Tixeuil (IUF, NPA)
Published: 8/10/2026, 4:00:00 AM
Categories: cs.DC, cs.LG

arXiv:2608.06946v1 Announce Type: cross Abstract: In gossip learning, a network of nodes trains a shared model collaboratively, without a central coordinator, by repeatedly exchanging parts of their local models. The state-of-the-art protocol, Partitioned Token Gossip Learning (PTGL) of Heged{"u}s ...

📖 Read original article


115. Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression ​

Author: Alex Kwon
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2608.06953v1 Announce Type: cross Abstract: Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory. We ask what governs whether it does. Matched notes carry the identical claim and...

📖 Read original article


116. CAi Copilot: Reducing Operational Workload in Molecular Design through Intent-Driven Agentic Workflows ​

Author: Zhu Wang, Jiangyu Chen, Yingjun Shang, Yuhui Yao, Laiao Lu, Tianfan Fu, Na Zou
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.06961v1 Announce Type: cross Abstract: Early-stage molecular design is an iterative process, not just a task of generating molecules. Researchers turn broad goals into design strategies, refine candidates, assess many properties, and gather evidence before synthesis and tests. AI methods ...

📖 Read original article


117. Mixture of Geodesic Factor Analyzers on Riemannian Homogeneous Spaces ​

Author: Hengchao Chen, Yuanyao Tan, Chao Huang, Hongtu Zhu, Qiang Sun
Published: 8/10/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.ST, stat.ME, stat.TH

arXiv:2608.06971v1 Announce Type: cross Abstract: This paper introduces Mixtures of Geodesic Factor Analyzers (MGFA) on Riemannian homogeneous spaces. MGFA uses a geodesic factor model within each mixture component, providing greater expressiveness than mixtures of Riemannian radial distributions an...

📖 Read original article


118. FedLBW: A Loss-Based Weighting Strategy for Federated Learning on Non-IID Data in Wireless Networks ​

Author: Majid Kundroo, Tinku Singh, Taehong Kim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.DC, cs.ET, cs.LG

arXiv:2608.07007v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative machine learning (ML) across distributed clients while preserving privacy. However, efficient model convergence in FL remains challenging, especially in wireless networks where non-independent and identic...

📖 Read original article


119. IceHorizon: A Dataset for Horizon Detection in Ice-Covered Maritime Environments and Comparative Evaluation of Detection Methods ​

Author: Alisa Pesotskaia, Emin Zerman
Published: 8/10/2026, 4:00:00 AM
Categories: eess.IV, cs.CV, cs.LG

arXiv:2608.07018v1 Announce Type: cross Abstract: Horizon detection in images of ice-covered waters is a challenging problem for maritime navigation due to low contrast between water and sky, cluttered ice structures, and varying illumination conditions. This paper presents a comparative evaluation ...

📖 Read original article


120. Limit Points of Reflow with Minibatch Optimal Transport ​

Author: Antonin Chambolle, Johannes Hertrich
Published: 8/10/2026, 4:00:00 AM
Categories: math.PR, cs.LG, cs.NA, math.NA

arXiv:2608.07042v1 Announce Type: cross Abstract: Rectified flows, also called flow matching or stochastic interpolants, are generative models that learn a time-dependent vector field steering a probability curve between two probability distributions, usually referred to as latent and target distrib...

📖 Read original article


121. Tensor Network Kernel Machines: A JAX Framework for Machine Learning and Nonlinear System Identification ​

Author: Albert Saiapin, Kim Batselier
Published: 8/10/2026, 4:00:00 AM
Categories: cs.MS, cs.LG, cs.SY, eess.SY

arXiv:2608.07043v1 Announce Type: cross Abstract: Developing nonlinear models that are both expressive and computationally efficient remains a challenge in machine learning and nonlinear system identification. Tensor network kernel machines (TNKM) address this challenge by combining nonlinear featur...

📖 Read original article


122. AutoIntervene: Calibrated Intervention for Action-Chunking Imitation Learning Policies ​

Author: Jinhe Tang, Weiming Zhi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.AI, cs.CV, cs.HC, cs.LG

arXiv:2608.07065v1 Announce Type: cross Abstract: Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather than single-step commands. Yet perception errors and execution drift can move the robot outside the demonstrati...

📖 Read original article


123. Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking ​

Author: Devin Pereira, Willem Zuidema
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.07077v1 Announce Type: cross Abstract: The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). Current models solve the standard formulation of the puzzle, but still struggle with the flat-to-flat variant (where initial a...

📖 Read original article


124. International Transfer of Stochastic Cortical Self-Reconstruction ​

Author: Fabian Bongratz, Zhizheng Zhuo, Chao Zhang, Yaou Liu, Dennis M. Hedderich, Christian Wachinger
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, q-bio.NC

arXiv:2608.07092v1 Announce Type: cross Abstract: Stochastic cortical self-reconstruction (SCSR) enables personalized mapping of gray matter atrophy, a hallmark of neurodegenerative disorders such as Alzheimer's disease (AD), onto high-resolution cortical surfaces. Unlike conventional normative mode...

📖 Read original article


125. Optimized Certainty Equivalent Risk Minimization Using Samples: Algorithms, Convergence Rates, and Applications ​

Author: Sumedh Gupte, Prashanth L. A., Sanjay P. Bhat
Published: 8/10/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2608.07113v1 Announce Type: cross Abstract: We consider the optimization of the Optimized Certainty Equivalent (OCE) risk, with applications including portfolio optimization in finance, and uncertainty quantification, classification, and regression in machine learning. Our contributions cover ...

📖 Read original article


126. Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory ​

Author: Taeil Kim, Kangsan Kim, Sung Ju Hwang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.07169v1 Announce Type: cross Abstract: Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which struggle to generate sufficient successful trajectories on their own. We propose Agent Memory Distillat...

📖 Read original article


127. Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes ​

Author: Luc Hazenoot, Zhaochun Ren, Amirhossein Zohrehvand
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, econ.GN, q-fin.EC

arXiv:2608.07208v1 Announce Type: cross Abstract: Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding similarities. They score the words a text uses, not the judgment a reader forms about it. Recent work has shown...

📖 Read original article


128. Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-Tuning ​

Author: Vasanth Iyer
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AR, cs.LG

arXiv:2608.07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class accelerators remains limited. This report presents a proof-of-concept deployment of distributed Nano...

📖 Read original article


129. Establishing Boundary KKT Convergence of Mirror Descent through Reparameterization ​

Author: Kuangyu Ding, Kim-Chuan Toh
Published: 8/10/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2608.07248v1 Announce Type: cross Abstract: We prove that mirror descent converges to a KKT point for the nonconvex problem without excluding boundary limits. The result holds under verifiable conditions that jointly couple the objective, the Legendre kernel, and the feasible geometry. The key...

📖 Read original article


130. High-dimensional ridgeless least squares interpolation under spiked covariance structures ​

Author: Zhijun Liu, Dandan Jiang
Published: 8/10/2026, 4:00:00 AM
Categories: math.ST, cs.LG, stat.ML, stat.TH

arXiv:2608.07281v1 Announce Type: cross Abstract: This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally. We consider a generalized spike...

📖 Read original article


131. Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons ​

Author: Guilin Zhang, Kai Zhao
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, math.ST, stat.TH

arXiv:2608.07303v1 Announce Type: cross Abstract: Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and they are easy to get wrong. We report a case study in which a simple AutoML engine, Orcetra, appeared...

📖 Read original article


132. Learning Fault-Tolerant Locomotion with Adaptive Gait Timing ​

Author: Giovanbattista Gravina, Luca Rossini, Carlo Rizzardo, Arturo Laurenzi, Nikos Tsagarakis
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.LG

arXiv:2608.07328v1 Announce Type: cross Abstract: Hardware failures require legged robots to rapidly reorganize coordination and gait timing to maintain stability and mobility. This is particularly challenging for larger quadrupeds, where increased mass and tighter actuation limits reduce the feasib...

📖 Read original article


133. Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding ​

Author: Karim Radouane, Jose G Moreno, Lynda Tamine
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR, cs.LG

arXiv:2608.07353v1 Announce Type: cross Abstract: Understanding concepts is fundamental to generalization. Despite their impressive performance on a wide range of tasks, Large Language Models (LLMs) still struggle with genuine concept understanding. Prior work has evaluated conceptual understanding ...

📖 Read original article


134. LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer's Disease Screening ​

Author: Xin Wang, Yingchao Huang, Yuhan Su, Shanshan Yao, Wei Peng
Published: 8/10/2026, 4:00:00 AM
Categories: eess.AS, cs.AI, cs.LG

arXiv:2608.07378v1 Announce Type: cross Abstract: Early diagnosis of Alzheimer's disease (AD) is critical for enabling timely interventions that may slow disease progression and improve patient outcomes. There is a growing need for AD detection methods that are non-invasive and cost-effective, espec...

📖 Read original article


135. Uncovering expert objectives in production planning via inverse optimization: An industrial case study ​

Author: Shivi Dixit, Rishabh Gupta, Adam Kelloway, John Wassick, Qi Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: math.OC, cs.HC, cs.LG

arXiv:2608.07398v1 Announce Type: cross Abstract: Production planning in the manufacturing industry often relies on the use of optimization models, but defining an appropriate objective function can be a challenge. In practice, planners must balance competing goals, manage uncertainty, and account f...

📖 Read original article


136. DynaCrys: Crystal Generation with Dynamic Space-Group Diffusion ​

Author: Zhuotao Jin, Xiaoyun Wang, Nicholas Brawand, Roman Zubatyuk, Atul Thakur, Eric Qu, Boris Kozinsky, Justin Smith
Published: 8/10/2026, 4:00:00 AM
Categories: cond-mat.mtrl-sci, cs.LG

arXiv:2608.07401v1 Announce Type: cross Abstract: The search for new crystalline materials spans an enormous compositional and structural space. Generating candidates in this space requires jointly modeling discrete crystallographic symmetry, elemental composition, and continuous geometry. We introd...

📖 Read original article


137. Addressable Memory for Video World Models ​

Author: Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky, Laura Leal-Taix'e, Despoina Paschalidou, Jonathan Lorraine, Aljo\v{s}a O\v{s}ep
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.07408v1 Announce Type: cross Abstract: We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames. However, we find that models can no longer reliably address stored con...

📖 Read original article


Author: Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno, Lynda Tamine
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CL, cs.IR, cs.LG

arXiv:2608.07411v1 Announce Type: cross Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmar...

📖 Read original article


139. Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement ​

Author: Xulin Fan, Juan Azcarreta, Ashutosh Pandey, Jesus Alvarez, Ke Tan, Jacob Donley, Ritwik Giri, Buye Xu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.SD, cs.LG

arXiv:2608.07423v1 Announce Type: cross Abstract: Low-latency, low-compute speech enhancement is essential for wearable devices with real-time communication requirements, but strict computational constraints significantly limit on-device performance. Knowledge Boosting has been proposed as an effect...

📖 Read original article


140. Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control ​

Author: Zhaoyu Zhu, Rui Gao, Shuang Li
Published: 8/10/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2608.07433v1 Announce Type: cross Abstract: Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space. We study entropy-regularized discounted linear-quadratic (LQ) control. A Bellman verification argument shows that the unrestricted problem has a...

📖 Read original article


141. Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers ​

Author: Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau, Anass Belfatmi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.07436v1 Announce Type: cross Abstract: Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on $(a+b) \bmod 113$ grok and later lose generalization. Across five seed...

📖 Read original article


Author: Md Tarek Hassan, Dmitry Zelenchuk, Muhammad Ali Babar Abbasi
Published: 8/10/2026, 4:00:00 AM
Categories: eess.SP, cs.ET, cs.LG

arXiv:2608.07444v1 Announce Type: cross Abstract: Accurate user equipment (UE) localization is critical for beam management in reconfigurable intelligent surface (RIS)-assisted millimeter-wave (mmWave) based sixth-generation (6G) networks, especially if the direct base-station-UE links are unavailab...

📖 Read original article


143. CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG ​

Author: Gyuwan Kim, Cheoneum Park, Tao Yang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.IR, cs.LG

arXiv:2608.07458v1 Announce Type: cross Abstract: Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the co...

📖 Read original article


144. MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation ​

Author: Youjun Zhao, Alex Warren, Gary K. L. Tam, Rynson W. H. Lau
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2608.07463v1 Announce Type: cross Abstract: Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. Existing VD...

📖 Read original article


145. Certified Interpolation Oversampling: Per-Instance Safety Guarantees for Imbalanced Learning ​

Author: Pankaj Yadav, Vivek Vijay
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2501.15790v2 Announce Type: replace Abstract: Synthetic minority oversampling is typically designed and evaluated against a predictive objective, generating samples that improve downstream classification. This paper pursues a second objective by generating samples that carry a stated safety pr...

📖 Read original article


146. Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription ​

Author: Benjamin Gutteridge, Matthew Thomas Jackson, Toni Kukurin, Xiaowen Dong
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2502.20295v3 Announce Type: replace Abstract: Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical to obtain for real-world problems, or rely on zero-shot tools such as OCR engines and multi-modal LLMs (MLL...

📖 Read original article


147. Improving Performance of Spike-based Deep Q-Learning using Ternary Neurons ​

Author: Aref Ghoreishee, Abhishek Mishra, John Walsh, Anup Das, Nagarajan Kandasamy
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.NE, cs.SY, eess.SY

arXiv:2506.03392v2 Announce Type: replace Abstract: We propose a new ternary spiking neuron model to improve the representation capacity of binary spiking neurons in deep Q-learning. Although a ternary neuron model has recently been introduced to overcome the limited representation capacity offered ...

📖 Read original article


148. Minimal Ingredients for Reward Assignment from Expert Demonstrations ​

Author: Zixuan Dong, Yumi Omori, Keith Ross
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2506.06793v2 Announce Type: replace Abstract: Reward assignment from scarce demonstrations is a key challenge in both offline and online imitation learning. A common and intuitive strategy assigns rewards according to how closely learner trajectories match expert demonstrations. Although this ...

📖 Read original article


149. Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers ​

Author: Nikolay Yudin, Sergei Kudriashov, Alexander Gaponov, Maxim Rakhuba
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.NA

arXiv:2507.07814v2 Announce Type: replace Abstract: We introduce a novel upper bound on the local Lipschitz constant of the dot-product self-attention block showing its dependence on the attention map distributions. The proposed bound is not only tighter than the prior art, but for the first time, r...

📖 Read original article


150. Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement ​

Author: Meihua Dang, Jiaqi Han, Minkai Xu, Kai Xu, Akash Srivastava, Stefano Ermon
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2507.08390v5 Announce Type: replace Abstract: Discrete diffusion models have recently emerged as strong alternatives to autoregressive language models, matching their performance through large-scale training. However, inference-time control remains relatively underexplored. In this work, we st...

📖 Read original article


151. Optimization-based Online Conformal Prediction for Multi-step Forecasting ​

Author: Ruipu Li, Daniel Menacho, Alexander Rodr'iguez
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2508.13362v3 Announce Type: replace Abstract: Conformal prediction (CP) provides distribution-free coverage guarantees, making it well suited for uncertainty quantification in time series forecasting. However, existing methods often struggle with multi-step settings: they either calibrate hori...

📖 Read original article


152. Optimization as a Dynamical System: Generative Schedules from Latent ODEs ​

Author: Matt L. Sampson, Peter Melchior
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2509.23052v2 Announce Type: replace Abstract: We present a new meta-learning method to determine the optimal learning rate schedule for gradient descent. It leverages training runs from a hyperparameter search to learn a latent representation of the training process, which is modeled as a dyna...

📖 Read original article


153. Provable Training Data Identification for Large Language Models ​

Author: Zhenlong Liu, Hao Zeng, Weiran Huang, Hongxin Wei
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2510.09717v3 Announce Type: replace Abstract: Identifying training data of large-scale models is critical for copyright litigation, privacy auditing, and ensuring fair evaluation. However, existing works typically treat this task as an instance-wise identification without controlling the error...

📖 Read original article


154. Stability of Transformers under Layer Normalization ​

Author: Kelvin Kan, Xingjian Li, Benjamin J. Zhang, Tuhin Sahai, Stanley Osher, Krishna Kumar, Markos A. Katsoulakis
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.OC

arXiv:2510.09904v2 Announce Type: replace Abstract: Despite their widespread use, training deep Transformers can be unstable. Layer normalization, a standard component, improves training stability, but its placement has often been ad-hoc. In this paper, we conduct a principled study on the forward (...

📖 Read original article


155. Iterative Training of Physics-Informed Neural Networks with Fourier-enhanced Features ​

Author: Yulun Wu, Miguel Aguiar, Karl H. Johansson, Matthieu Barreau
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2510.19399v2 Announce Type: replace Abstract: Spectral bias, the tendency of neural networks to learn low-frequency features first, is a well-known issue with many training algorithms for physics-informed neural networks (PINNs). To overcome this issue, we propose IFeF-PINN, an algorithm for i...

📖 Read original article


156. Defining Energy Indicators for Impact Identification on Aerospace Composites: A Structured Feature Selection Approach Guided by Domain Knowledge ​

Author: Nat'alia Ribeiro Marinho, Richard Loendersloot, Frank Grooteman, Jan Willem Wiegman, Uraz Odyurt, Tiedo Tinga
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, physics.app-ph

arXiv:2511.01592v2 Announce Type: replace Abstract: Energy estimation is critical to impact identification on aerospace composites, where low-velocity impacts can induce internal damage that is undetectable at the surface. Data sparsity, signal noise, complex feature interdependencies, non-linear dy...

📖 Read original article


157. In Situ Training of Implicit Neural Compressors for Scientific Simulations via Sketch-Based Regularization ​

Author: Cooper Simpson, Stephen Becker, Alireza Doostan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CE, cs.NA, math.NA

arXiv:2511.02659v4 Announce Type: replace Abstract: Focusing on implicit neural representations, we present a novel in situ training protocol that employs limited memory buffers of full and sketched data samples, where the sketched data are leveraged to prevent catastrophic forgetting. The theoretic...

📖 Read original article


158. Equivariant Sparse Autoencoders: Mechanistic Interpretability of Neural Networks on Symmetric Data ​

Author: Ege Erdogan, Ana Lucic
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2511.09432v2 Announce Type: replace Abstract: Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity. In particular, their activations entangle many concepts into fewer dimensions, a phenomenon known as superposition. Mechanis...

📖 Read original article


159. PCAE: Learning Ordered Representations in Latent Space for Intrinsic Dimension Estimation via Principal Component Autoencoder ​

Author: Qipeng Zhan, Zhuoping Zhou, Zexuan Wang, Li Shen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2601.19179v2 Announce Type: replace Abstract: Autoencoders have long been considered a nonlinear extension of Principal Component Analysis (PCA). Prior studies have demonstrated that linear autoencoders (LAEs) can recover the ordered, axis-aligned principal components of PCA by incorporating n...

📖 Read original article


160. Self-Distillation Enables Continual Learning ​

Author: Idan Shenfeld, Mehul Damani, Jonas H"ubotter, Pulkit Agrawal
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2601.19897v2 Announce Type: replace Abstract: Continual learning, enabling models to acquire new skills and knowledge without degrading existing capabilities, remains a fundamental challenge for foundation models. While on-policy reinforcement learning can reduce forgetting, it requires explic...

📖 Read original article


161. Parameter-free Dynamic Regret: Time-varying Movement Costs, Delayed Feedback, and Memory ​

Author: Hao Qiu, Andrew Jacobsen, Emmanuel Esposito, Mengxiao Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, stat.ML

arXiv:2602.06902v3 Announce Type: replace Abstract: In this paper, we study dynamic regret in unconstrained online convex optimization (OCO) with movement costs. Specifically, we generalize the standard setting by allowing the movement cost coefficients $\lambda_t$ to vary arbitrarily over time. Our...

📖 Read original article


162. Large Causal Models for Temporal Causal Discovery ​

Author: Nikolaos Kougioulis, Nikolaos Gkorgkolis, MingXue Wang, Bora Caglayan, Dario Simionato, Andrea Tonon, Ioannis Tsamardinos
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2602.18662v3 Announce Type: replace Abstract: Causal discovery for both cross-sectional and temporal data has traditionally followed a dataset-specific paradigm, where a new model is fitted for each individual dataset. Such an approach limits the potential of multi-dataset pretraining. The con...

📖 Read original article


163. MAC: A Conversion Rate Prediction Benchmark Featuring Labels Under Multiple Attribution Mechanisms ​

Author: Jinqi Wu, Sishuo Chen, Zhangming Chan, Yong Bai, Lei Zhang, Sheng Chen, Chenghuan Hou, Xiang-Rong Sheng, Han Zhu, Jian Xu, Bo Zheng, Chaoyou Fu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.02184v2 Announce Type: replace Abstract: Multi-attribution learning (MAL), which enhances model performance by learning from conversion labels yielded by multiple attribution mechanisms, has emerged as a promising learning paradigm for conversion rate (CVR) prediction. However, the conver...

📖 Read original article


164. Seeking SOTA: Time-Series Forecasting Must Adopt Taxonomy-Specific Evaluation to Dispel Illusory Gains ​

Author: Raeid Saqur, Christoph Bergmeir, Blanka Horvath, Daniel Schmidt, Frank Rudzicz, Terry Lyons
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2603.15506v2 Announce Type: replace Abstract: We argue that the current practice of evaluating AI/ML time-series forecasting models, predominantly on benchmarks characterized by strong, persistent periodicities and seasonalities, obscures real progress by overlooking the performance of efficie...

📖 Read original article


165. Conditioning Protein Generation via Hopfield Pattern Multiplicity ​

Author: Jeffrey D. Varner
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, q-bio.BM, q-bio.QM

arXiv:2603.20115v2 Announce Type: replace Abstract: Small protein-family alignments often contain a subset of interest but not enough labeled data to train a conditional generator. We condition a training-free stochastic-attention sampler by adding one multiplicity ratio to its logits. Increasing th...

📖 Read original article


166. TMTE: Effective Multimodal Graph Learning with Task-aware Modality and Topology Co-evolution ​

Author: Yinlin Zhu, Xunkai Li, Di Wu, Wang Luo, Miao Hu, Guocong Quan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2603.27723v2 Announce Type: replace Abstract: Multimodal-attributed graphs (MAGs) are a fundamental data structure for multimodal graph learning (MGL), enabling both graph-centric and modality-centric tasks. However, our empirical analysis reveals inherent topology quality limitations in real-...

📖 Read original article


167. CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment ​

Author: Anurag Kumar, Raghuveer Peri, Jon Burnsky, Alexandru Nelus, Rohit Paturi, Srikanth Vishnubhotla, Yanjun Qi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.00310v2 Announce Type: replace Abstract: Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Models aligned on text alone show a higher rate of successful attacks when extended to two or more modalities...

📖 Read original article


168. Embedded Variational Neural Stochastic Differential Equations for Learning Heterogeneous Dynamics ​

Author: Sandeep Kumar Samota, Reema Gupta, Snehashish Chakraverty
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, math.DS

arXiv:2604.00669v2 Announce Type: replace Abstract: This study examines the challenges of modeling complex and noisy data related to socioeconomic factors over time, with a focus on data from various districts in Odisha, India. Traditional time-series models struggle to capture both trends and varia...

📖 Read original article


169. Cluster Attention for Graph Machine Learning ​

Author: Oleg Platonov, Liudmila Prokhorenkova
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2604.07492v2 Announce Type: replace Abstract: Message Passing Neural Networks have recently become the most popular approach to graph machine learning tasks; however, their receptive field is limited by the number of message passing layers. To increase the receptive field, Graph Transformers w...

📖 Read original article


170. Intersectional Disentangling of Temporal and Acquisition Bias in Fetal Ultrasound ​

Author: Aya Elgebaly, Joris Fournel, Benjamin Laine J{\o}nch Jurgensen, Kamil Mikolaj, Anders Christensen, Martin Tolsgaard, Claes Ladefoged, Aasa Feragen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CV, eess.IV

arXiv:2605.02942v3 Announce Type: replace Abstract: Fairness studies of medical imaging AI often explain subgroup performance gaps through under-representation in the training data. We show that intersectional analysis can disentangle fairness and performance gaps arising from clinical and acquisiti...

📖 Read original article


171. DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts ​

Author: Jiading Gai, Shuai Zhang, Xiang Song, Bernie Wang, George Karypis
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.15422v4 Announce Type: replace Abstract: Modern RL post-training methods such as GRPO and DAPO train on N response sequences of R tokens sampled from a shared prompt of P tokens, but standard FlashAttention replicates all P prompt tokens N times across both forward and backward passes -- ...

📖 Read original article


172. The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought ​

Author: Moritz Br"osamle, Stephan Eckstein
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CC, cs.CL

arXiv:2605.18079v2 Announce Type: replace Abstract: Existing expressivity results for transformers typically rely on hardmax attention, high precision, and other architectural modifications that disconnect them from the models used in practice. We bridge this gap by analyzing standard transformer de...

📖 Read original article


173. Mitigating Gradient Pathology in PINNs through Aligned Constraint ​

Author: Yichen Luo, Peiyu Zhu, Dongxiao Hu, Jia Wang, Tailin Wu, Dapeng Lan, Yu Liu, Zhibo Pang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.25001v2 Announce Type: replace Abstract: While Physics-Informed Neural Networks (PINNs) are powerful for solving Partial Differential Equations (PDEs), their training is often paralyzed by gradient pathology. The gradients from the PDE residuals and boundary constraints oppose each other,...

📖 Read original article


174. The Challenges of Using Reinforcement Learning for Controlling Industrial Energy Systems ​

Author: Tobias Lademann, Th'eo Vincent, Jan Peters, Matthias Weigold
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2605.31044v2 Announce Type: replace Abstract: Reinforcement learning has shown promising results for optimizing the control of industrial energy systems, yet most existing studies remain limited to the application in simulation environments. We investigate the challenges of deploying reinforce...

📖 Read original article


175. Rethinking Evaluation Paradigms in IBP-based Certified Training ​

Author: Konstantin Kaulen, Hadar Shavit, Holger H. Hoos
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.CV

arXiv:2606.02134v2 Announce Type: replace Abstract: Deep neural networks achieve strong performance on many supervised learning tasks but remain vulnerable to adversarial perturbations. Neural network verification provides mathematically rigorous robustness guarantees, yet at substantial computation...

📖 Read original article


176. TiWeaver: Unified Temporal Dynamics Modeling via Contextual Patching ​

Author: Zhe Li, Jindong Tian, Hao Miao, Zhi Lei, Chenjuan Guo, Bin Yang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2606.03121v2 Announce Type: replace Abstract: Multivariate time series forecasting plays a critical role in real-world applications, including weather prediction, stock analysis, and health monitoring. Due to the diversity of data sources, time series exhibit diverse temporal dynamics, often a...

📖 Read original article


177. Beyond Structural Symmetries: Linear Mode Connectivity via Neuron Identifiability ​

Author: Vincent B"urgin, Daniel Herbst, Ya-Wei Eileen Lin, Stefanie Jegelka
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2606.04754v2 Announce Type: replace Abstract: Many striking phenomena in deep learning, such as linear mode connectivity and the structured behavior of training dynamics, are closely tied to parameter symmetries: transformations that leave the realized function unchanged. Despite growing atten...

📖 Read original article


178. AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents ​

Author: Hao Bai, Rui Yang, Chenlu Ye, Spencer Whitehead, Aviral Kumar, Tong Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2606.05597v3 Announce Type: replace Abstract: Training vision-language web agents with multi-step RL is compute-intensive, with two dominant forms of inefficiency: idle GPUs in synchronous RL, and trajectories that use more steps and tokens than necessary. We present AsyncWebRL, which addresse...

📖 Read original article


179. Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path ​

Author: Thomas Sesmat, Gabriel Meseguer-Brocal, Geoffroy Peeters
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.SD

arXiv:2606.07271v2 Announce Type: replace Abstract: Understanding memorization in generative models remains challenging, with implications for copyright and privacy. Beyond verbatim reproduction, models can encode subtler traces of their training data that never surface in their outputs yet remain e...

📖 Read original article


180. An Empirical Study of openPangu Quantization on Ascend NPUs ​

Author: Tong Shi, Jiacheng Wang, Hui Xie, Ying Li, Aishan Liu, Jinyang Guo, Xianglong Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2606.21257v4 Announce Type: replace Abstract: openPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive post-training quantization on Ascend NPUs has not been systematically characterized. This paper conducts a contr...

📖 Read original article


181. Training-free Task Classification for Multi-Task Model Merging ​

Author: Jungyong Son, Jinwook Jung, Sungyong Baik
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2606.22589v2 Announce Type: replace Abstract: Ever since the advent of foundation models and the pre-training-finetuning paradigm, there have been numerous efforts to merge multiple task-specific experts into a single multi-task model. Prior work largely focuses on finding a single merged mode...

📖 Read original article


182. Compositional Behavioral Semantics for State Abstraction in Reinforcement Learning ​

Author: Yivan Zhang, Ziyan Luo, Manuel Baltieri
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, math.CT

arXiv:2606.25357v2 Announce Type: replace Abstract: State abstraction plays a key role in scaling reinforcement learning to complex but structured systems. In studying such systems, a wide range of behavioral structures have been studied in reinforcement learning, including value functions, invarian...

📖 Read original article


183. Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models ​

Author: Jiacheng Zheng, Chang Guo, Zixuan Wang, Xinyu Liu, Hao Chen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, q-bio.QM

arXiv:2607.10729v3 Announce Type: replace Abstract: Molecular property models are commonly evaluated by holding out Bemis-Murcko scaffolds, yet a scaffold identifier is only one notion of chemical unfamiliarity. We introduce a label-free structural-frontier split that reserves the sparsest and most ...

📖 Read original article


184. When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal ​

Author: James E. Allchin
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.11953v3 Announce Type: replace Abstract: Does a reinforcement-learning agent that earns high reward actually learn its task's hidden state, or only a shortcut that correlates with reward? We build an instrument that makes this question directly measurable: the task is a hidden finite auto...

📖 Read original article


185. Reducing information dependency does not cause training data privacy. Adversarially non-robust features do ​

Author: Rasmus Torp, Shailen K. Smith, Adam Breuer
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.12354v2 Announce Type: replace Abstract: In this paper, we challenge the prevailing view that information dependency (including rote memorization) drives training data exposure to image reconstruction attacks. We show that extensive exposure can persist without rote memorization and is in...

📖 Read original article


186. Energy-Based Physics-Informed Form Finding for Clustered Tensegrity Structures ​

Author: Jing Qin, Muhao Chen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.NA, math.NA

arXiv:2607.12888v2 Announce Type: replace Abstract: Tensegrity form-finding and physical property prediction are fundamental inverse problems in structural mechanics, which aim to determine equilibrium configurations and internal force distributions. These problems are challenging due to strong nonl...

📖 Read original article


187. DeepLoop: Depth Scaling for Looped Transformers ​

Author: Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.13491v2 Announce Type: replace Abstract: Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Trans...

📖 Read original article


188. Counterfactual Shapley Credit Assignment ​

Author: Mingxuan Li, Kai-Zhan Lee, Elias Bareinboim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.16999v2 Announce Type: replace Abstract: The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing frameworks, whether relying on temporal contiguity or hindsight-conditioned reward reweighting, frequently fail ...

📖 Read original article


189. IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning ​

Author: Wei Zhang, Xinwu Liu, Yihang Cheng
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2607.22251v2 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT) of LLMs whose effectiveness depends on rank allocation. Existing adaptive LoRA methods derive ranks from local gradient, activation, or matrix statistics...

📖 Read original article


190. Neurai-VN Benchmark: Standardized Machine Learning Models for Multimodal Digital Phenotyping in Mental Health Classification ​

Author: Quoc-Cuong Pham, Hoang-Thuy-Duong Vu, Thi-Thanh-Huong Ha, Huy-Hieu Pham
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.25232v2 Announce Type: replace Abstract: Digital phenotyping (DP) using smartphones and wearable devices has shown considerable potential for mental health monitoring. However, progress remains difficult to evaluate due to heterogeneous datasets, inconsistent preprocessing pipelines. In t...

📖 Read original article


191. MUGEN: A Unified Framework for Efficient Motion Understanding and Generation ​

Author: Zhankai Ye, Yukai Jin, Bingyang Wei, Bofan Li, Yusen Wu, Fangyi Li, Shangqian Gao, Xin Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2607.27581v2 Announce Type: replace Abstract: Grounding human motion in language, and language in motion, is a central step toward physical AI systems that can understand, generate, and communicate human behavior. Unified motion--language systems first coupled the two directions through a shar...

📖 Read original article


192. ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow ​

Author: Dongxiu Liu, Haoyi Niu, Peng Cheng, Yuan Gao, Xirui Kang, Sangli Teng, Koushil Sreenath, Xianyuan Zhan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CV, cs.RO

arXiv:2607.27924v2 Announce Type: replace Abstract: In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are largely confined to discrete-time prediction, thereby exhibiting significant inefficiency in capturin...

📖 Read original article


193. Topology-Aware Data Movement for Disaggregated GPU Inference ​

Author: Sanjeev Rao Ganjihal
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, cs.PF

arXiv:2607.28633v2 Announce Type: replace Abstract: Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run on separate GPU pools, the KV cache must be transferred between them. For a 70B model this is 1.3 GB per reque...

📖 Read original article


194. Provably Learning Multi-Head Attention with Queries ​

Author: Sunyeop Kim, Insung Kim, Jian Guo
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CR

arXiv:2608.03294v2 Announce Type: replace Abstract: We study the problem of learning multi-head softmax attention from black-box input-output access. The learner may query arbitrary real-valued token sequences and observe only the scalar output at the final token. Recent work gives an algorithm usin...

📖 Read original article


195. A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation ​

Author: Abdallah Khemais
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.03620v2 Announce Type: replace Abstract: Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different objects: one forward pass versus the parameters behind every forward pass. We ask when they agree. We study an id...

📖 Read original article


196. SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant ​

Author: Adel Javanmard, David P. Woodruff, Vahab Mirrokni
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI, stat.ML

arXiv:2608.05127v2 Announce Type: replace Abstract: Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD, use high-dimensional geometric constructions but incur unfavorable d...

📖 Read original article


197. Disentangling 3D Modeling from Spatial Reasoning ​

Author: Haoze Sun, Jiequan Cui, Qingshan Xu, Richang Hong
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.CV

arXiv:2608.05242v2 Announce Type: replace Abstract: In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perception and reasoning through large-scale training. Our key observation is...

📖 Read original article


198. PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis ​

Author: Xiaomin He, Dongling Xiao, Jiahao Xie, Ruiqi Lu, Qianle Wang, Zhongbin Guo, Wanxuan Sun
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.05249v2 Announce Type: replace Abstract: Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering one self-contained question. We study this gap through rubric comprehen...

📖 Read original article


199. Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading ​

Author: Rasul Khanbayov, Hasan Kurban
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG

arXiv:2608.05675v2 Announce Type: replace Abstract: Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, whic...

📖 Read original article


200. Is Self-Pretraining really useful to improve diagnosis in medical Time Series? ​

Author: Omar Coser, Antonio Orvieto, Paolo Soda, Loredana Zollo
Published: 8/10/2026, 4:00:00 AM
Categories: cs.LG, cs.AI

arXiv:2608.06122v2 Announce Type: replace Abstract: Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series. Our ...

📖 Read original article


201. Quantum Generative Diffusion Model: A Fully Quantum-Mechanical Model for Generating Quantum State Ensemble ​

Author: Chuangtao Chen, Qinglin Zhao, MengChu Zhou, Zhimin He, Zhili Sun, Haozhen Situ
Published: 8/10/2026, 4:00:00 AM
Categories: quant-ph, cs.LG

arXiv:2401.07039v5 Announce Type: replace-cross Abstract: Mixed quantum states are the native description of many physically important quantum systems, making their generation a fundamental task in quantum information processing. However, constructing a diffusion process that generates density opera...

📖 Read original article


202. Boundary Density Likelihood for Direct Event-Time Supervision ​

Author: Clark Peng, Tolga Din\c{c}er
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG, stat.ML

arXiv:2408.12792v2 Announce Type: replace-cross Abstract: Event detection turns long recordings into a sparse set of ranked timestamps. Yet many sequence models are trained for samplewise segmentation and only convert predicted states into events after training. We ask whether training directly for ...

📖 Read original article


203. Decentralized Indoor Localization Based on A Sparse Gaussian Process with Reduced-Dimensional Inputs for Real-Time Sensing and Training on IoT Devices ​

Author: Zhe Tang, Sihao Li, Zichen Huang, Guandong Yang, Kyeong Soo Kim, Jeremy S. Smith, Zhaowei Zhu, Qi Xuan
Published: 8/10/2026, 4:00:00 AM
Categories: eess.SP, cs.LG, cs.NI

arXiv:2409.00078v2 Announce Type: replace-cross Abstract: As a large number of Internet of Things (IoT) devices are deployed in the field, there arises huge potential of edge computing for indoor localization on those devices. Conventional indoor localization based on a centralized server with subst...

📖 Read original article


204. Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions ​

Author: Iskander Azangulov, George Deligiannidis, Judith Rousseau
Published: 8/10/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, math.ST, stat.TH

arXiv:2409.18804v3 Announce Type: replace-cross Abstract: Denoising Diffusion Probabilistic Models (DDPM) are powerful state-of-the-art methods used to generate synthetic data from high-dimensional data distributions and are widely used for image, audio, and video generation as well as many more app...

📖 Read original article


205. Harnessing the Synergy between LLM Agents and Knowledge Graphs for Urban Socioeconomic Prediction ​

Author: Zhilun Zhou, Jingyang Fan, Yu Liu, Fengli Xu, Depeng Jin, Yong Li
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.SI

arXiv:2411.00028v3 Announce Type: replace-cross Abstract: Socioeconomic prediction aims to leverage various urban data to predict the socioeconomic indicators of regions such as population and commercial activity level, which plays an important role in understanding urban regions and supporting deci...

📖 Read original article


206. CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM ​

Author: Qingyuan Liu, Liyan Chen, Haocheng Wang, Yanning Yang, Dong Du, Zhigang Mao, Naifeng Jing, Yubin Xia, Haibo Chen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AR, cs.LG

arXiv:2504.17584v2 Announce Type: replace-cross Abstract: Attention-FC Disaggregated (AFD) LLM inference systems offload memory-bound Attention operations to memory-rich accelerators (e.g., CPUs, HBM-PIM) while retaining compute-bound Fully-Connected (FC) operations on GPUs. In this paper, we first ...

📖 Read original article


207. PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks ​

Author: Ziyuan Li, Uwe Jaekel, Babette Dellen
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG, eess.IV

arXiv:2505.04397v3 Announce Type: replace-cross Abstract: Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexplored. Product units offer a direct approach to modeling such interactions, but their use in deep archit...

📖 Read original article


208. Symbolic Graphics Programming with Large Language Models ​

Author: Yamei Chen, Haoquan Zhang, Yangyi Huang, Zeju Qiu, Kaipeng Zhang, Yandong Wen, Weiyang Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2509.05208v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at program synthesis, yet their ability to produce symbolic graphics programs (SGPs) that render into precise visual content remains underexplored. We study symbolic graphics programming, where the goal is t...

📖 Read original article


209. Using Reinforcement Learning to Optimize the Global and Local Crossing Number ​

Author: Timo Brand, Henry F"orster, Stephen Kobourov, Daniel Kohrt, Robin Schukrafft, Markus Wallinger, Johannes Zink
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CG, cs.LG

arXiv:2509.06108v4 Announce Type: replace-cross Abstract: Graph drawing concerns the algorithmic visualization of graphs. A good drawing of a graph is easy to read and facilitates solving tasks on the graph. Several properties have been identified to occur in good drawings of graphs. Such properties...

📖 Read original article


210. Robot guide with multi-agent control and automatic scenario generation with LLM ​

Author: Elizaveta D. Moskovskaya, Anton D. Moscowsky
Published: 8/10/2026, 4:00:00 AM
Categories: cs.RO, cs.LG

arXiv:2509.10317v2 Announce Type: replace-cross Abstract: The article describes the development of a hybrid social robot control architecture to overcome the limitations of traditional approaches, where behavior scripts manually synchronize the robot's actions and text, and existing methods focus pr...

📖 Read original article


211. Let's Unlearn Stereotypes Before Decision-Making: Assessing the Impact of Intrinsic Bias Mitigation on Downstream Fairness in LLMs ​

Author: Mina Arzaghi, Alireza Dehghanpour Farashah, Florian Carichon, Jean-Fran\c{c}ois Plante, Golnoosh Farnadi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.CY, cs.LG

arXiv:2509.16462v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used in high-stakes decision-making systems, where biased predictions can reinforce social and economic disparities. Although prior work has examined intrinsic representational bias and unfair dow...

📖 Read original article


212. Mean-square and sublinear convergence of a stochastic proximal point algorithm in metric spaces of nonpositive curvature ​

Author: Nicholas Pischke
Published: 8/10/2026, 4:00:00 AM
Categories: math.OC, cs.LG

arXiv:2510.10697v2 Announce Type: replace-cross Abstract: We define a stochastic variant of the proximal point algorithm in the general setting of nonlinear Hadamard spaces for approximating zeros of the mean of a stochastically perturbed monotone vector field. Generalizing previous work by P. Bianc...

📖 Read original article


213. Free Denoising Diffusion Models ​

Author: Swagatam Das
Published: 8/10/2026, 4:00:00 AM
Categories: math.PR, cs.LG, stat.ML

arXiv:2510.22778v3 Announce Type: replace-cross Abstract: We develop a free-probabilistic framework for denoising diffusion, in which the data is a self-adjoint operator and its law a spectral distribution. The forward process is the free Ornstein--Uhlenbeck diffusion, whose spectral marginals solve...

📖 Read original article


214. Robust inference using density-powered Stein operators ​

Author: Shinto Eguchi
Published: 8/10/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2511.03963v3 Announce Type: replace-cross Abstract: We introduce a density-power weighted variant of the Stein operator, called the $\gamma$-Stein operator, for robust inference with unnormalized probability models. The operator is motivated by the first variation of the $\gamma$-divergence un...

📖 Read original article


215. Intelligence per Watt: Measuring Intelligence Efficiency of Local AI ​

Author: Jon Saad-Falcon, Avanika Narayan, Hakki Orhun Akengin, J. Wes Griffin, Herumb Shandilya, Adrian Gamarra Lafuente, Medhya Goel, Rebecca Joseph, Shlok Natarajan, Etash Kumar Guha, Shang Zhu, Ben Athiwaratkun, John Hennessy, Azalia Mirhoseini, Christopher R'e
Published: 8/10/2026, 4:00:00 AM
Categories: cs.DC, cs.AI, cs.CL, cs.LG

arXiv:2511.07885v5 Announce Type: replace-cross Abstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strains this paradigm faster than providers can scale. Two advances create an opportunity to rethink it: smal...

📖 Read original article


216. BDD2Seq: Enabling Scalable Reversible-Circuit Synthesis via Graph-to-Sequence Learning ​

Author: Mingkai Miao, Jianheng Tang, Guangyu Hu, Hongce Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AR, cs.LG

arXiv:2511.08315v2 Announce Type: replace-cross Abstract: Binary Decision Diagrams (BDDs) are instrumental in many electronic design automation (EDA) tasks thanks to their compact representation of Boolean functions. In BDD-based reversible-circuit synthesis, which is critical for quantum computing,...

📖 Read original article


217. Sampling via Stochastic Interpolants by Langevin-based Velocity and Initialization Estimation in Flow ODEs ​

Author: Chenguang Duan, Yuling Jiao, Gabriele Steidl, Christian Wald, Jerry Zhijian Yang, Ruizhe Zhang
Published: 8/10/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA, math.PR, stat.ML

arXiv:2601.08527v3 Announce Type: replace-cross Abstract: We propose a novel method for sampling from unnormalized Boltzmann densities based on a probability flow ordinary differential equation (ODE) derived from linear stochastic interpolants. The key innovation of our approach is the use of a sequ...

📖 Read original article


218. Kimi K2.5: Visual Agentic Intelligence ​

Author: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, S. H. Cai, Yuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Cheng Chen, Guanduo Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kefan Chen, Liang Chen, Ruijue Chen, Xinhao Chen, Yanru Chen, Yanxu Chen, Yicun Chen, Yimin Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Ziwei Chen, Dazhi Cheng, Yean Cheng, Minghan Chu, Jialei Cui, Jiaqi Deng, Muxi Diao, Hao Ding, Mengfan Dong, Mengnan Dong, Yuxin Dong, Yuhao Dong, Angang Du, Chenzhuang Du, Dikang Du, Lingxiao Du, Yulun Du, Yu Fan, Shengjun Fang, Qiulin Feng, Yichen Feng, Garimugai Fu, Kelin Fu, Hongcheng Gao, Tong Gao, Yuyao Ge, Shangyi Geng, Chengyang Gong, Xiaochen Gong, Zhuoma Gongque, Qizheng Gu, Xinran Gu, Yicheng Gu, Longyu Guan, Shuhao Guan, Yuanying Guo, Xiaoru Hao, Dailan He, Tianhong He, Weiran He, Wenyang He, Yibo He, Yunjia He, Chao Hong, Hao Hu, Jiaxi Hu, Yangyang Hu, Zhenxing Hu, Ke Huang, Ruiyuan Huang, Weixiao Huang, Zhiqi Huang, Chaobo Jia, Tao Jiang, Zhejun Jiang, Xinyi Jin, Yu Jing, Guokun Lai, Aidi Li, C. Li, Cheng Li, Fang Li, Guanghe Li, Guanyu Li, Haitao Li, Haoyang Li, Jia Li, Jingwei Li, Junxiong Li, Lincan Li, Mo Li, Weihong Li, Wentao Li, Xinhang Li, Xinhao Li, Yang Li, Yanhao Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zhaoxi Li, Zheming Li, Weilong Liao, Jiawei Lin, Xiaohan Lin, Yibo Lin, Zhishan Lin, Zichao Lin, Cheng Liu, Chenyu Liu, Hongzhang Liu, Liang Liu, Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Tianyu Liu, Weizhou Liu, Xiangyan Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yuanxin Liu, Zhengying Liu, Zhongnuo Liu, Enzhe Lu, Haoyu Lu, Zhiyuan Lu, G. Luo, Junyu Luo, Tongxu Luo, Yashuo Luo, Long Ma, Shaoguang Mao, Yuan Mei, Xin Men, Fanqing Meng, Zhiyong Meng, Yibo Miao, Minqing Ni, Kun Ouyang, Siyuan Pan, Bo Pang, Yuchao Qian, Ruoyu Qin, Zeyu Qin, Jiezhong Qiu, Bowen Qu, Zeyu Shang, Youbo Shao, Tianxiao Shen, Zhennan Shen, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Feifan Song, Pengwei Song, Tianhui Song, Xiaoxi Song, Hongjin Su, Jianlin Su, Zhaochen Su, Lin Sui, Jinsong Sun, Junyao Sun, Tongyu Sun, Flood Sung, Yunpeng Tai, Chuning Tang, Heyi Tang, Xiaojuan Tang, Zhengyang Tang, Jiawen Tao, Shiyuan Teng, Chaoran Tian, Pengfei Tian, Bowen Wang, Chensi Wang, Chuang Wang, Congcong Wang, Dingkun Wang, Dinglu Wang, Dongliang Wang, Feng Wang, Hailong Wang, Haiming Wang, Hao Wang, Hengzhi Wang, Huaqing Wang, Hui Wang, Jiahao Wang, Jinhong Wang, Jiuzheng Wang, Kaixin Wang, Linian Wang, Qibin Wang, Shengjie Wang, Shuyi Wang, Si Wang, Wei Wang, Xiaochen Wang, Xinyuan Wang, Yao Wang, Yejie Wang, Yipu Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhexu Wang, Zifan Wang, Zihan Wang, Zizhe Wang, Chu Wei, Ming Wei, Chuan Wen, Zichen Wen, Chengjie Wu, Haoning Wu, Junyan Wu, Rucong Wu, Wenhao Wu, Yuefeng Wu, Yuhao Wu, Yuxin Wu, Zijian Wu, Chenjun Xiao, Jin Xie, Xiaotong Xie, Yuchong Xie, Bowei Xing, Boyu Xu, Jianfan Xu, Jing Xu, Jinjing Xu, L. H. Xu, Lin Xu, Suting Xu, Weixin Xu, Xinbo Xu, Xinran Xu, Yangchuan Xu, Yichang Xu, Yuemeng Xu, Zelai Xu, Ziyao Xu, Junjie Yan, Yuzi Yan, Guangyao Yang, Hao Yang, Junwei Yang, Kai Yang, Ningyuan Yang, Xiaofei Yang, Xinlong Yang, Xinyu Yang, Ying Yang, Yi Yang, Yi Yang, Zhen Yang, Zhilin Yang, Zonghan Yang, Haotian Yao, Dan Ye, Haoran Ye, Wenjie Ye, Zhuorui Ye, Peng Yebo, Bohong Yin, Chengzhen Yu, Longhui Yu, Tao Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Xiaokun Yuan, Yang Yue, Weihao Zeng, Dunyuan Zha, Haobing Zhan, Dehao Zhang, Hao Zhang, Jin Zhang, Puqi Zhang, Qiao Zhang, Rui Zhang, Xiaobin Zhang, Xiaoyun Zhang, Y. Zhang, Yadong Zhang, Yangkun Zhang, Yichi Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yushun Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Chenguang Zhao, Feifan Zhao, Jinxiang Zhao, Shuai Zhao, Xiangyu Zhao, Xuanle Zhao, Yikai Zhao, Zijia Zhao, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Junfeng Zhong, Longguang Zhong, Weiming Zhong, M. Zhou, Runjie Zhou, Xinyu Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yuxuan Zhu, Zhen Zhu, Jingze Zhuang, Weiyu Zhuang, Ying Zou, Xinxing Zu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2602.02276v2 Announce Type: replace-cross Abstract: We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of t...

📖 Read original article


219. Transferable FB-GNN-MBE Framework for Potential Energy Surfaces: Data-Adaptive Transfer Learning in Deep Learned Many-Body Expansion Theory ​

Author: Siqi Chen, Zhiqiang Wang, Yili Shen, Xianqi Deng, Xi Cheng, Cheng-Wei Ju, Jun Yi, Guo Ling, Dieaa Alhmoud, Hui Guan, Zhou Lin
Published: 8/10/2026, 4:00:00 AM
Categories: physics.chem-ph, cs.LG

arXiv:2604.09320v4 Announce Type: replace-cross Abstract: Mechanistic understanding and rational design of complex chemical systems depend on fast and accurate predictions of electronic structures beyond individual building blocks. However, if the system exceeds hundreds of atoms, first-principles q...

📖 Read original article


220. Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages ​

Author: Kevin Guan, Happy Buzaaba, Christiane Fellbaum
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG

arXiv:2605.02608v2 Announce Type: replace-cross Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers---the Biaffine LSTM,...

📖 Read original article


221. {\Omega}-QVLA: Robust Quantization for Vision-Language-Action Models via Composite Rotation and Per-step Scaling ​

Author: Xinyu Wang, Mingze Li, Sicheng Lyu, Dongxiu Liu, Kaicheng Yang, Ziyu Zhao, Yufei Cui, Xiao-Wen Chang, Peng Lu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.LG

arXiv:2605.28803v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models unify perception, reasoning, and control within a single policy, yet their multi-billion-parameter backbones and diffusion-based action heads make on-device deployment prohibitively expensive. Prior quantiz...

📖 Read original article


222. Vector Space of Cycles ​

Author: Moo K. Chung, Anass B. El-Yaagoubi, Hernando Ombao
Published: 8/10/2026, 4:00:00 AM
Categories: stat.ML, cs.LG, physics.data-an, q-bio.NC

arXiv:2606.08202v2 Announce Type: replace-cross Abstract: Most statistical and machine learning methods for directed interactions focus on pairwise effects among variables. Even existing cyclic models represent feedback primarily through node-level dependencies, making large-scale recurrent organiza...

📖 Read original article


223. The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products ​

Author: Hao-Ping Lee, Jessica He, David Piorkowski, Thomas Serban von Davier, Jodi Forlizzi, Sauvik Das
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.HC, cs.LG, cs.SE

arXiv:2606.15485v2 Announce Type: replace-cross Abstract: Agentic AI systems act autonomously, use tools, adapt to context, and operate in complex real-world environments. However, these same characteristics can create or exacerbate product risks. We studied how industry developers (n=35) perceive, ...

📖 Read original article


224. Diffusion-MF: Approximate Structured Diffusion for Sequence Labelling ​

Author: Nicolas Floquet, Joseph Le Roux, Nadi Tomeh
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2606.18856v3 Announce Type: replace-cross Abstract: We introduce Diffusion-MF, a discrete diffu- sion sequence labeller that places a linear-chain conditional random field (LCRF) inside the denoising loop. Unlike prior diffusion labellers, it performs structured inference at every step; parall...

📖 Read original article


225. Dirac-Frenkel dynamics with inertia for nonlinearly parametrized solutions of evolution problems ​

Author: Matteo Raviola, Benjamin Peherstorfer
Published: 8/10/2026, 4:00:00 AM
Categories: math.NA, cs.LG, cs.NA

arXiv:2606.24769v3 Announce Type: replace-cross Abstract: Even when Dirac-Frenkel dynamics determine a well-defined evolution in function space, the corresponding parameter dynamics can be non-unique or ill-conditioned for redundant nonlinear parametrizations, such as typical neural networks or mixt...

📖 Read original article


226. RenderFormer++: Scalable and Physics-Informed Feed-Forward Neural Rendering ​

Author: Huangsheng Du, Haoran Zhu, Youcheng Cai, Jingyang Meng, Ligang Liu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.GR, cs.CV, cs.LG

arXiv:2606.30380v2 Announce Type: replace-cross Abstract: We present RenderFormer++, a scalable and physics-informed feed-forward neural rendering framework for global illumination in mesh scenes. Existing Transformer-based neural rendering methods such as RenderFormer achieve promising cross-scene ...

📖 Read original article


227. LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models ​

Author: Sojung An, Junha Lee, Sujeong You, Nam Ik Cho, Donghyun Kim
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CV, cs.AI, cs.LG

arXiv:2607.06918v2 Announce Type: replace-cross Abstract: Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM adaptation stems from the prohibitive costs of full fine-tuning and catastrophic forgetting. To address t...

📖 Read original article


228. Kimi K3: Open Frontier Intelligence ​

Author: Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, Yuxuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Guanduo Chen, Guangyu Chen, Guanzheng Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kexin Chen, Peng Chen, Ruijue Chen, Wentao Chen, Xin Chen, Yang Chen, Yanru Chen, Yifei Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Dazhi Cheng, Yean Cheng, Jialei Cui, Jingbing Cui, Anqi Dai, Jiaqi Deng, Hao Ding, Rui Ding, Shaofeng Ding, Mengfan Dong, Mengnan Dong, Yuhao Dong, Yuxin Dong, Angang Du, Chenzhuang Du, Dikang Du, Jusen Du, Yulun Du, Yu Fan, Jing Feng, Qiulin Feng, Yichen Feng, Kelin Fu, Qiang Fu, Fuxuan Gao, Hongcheng Gao, Jingyue Gao, Tong Gao, Weijia Gao, Shangyi Geng, Jie Gong, Linhu Gong, Shengao Gong, Xiaochen Gong, Qizheng Gu, Yicheng Gu, Shuhao Guan, Haiqing Guo, Shiqi Guo, Xiang Guo, Zhengyan Guo, Beixi Hao, Wenxin Hao, Xiaoru Hao, Dailan He, Haotian He, Lehan He, Qi He, Weiran He, Xinran He, Xinyi He, Yibo He, Yunjia He, Chao Hong, Tiange Hong, Hao Hu, Jiaxi Hu, Ruikun Hu, Weiming Hu, Yangyang Hu, Zhenxing Hu, Liang Hua, Jinbin Huang, Ke Huang, Ruiyuan Huang, Siying Huang, Weixiao Huang, Yan Huang, Zhengjie Huang, Zhiqi Huang, Yulong Hui, Chaobo Jia, Yutong Jiang, Zhejun Jiang, Zuoyou Jiang, Wenyi Jin, Xinyi Jin, Yu Jing, Huanjun Kong, Guokun Lai, Aidi Li, Cheng Li, Chengyuan Li, Cong Li, Fang Li, Guanyu Li, Haoyang Li, Jia Li, Junxiong Li, Lei Li, Letian Li, Lincan Li, Weihong Li, Wentao Li, Xintong Li, Yang Li, Yishen Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zhaoxi Li, Zheming Li, Zhengxiao Li, Zhiyuan Li, Jiawei Lin, Xiaohan Lin, Yibo Lin, Zichao Lin, Ziyan Lin, Bill Liu, Boxiao Liu, Chuan Liu, Liang Liu, Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Weizhou Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yipeng Liu, Zhengying Liu, Zhiheng Liu, Enzhe Lu, Haoyu Lu, Linqiang Lu, Tingzhan Lu, Zhiyuan Lu, Aotian Luo, G. Luo, Junyu Luo, Yifan Luo, B. Lyu, Wenzhou Lyu, Shaoguang Mao, Yuan Mei, Xin Men, Minqing Ni, Yixuan Niu, Siyuan Pan, Shujun Peng, Zhangyang Qi, Ruoyu Qin, ZeChao Qin, Zeyu Qin, Haiquan Qiu, Jianxin Qiu, Jiezhong Qiu, Bowen Qu, Yuhao Qu, Zeyu Shang, Youbo Shao, Han Shen, Jincheng Shi, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Wingchun Siu, Pengwei Song, Xiaoxi Song, Jianlin Su, Yunfeng Su, Zhaochen Su, Lin Sui, Jingsong Sun, Junyao Sun, Shaoning Sun, Shuzhe Sun, Tongyu Sun, Yujun Sun, Yunpeng Tai, Chuning Tang, Heyi Tang, Sirui Tang, Zecheng Tang, Chaoran Tian, Rongpeng Tian, Yu Tian, Wei Tu, Chensi Wang, Chuang Wang, Chunjie Wang, Dinglu Wang, Feng Wang, Hailong Wang, Haiming Wang, Hao Wang, Hao Wang, Huaqing Wang, Hui Wang, Jiayi Wang, Jinglong Wang, Jinhong Wang, Jiuzheng Wang, Linian Wang, Shaobo Wang, Shenzhi Wang, Shuyi Wang, Si Wang, Siyuan Wang, Tianfu Wang, Wenjue Wang, Xingran Wang, Xinmei Wang, Xinyuan Wang, Xusheng Wang, Yalin Wang, Yangkun Wang, Yao Wang, Yaoyu Wang, Yejie Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhenhao Wang, Zhongsheng Wang, Zifan Wang, Chu Wei, Ming Wei, Shouxin Wei, Zichen Wen, Fan Wu, Haoning Wu, Rucong Wu, Wenhao Wu, Xiaoxue Wu, Yingcong Wu, Yongqi Wu, Yuxin Wu, Zijian Wu, Xinglang Xian, Chenxuan Xiang, Yuye Xiang, Bocheng Xiao, Chenjun Xiao, Xin Xiao, Jin Xie, Xiaotong Xie, Yifeng Xie, Zhe Xie, Bowei Xing, Yiming Xiong, Baosheng Xu, Boyu Xu, Jiale Xu, Jianfan Xu, Jing Xu, Jinjing Xu, L. H. Xu, Qingtao Xu, Shuyao Xu, Suting Xu, Tiantian Xu, Tianxiang Xu, Weixin Xu, Xinran Xu, Yangchuan Xu, Ye Xu, Yueni Xu, Ziyao Xu, Haonan Xue, Junjie Yan, Yaoyao Yan, Fan Yang, Guangyao Yang, Hao Yang, Junwei Yang, Ruoyu Yang, Wenjie Yang, Xiaofei Yang, Xinyu Yang, Yi Yang, Yiling Yang, Ying Yang, Yuchen Yang, Zhen Yang, Zhilin Yang, Zian Yang, Zuhao Yang, Haotian Yao, Dan Ye, Haoran Ye, Wenjie Ye, Zhanbo Ye, Bohong Yin, Haoxiang Yin, Xietong Yin, Chengzhen Yu, Haozhen Yu, Longhui Yu, Shengnan Yu, Shuying Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Tongtian Yue, Wei Yue, Yang Yue, Dunyuan Zha, Haobing Zhan, B. H. Zhang, Dehao Zhang, Fei Zhang, Hao Zhang, Haoyuan Zhang, Huanyu Zhang, Jiapei Zhang, Jiaxuan Zhang, Jin Zhang, Kaiyi Zhang, Miaozhen Zhang, Puqi Zhang, Qinglei Zhang, Rong Zhang, Rui Zhang, Shaoshuai Zhang, Shiyi Zhang, Xiaobin Zhang, Xiaoyun Zhang, Y. Zhang, Yangkun Zhang, Ye Zhang, Yichi Zhang, Yikun Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Zijing Zhang, Bin Zhao, Chenguang Zhao, Feifan Zhao, Jinglun Zhao, Jinxiang Zhao, Shuai Zhao, Wenshuo Zhao, Xiangyu Zhao, Xuanle Zhao, Yikai Zhao, Zijia Zhao, Haozhi Zheng, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Haofeng Zhong, Lei Zhong, Longguang Zhong, M. Zhou, Qiankang Zhou, Runjie Zhou, Ruozhang Zhou, Xinyu Zhou, Yiqiao Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yangjunfeng Zhu, Yuxuan Zhu, Zhen Zhu, Chen Zhuang, Weiyu Zhuang, Xinxing Zu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2607.24653v2 Announce Type: replace-cross Abstract: We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which i...

📖 Read original article


229. Can AI agents conduct open-ended AI research? Early evidence from two case studies ​

Author: Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.CY, cs.LG

arXiv:2607.27191v2 Announce Type: replace-cross Abstract: Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open...

📖 Read original article


230. H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases ​

Author: Shusen Zhang, Junyi Hu, Ye Feng, Ziteng Wang, Zhaoyuan Pan, Xiaojun Yuan, Jiangshou Hong, Guosheng Dong, Xiangzhi Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.00065v3 Announce Type: replace-cross Abstract: Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constraints, and compositional concepts. However, existing representations lie at two extremes: single-vector...

📖 Read original article


231. LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference ​

Author: Zhichen Liu, Ruihan Sun, Hengjie Yang, Zipeng Wu, Zhaohan Chen, Xiaofan Zhang, Yang Xu
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.LG

arXiv:2608.02515v2 Announce Type: replace-cross Abstract: Long-running assistants and agents consume interaction streams that eventually outgrow the context. Existing context retention, summarization, and retrieval preserve access to selected history, but do not provide a persistent state over the f...

📖 Read original article


232. Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model ​

Author: Abdallah Khemais
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.03629v2 Announce Type: replace-cross Abstract: A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream. For the one composition in that model where two car...

📖 Read original article


233. A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper ​

Author: Ali Shendabadi, Parnia Izadirad, Mostafa Salehi
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CL, cs.AI, cs.LG, cs.SD

arXiv:2608.05165v2 Announce Type: replace-cross Abstract: Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we study the use of Whisper for Persian SER with a particular focus on representation dimensionality reduction...

📖 Read original article


234. Challenges for Musical Education in the Age of AI and Digital Transformation ​

Author: Jean-Pierre Briot
Published: 8/10/2026, 4:00:00 AM
Categories: cs.CY, cs.AI, cs.LG

arXiv:2608.05176v2 Announce Type: replace-cross Abstract: Music education has never been a static discipline. Each major technological shift has forced educators and institutions to reconsider what they teach, how they teach it, and why. We now stand at what may be the most consequential of such tur...

📖 Read original article


235. SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse ​

Author: Jialuo Chen, Minghe Wang, Lingqi Jiang, Jianan Ma, Xinhao Deng, Xiaohu Du, Ruixiao Lin, Yunhao Feng, Linkang Du, Jingyi Wang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05204v2 Announce Type: replace-cross Abstract: LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. As skills become marketplace artifacts, auditing their re...

📖 Read original article


236. Recursive Synthesis for Long-Horizon Terminal Tasks ​

Author: Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang
Published: 8/10/2026, 4:00:00 AM
Categories: cs.AI, cs.LG

arXiv:2608.05466v2 Announce Type: replace-cross Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually ...

📖 Read original article


237. Deep Generalised Mixed Models: a Novel Neural Network Structure for Analysing Hierarchical Data ​

Author: Nina van Gerwen, Dimitris Rizopoulos, Manon Hillegers, Loes Keijsers, Sten Willemsen
Published: 8/10/2026, 4:00:00 AM
Categories: stat.ML, cs.LG

arXiv:2608.05930v2 Announce Type: replace-cross Abstract: The experience sampling method (ESM) is a longitudinal research design where participants report their thoughts, emotional states and behaviours multiple times a day. Our work is motivated by such data collected by the GrowIt! app, which was ...

📖 Read original article