{"ok":true,"snapshot":{"date":"2026-10-03","capturedAt":"2026-10-03T11:31:23.195Z","total_papers":50,"categories_queried":["cs.AI","cs.LG","cs.CL","cs.CV"],"raw_count":100,"papers":[{"arxivId":"2610.02210","version":"v1","title":"Moore, Escher, Penrose: A Conformal Golden Braid","abstract":"I don't think I have ever done anything as peculiar in my life. Among other things, it shows a young man looking with interest at a print on the wall of an exhibition that features himself. How can this be? Perhaps I am not far removed from Einstein's curved universe.'' So wrote M.C. Escher about his 1956 lithograph Print Gallery. Nearly half a century later, a mathematical analysis related its geometry to an untwisted source image through a conformal power map $z \\mapsto z^α$, $α\\in \\mathbb{C}$. Building on this construction, we use a frozen text-to-image diffusion model to generate new self-referential scenes. Prompting alone does not enforce the recursion, while a post-hoc transformation can leave structures poorly connected. Applying the transformation during sampling is also insuffic…","authors":["Sophia Feldman","Assaf Shocher"],"primaryCategory":"cs.CV","categories":["cs.CV"],"publishedAt":"2026-10-01T17:59:59Z","updatedAt":"2026-10-01T17:59:59Z","htmlUrl":"https://arxiv.org/abs/2610.02210v1","pdfUrl":"https://arxiv.org/pdf/2610.02210v1","doi":null},{"arxivId":"2610.02207","version":"v1","title":"One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars","abstract":"3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architecture…","authors":["Ramazan Fazylov","Stamatis Lefkimmiatis","Ivan Laptev"],"primaryCategory":"cs.CV","categories":["cs.CV","cs.AI","cs.HC","cs.LG"],"publishedAt":"2026-10-01T17:59:58Z","updatedAt":"2026-10-01T17:59:58Z","htmlUrl":"https://arxiv.org/abs/2610.02207v1","pdfUrl":"https://arxiv.org/pdf/2610.02207v1","doi":null},{"arxivId":"2610.02208","version":"v1","title":"Sphere Encoder 2","abstract":"Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at \\href{https://github.…","authors":["Kaiyu Yue","Sean McLeish","Ruchit Rawal","Brian Bartoldson","Menglin Jia","Tom Goldstein"],"primaryCategory":"cs.CV","categories":["cs.CV"],"publishedAt":"2026-10-01T17:59:58Z","updatedAt":"2026-10-01T17:59:58Z","htmlUrl":"https://arxiv.org/abs/2610.02208v1","pdfUrl":"https://arxiv.org/pdf/2610.02208v1","doi":null},{"arxivId":"2610.02206","version":"v1","title":"KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards","abstract":"LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is con…","authors":["Pengfei Li","Naufal Suryanto","Sicheng Zhang","Muzammal Naseer"],"primaryCategory":"cs.CL","categories":["cs.CL","cs.AI","cs.CR"],"publishedAt":"2026-10-01T17:59:55Z","updatedAt":"2026-10-01T17:59:55Z","htmlUrl":"https://arxiv.org/abs/2610.02206v1","pdfUrl":"https://arxiv.org/pdf/2610.02206v1","doi":null},{"arxivId":"2610.02205","version":"v1","title":"ROWBench: Do Video Models Render What the Program Specifies?","abstract":"Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instruction or physical adherence, but rarely test fidelity to fine-grained, program-specified world events. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. PROWBench logs entity states and timestamped events, including those outside the camera's field of view, as replayable world records, from which it renders synchronized views and proxy representations. This enables generated videos to be ch…","authors":["Zheng-Hui Huang","Guixu Lin","Yu-Ju Tsai","Jian-Kai Zhu","Fengbo Lan","Yu-Lun Liu","Yung-Yu Chuang","Kaipeng Zhang"],"primaryCategory":"cs.CV","categories":["cs.CV"],"publishedAt":"2026-10-01T17:59:53Z","updatedAt":"2026-10-01T17:59:53Z","htmlUrl":"https://arxiv.org/abs/2610.02205v1","pdfUrl":"https://arxiv.org/pdf/2610.02205v1","doi":null},{"arxivId":"2610.02204","version":"v1","title":"Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents","abstract":"Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for re…","authors":["Yen-Jen Wang","Haozhe Jiang","Shuying Deng","Haoru Xue","Weirui Ye","Rocky Duan","Nika Haghtalab","S. Shankar Sastry"],"primaryCategory":"cs.RO","categories":["cs.RO","cs.AI","eess.SY"],"publishedAt":"2026-10-01T17:59:50Z","updatedAt":"2026-10-01T17:59:50Z","htmlUrl":"https://arxiv.org/abs/2610.02204v1","pdfUrl":"https://arxiv.org/pdf/2610.02204v1","doi":null},{"arxivId":"2610.02203","version":"v1","title":"Embedding Prediction Helps Image Generation","abstract":"In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\\times2…","authors":["Sihan Xu","Ji Xie","Zilin Wang","Hui Shen","Stella X. Yu"],"primaryCategory":"cs.CV","categories":["cs.CV","cs.LG"],"publishedAt":"2026-10-01T17:59:49Z","updatedAt":"2026-10-01T17:59:49Z","htmlUrl":"https://arxiv.org/abs/2610.02203v1","pdfUrl":"https://arxiv.org/pdf/2610.02203v1","doi":null},{"arxivId":"2610.02202","version":"v1","title":"ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research","abstract":"What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature a…","authors":["Sohyeon Kim","Yoonho Lee","Bo Liu","Dayoon Ko","Rulin Shao","Seungone Kim","Graham Neubig","Pang Wei Koh"],"primaryCategory":"cs.AI","categories":["cs.AI","cs.CL","cs.IR"],"publishedAt":"2026-10-01T17:59:47Z","updatedAt":"2026-10-01T17:59:47Z","htmlUrl":"https://arxiv.org/abs/2610.02202v1","pdfUrl":"https://arxiv.org/pdf/2610.02202v1","doi":null},{"arxivId":"2610.02201","version":"v1","title":"SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation","abstract":"High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-…","authors":["Tianjiao Yu","Xinzhuo Li","Yifan Shen","Ying Shen","Kiet A. Nguyen","Adheesh Sunil Juvekar","Ismini Lourentzou"],"primaryCategory":"cs.CV","categories":["cs.CV","cs.AI","cs.LG"],"publishedAt":"2026-10-01T17:59:46Z","updatedAt":"2026-10-01T17:59:46Z","htmlUrl":"https://arxiv.org/abs/2610.02201v1","pdfUrl":"https://arxiv.org/pdf/2610.02201v1","doi":null},{"arxivId":"2610.02200","version":"v1","title":"VISTA: A Visual Harness for Reasoning in an Interactive World","abstract":"We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA'…","authors":["Qiushi Han","Keya Hu","Linlu Qiu","Cathy Wu","Kaiming He"],"primaryCategory":"cs.AI","categories":["cs.AI","cs.CV"],"publishedAt":"2026-10-01T17:59:45Z","updatedAt":"2026-10-01T17:59:45Z","htmlUrl":"https://arxiv.org/abs/2610.02200v1","pdfUrl":"https://arxiv.org/pdf/2610.02200v1","doi":null},{"arxivId":"2610.02199","version":"v1","title":"TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning","abstract":"Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently introduced Muon optimizer reduces optimizer memory through matrix-valued updates. Still, its geometry differs from AdamW and can lead to performance degradation when fine-tuning AdamW-pretrained models. To reduce optimizer memory without sacrificing accuracy or computational efficiency in LLM fine-tuning, we propose Ternary Absolute-max Column-wise One-sparse optimizer, or TACO, which follows Muon's operator-norm steepest-descent view but takes the geometric route further. TACO…","authors":["Jichao Jiang","Cristian McGee","El Houcine Bergou","Hanqin Cai","Aritra Dutta"],"primaryCategory":"cs.LG","categories":["cs.LG","math.OC"],"publishedAt":"2026-10-01T17:59:42Z","updatedAt":"2026-10-01T17:59:42Z","htmlUrl":"https://arxiv.org/abs/2610.02199v1","pdfUrl":"https://arxiv.org/pdf/2610.02199v1","doi":null},{"arxivId":"2610.02198","version":"v1","title":"FERPO: Forward Entropy-Regularized Policy Optimization","abstract":"Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL…","authors":["Sebastian Sanokowski","Alireza Sarmadi","Majid Khadiv"],"primaryCategory":"cs.LG","categories":["cs.LG","cs.AI","cs.RO","stat.ML"],"publishedAt":"2026-10-01T17:59:41Z","updatedAt":"2026-10-01T17:59:41Z","htmlUrl":"https://arxiv.org/abs/2610.02198v1","pdfUrl":"https://arxiv.org/pdf/2610.02198v1","doi":null},{"arxivId":"2610.02197","version":"v1","title":"HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation","abstract":"Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, \"a balloon floating upward while steam rises from a pot\" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enfor…","authors":["Tahira Kazimi","Shubhankar Borse","Munawar Hayat","Fatih Porikli","Pinar Yanardag"],"primaryCategory":"cs.CV","categories":["cs.CV"],"publishedAt":"2026-10-01T17:59:41Z","updatedAt":"2026-10-01T17:59:41Z","htmlUrl":"https://arxiv.org/abs/2610.02197v1","pdfUrl":"https://arxiv.org/pdf/2610.02197v1","doi":null},{"arxivId":"2610.02196","version":"v1","title":"InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation","abstract":"We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new …","authors":["Zhuo Lin","Sirui Xu","Liuyu Bian","Yu-Xiong Wang","Liang-Yan Gui"],"primaryCategory":"cs.RO","categories":["cs.RO","cs.CV","cs.GR"],"publishedAt":"2026-10-01T17:59:40Z","updatedAt":"2026-10-01T17:59:40Z","htmlUrl":"https://arxiv.org/abs/2610.02196v1","pdfUrl":"https://arxiv.org/pdf/2610.02196v1","doi":null},{"arxivId":"2610.02195","version":"v1","title":"Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control","abstract":"The generalized Schrödinger bridge on a graph moves mass between two distributions while charging a cost for the states visited. It has been approached by learning the rates of a controlled continuous-time Markov chain, with a temporal-difference penalty that restores the cost. A state cost folds into the reference process as a Feynman-Kac tilt. The cost-augmented bridge is then a plain bridge against the tilted reference, and the penalty is unnecessary. The bridge is computed exactly by alternating two endpoint rescalings, each one sparse matrix-exponential application; nothing is discretized in time or learned. The alternation converges at a rate set by the endpoint coupling alone. For a quadratic congestion cost on time-averaged occupancies, damped best response around the exact bridge…","authors":["Akshay Balsubramani"],"primaryCategory":"cs.LG","categories":["cs.LG"],"publishedAt":"2026-10-01T17:59:39Z","updatedAt":"2026-10-01T17:59:39Z","htmlUrl":"https://arxiv.org/abs/2610.02195v1","pdfUrl":"https://arxiv.org/pdf/2610.02195v1","doi":null},{"arxivId":"2610.02193","version":"v1","title":"Hierarchical Continuous Diffusion Language Models","abstract":"Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training …","authors":["Hui Ren","Zihan Li","Chang Liu","Huidong Liu","Alexander Schwing"],"primaryCategory":"cs.CL","categories":["cs.CL","cs.AI","cs.LG"],"publishedAt":"2026-10-01T17:59:39Z","updatedAt":"2026-10-01T17:59:39Z","htmlUrl":"https://arxiv.org/abs/2610.02193v1","pdfUrl":"https://arxiv.org/pdf/2610.02193v1","doi":null},{"arxivId":"2610.02191","version":"v1","title":"The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models","abstract":"While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \\hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock s…","authors":["Shuo Xing","Zilin Dai","Chengyuan Qian","Fangzhou Lin","Wenjing Chen","Ping He","Pan Lu","Alvaro Velasquez"],"primaryCategory":"cs.LG","categories":["cs.LG"],"publishedAt":"2026-10-01T17:59:32Z","updatedAt":"2026-10-01T17:59:32Z","htmlUrl":"https://arxiv.org/abs/2610.02191v1","pdfUrl":"https://arxiv.org/pdf/2610.02191v1","doi":null},{"arxivId":"2610.02190","version":"v1","title":"Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning","abstract":"Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \\textbf{Z}ero-and-\\textbf{F}irst-\\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-select…","authors":["Cristian McGee","El Houcine Bergou","Aritra Dutta"],"primaryCategory":"cs.LG","categories":["cs.LG","math.OC"],"publishedAt":"2026-10-01T17:59:28Z","updatedAt":"2026-10-01T17:59:28Z","htmlUrl":"https://arxiv.org/abs/2610.02190v1","pdfUrl":"https://arxiv.org/pdf/2610.02190v1","doi":null},{"arxivId":"2610.02189","version":"v1","title":"Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features","abstract":"Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcem…","authors":["Jason X. Liu","Sebastian Ibarraran","Frank Hu","Soojung Yang","Xinyu A. Feng","Abigail Park","Anagha Aneesh","Lacramioara Bintu"],"primaryCategory":"cs.LG","categories":["cs.LG"],"publishedAt":"2026-10-01T17:59:20Z","updatedAt":"2026-10-01T17:59:20Z","htmlUrl":"https://arxiv.org/abs/2610.02189v1","pdfUrl":"https://arxiv.org/pdf/2610.02189v1","doi":null},{"arxivId":"2610.02188","version":"v1","title":"DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation","abstract":"Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking disc…","authors":["Zhengming Yu","Junkun Yuan","Haotian Yang","Gordon Guocheng Qian","Yizhi Wang","Angtian Wang","Yiding Yang","Bo Liu"],"primaryCategory":"cs.CV","categories":["cs.CV","cs.AI"],"publishedAt":"2026-10-01T17:59:01Z","updatedAt":"2026-10-01T17:59:01Z","htmlUrl":"https://arxiv.org/abs/2610.02188v1","pdfUrl":"https://arxiv.org/pdf/2610.02188v1","doi":null},{"arxivId":"2610.02186","version":"v1","title":"Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry","abstract":"Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding th…","authors":["Yiming Huang","Yujie Zeng","Vijay Prakash Dwivedi","Simone Foti","Jianmin Wang","Jure Leskovec","Tolga Birdal"],"primaryCategory":"cs.LG","categories":["cs.LG","cs.AI"],"publishedAt":"2026-10-01T17:58:45Z","updatedAt":"2026-10-01T17:58:45Z","htmlUrl":"https://arxiv.org/abs/2610.02186v1","pdfUrl":"https://arxiv.org/pdf/2610.02186v1","doi":null},{"arxivId":"2610.02185","version":"v1","title":"Decoding Looped Transformers Better for (Almost) Free","abstract":"Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gai…","authors":["Weihao Liu","Huangjie Zheng","Tianrong Chen","Rohit Dilip","Richard He Bai","Yizhu Jiao","Yuyang Wang","Ruixiang Zhang"],"primaryCategory":"cs.LG","categories":["cs.LG"],"publishedAt":"2026-10-01T17:58:38Z","updatedAt":"2026-10-01T17:58:38Z","htmlUrl":"https://arxiv.org/abs/2610.02185v1","pdfUrl":"https://arxiv.org/pdf/2610.02185v1","doi":null},{"arxivId":"2610.02182","version":"v1","title":"SoftServe: A Scalable Quasi-Newton Method for Deep Learning","abstract":"Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limited their use in deep learning: non-convexity and enormous parameter sizes. We introduce SoftServe, a family of QN methods designed to overcome these obstacles without line searches or ad hoc curvature corrections. SoftServe derives positivedefinite curvature estimates from the variational objective of Berglund et al. (2025), even in the presence of negative curvature. We develop diagonal and Kroneckerfactored variants that preserve positive definiteness by construction and scale to massive neural networks. Finally, SoftServe relies on the stable coupled Newton-Schulz iteration for the required matrix operations, replacing costly matrix decomp…","authors":["Joohwan Ko","Tetiana Parshakova","Diana Cai","Robert M. Gower"],"primaryCategory":"cs.LG","categories":["cs.LG","cs.AI"],"publishedAt":"2026-10-01T17:58:24Z","updatedAt":"2026-10-01T17:58:24Z","htmlUrl":"https://arxiv.org/abs/2610.02182v1","pdfUrl":"https://arxiv.org/pdf/2610.02182v1","doi":null},{"arxivId":"2610.02181","version":"v1","title":"OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning","abstract":"We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleav…","authors":["Haibo Wang","Jiteng Mu","Jialu Li","Jingru Yi","Yuanjun Xiong","Jianming Zhang","Lifu Huang","Mingze Xu"],"primaryCategory":"cs.CV","categories":["cs.CV"],"publishedAt":"2026-10-01T17:58:16Z","updatedAt":"2026-10-01T17:58:16Z","htmlUrl":"https://arxiv.org/abs/2610.02181v1","pdfUrl":"https://arxiv.org/pdf/2610.02181v1","doi":null},{"arxivId":"2610.02180","version":"v1","title":"Generative Cinematographer: Composing Camera and Object Motion in 3D","abstract":"Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video…","authors":["Jiahan Zhang","Chaohao Yang","Namitha Guruprasad","Vivekjyoti Banerjee","Trong-Tung Nguyen","Alan Yuille","Anand Bhattad"],"primaryCategory":"cs.CV","categories":["cs.CV","cs.AI"],"publishedAt":"2026-10-01T17:58:09Z","updatedAt":"2026-10-01T17:58:09Z","htmlUrl":"https://arxiv.org/abs/2610.02180v1","pdfUrl":"https://arxiv.org/pdf/2610.02180v1","doi":null},{"arxivId":"2610.02179","version":"v1","title":"From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation","abstract":"Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences i…","authors":["Siqi Zhu","Suozhi Huang","Kaixuan Zhang","Yuheng Yang","Zhanyang Jin","Yihang Sun","Jiaxuan You"],"primaryCategory":"cs.LG","categories":["cs.LG"],"publishedAt":"2026-10-01T17:57:44Z","updatedAt":"2026-10-01T17:57:44Z","htmlUrl":"https://arxiv.org/abs/2610.02179v1","pdfUrl":"https://arxiv.org/pdf/2610.02179v1","doi":null},{"arxivId":"2610.02175","version":"v1","title":"Effective Resistance and Graph Neural Network Reliability in Tissue-Specific Interactomes","abstract":"Protein function annotation needs to know which predictions to distrust, not only what a model predicts. We ask whether tissue-specific interaction structure carries that information. Our candidate signal is effective resistance, used previously to relieve over-squashing by rewiring. Across 24 tissue-specific interactomes it is dominated by inverse degree, and the degeneration deepens as the co-expression filtered network grows, with a Spearman correlation of -0.955. The residual departure from that limit exceeds degree-preserving null graphs in all 24 networks. Controlling for predictive entropy, degree, annotation cardinality, local structure and feature-only difficulty, the residual explains additional per-node loss in 19 of 24 held-out networks once a permutation floor is subtracted, …","authors":["Jianru Shen"],"primaryCategory":"cs.LG","categories":["cs.LG","q-bio.MN"],"publishedAt":"2026-10-01T17:57:10Z","updatedAt":"2026-10-01T17:57:10Z","htmlUrl":"https://arxiv.org/abs/2610.02175v1","pdfUrl":"https://arxiv.org/pdf/2610.02175v1","doi":null},{"arxivId":"2610.02173","version":"v1","title":"Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair","abstract":"Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $λ$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(λ)=\\mathrm{own}_r+γ_rλ$. The slope $γ_r$ is a fixed coefficient that consistently in…","authors":["Areeb Ahmad","Pratinav Seth","Vinay Kumar Sankarapu"],"primaryCategory":"cs.LG","categories":["cs.LG","cs.CL"],"publishedAt":"2026-10-01T17:57:04Z","updatedAt":"2026-10-01T17:57:04Z","htmlUrl":"https://arxiv.org/abs/2610.02173v1","pdfUrl":"https://arxiv.org/pdf/2610.02173v1","doi":null},{"arxivId":"2610.02170","version":"v1","title":"Watch, Infer, Coordinate: Inferring Robot Partner Constraints for Zero-Shot Coordination","abstract":"Robots operating in the physical world will increasingly need to coordinate with other robots, particularly in manipulation tasks where an object may be too large or heavy for a single robot to carry alone. Physical limitations caused by hardware degradation or actuator faults can restrict the actions a robot can reliably execute, yet these limitations may be unknown to its partner. We study whether a helper can infer a robot partner's physical constraints from observing it coordinate with another robot, then use the inferred capability to coordinate with the same partner on a new task. This is difficult because a demonstration shows what the constrained robot did, but not what it could have done. In physically coupled tasks, the other robot may also compensate for its limitations, making…","authors":["Suyu Ye","Zheyuan Zhang","Vaishnav Tadiparthi","Hossein Nourkhiz Mahjoub","Ehsan Moradi Pari","Tianmin Shu","Homanga Bharadhwaj","Nakul Agarwal"],"primaryCategory":"cs.RO","categories":["cs.RO","cs.AI","cs.MA"],"publishedAt":"2026-10-01T17:56:04Z","updatedAt":"2026-10-01T17:56:04Z","htmlUrl":"https://arxiv.org/abs/2610.02170v1","pdfUrl":"https://arxiv.org/pdf/2610.02170v1","doi":null},{"arxivId":"2610.02163","version":"v1","title":"AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents","abstract":"Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fin…","authors":["Xuan Zhang","Longtao Zheng","Cunxiao Du","Bo An","Xin Dong"],"primaryCategory":"cs.CL","categories":["cs.CL"],"publishedAt":"2026-10-01T17:54:34Z","updatedAt":"2026-10-01T17:54:34Z","htmlUrl":"https://arxiv.org/abs/2610.02163v1","pdfUrl":"https://arxiv.org/pdf/2610.02163v1","doi":null},{"arxivId":"2610.02162","version":"v1","title":"World Observer: Joint Actor-Observer Generation for Persistent World Modeling","abstract":"How can a world model continuously observe regions beyond the actor's current view? Video world models simulate how an environment evolves from an agent's actions, yet remain actor-centric. Once an object leaves the actor's view, they lose direct evidence of its evolution, often failing to preserve its state and dynamics upon re-entry. To address this, we introduce World Observer, which decouples observing from acting by jointly generating a perspective actor for the agent-centric view with one or more panoramic observers that watch selected world regions. This allows objects that leave the actor's view to remain visually evolving in an observer, so their updated states are reflected when they re-enter. We ground the actor and observers by warping from a shared panoramic source for explic…","authors":["Hyunwook Choi","Dahyun Chung","Hyunsung Kim","Siyoon Jin","Jinhyeok Choi","Junyoung Seo","Seungryong Kim"],"primaryCategory":"cs.CV","categories":["cs.CV"],"publishedAt":"2026-10-01T17:53:20Z","updatedAt":"2026-10-01T17:53:20Z","htmlUrl":"https://arxiv.org/abs/2610.02162v1","pdfUrl":"https://arxiv.org/pdf/2610.02162v1","doi":null},{"arxivId":"2610.02161","version":"v1","title":"DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication","abstract":"Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then gene…","authors":["Hanchu Zhou","Dechen Gao","Hang Wang","Brendan Lynch","Boqi Zhao","Qiyao Ma","Raman Goyal","Junshan Zhang"],"primaryCategory":"cs.RO","categories":["cs.RO","cs.AI"],"publishedAt":"2026-10-01T17:53:09Z","updatedAt":"2026-10-01T17:53:09Z","htmlUrl":"https://arxiv.org/abs/2610.02161v1","pdfUrl":"https://arxiv.org/pdf/2610.02161v1","doi":null},{"arxivId":"2610.02160","version":"v1","title":"4Director: Controlling Video World Models with Rigid 3D Geometry","abstract":"Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We introduce 4Director, a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh and moved by one prescribed rigid transformation per frame. This representation provides an intuitive 3D control interface and prevents unobserved geometry from being regenerated independently in every frame. We render the controlled scene as a depth video and introduce a Motion Adapter that transforms this geome…","authors":["Wei Cao","Hao Zhang","Vikram Voleti","Yuqun Wu","Mallikarjun B R","Shimon Vainer","Mark Boss","Yaoyao Liu"],"primaryCategory":"cs.CV","categories":["cs.CV"],"publishedAt":"2026-10-01T17:52:58Z","updatedAt":"2026-10-01T17:52:58Z","htmlUrl":"https://arxiv.org/abs/2610.02160v1","pdfUrl":"https://arxiv.org/pdf/2610.02160v1","doi":null},{"arxivId":"2610.02159","version":"v1","title":"When Do Intrinsic Rewards Lead to Exploration?","abstract":"Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish conditions under which existing intrinsic rewards suc…","authors":["Scott W. Viteri","Laura Gomezjurado Gonzalez","Clark Barrett"],"primaryCategory":"cs.LG","categories":["cs.LG"],"publishedAt":"2026-10-01T17:52:41Z","updatedAt":"2026-10-01T17:52:41Z","htmlUrl":"https://arxiv.org/abs/2610.02159v1","pdfUrl":"https://arxiv.org/pdf/2610.02159v1","doi":null},{"arxivId":"2610.02158","version":"v1","title":"Muon meets Tamed Langevin: Momentum Preconditioning beyond Convex and gradient-Lipschitz Potentials","abstract":"We consider the problem of sampling from Gibbs distributions on matrix spaces whose potential energies are neither convex nor globally gradient-Lipschitz. We introduce a family of non-quadratic kinetic energies that lead to a new underdamped Langevin system with momentum preconditioning, in which the gradient of the kinetic energy acts as a smooth spectral taming of the momentum. We prove that, under these relaxed assumptions on the potential, the resulting dynamics leaves the target Gibbs measure invariant, and we establish exponential convergence to equilibrium in a weighted total variation distance. Finally, we show that the corresponding Euler-Maruyama discretization admits moment bounds that are uniform in time, without any modification of the potential gradient, which ensures the st…","authors":["Nikolaos Makras","Sotirios Sabanis"],"primaryCategory":"cs.LG","categories":["cs.LG","math.OC","math.PR","stat.ML"],"publishedAt":"2026-10-01T17:52:36Z","updatedAt":"2026-10-01T17:52:36Z","htmlUrl":"https://arxiv.org/abs/2610.02158v1","pdfUrl":"https://arxiv.org/pdf/2610.02158v1","doi":null},{"arxivId":"2610.02153","version":"v1","title":"MosaiChunk: Compositing Spatio-Temporal Memory for Autoregressive Video Generation","abstract":"Long-horizon autoregressive video generation is limited by a finite context window. When an object or scene falls out of context, its fine-grained visual details may be lost and difficult to recover upon reappearance. To retain access to such visual details, we introduce MosaiChunk, a spatio-temporal memory mechanism that composes a mosaic of selected historical key-value (KV) entries across space and time. Our approach is motivated by the observation that a frozen video generator can directly consume such non-contiguous historical KV and recover the corresponding visual content. We therefore keep the generator fixed and learn only a lightweight router that determines which historical sections to include in the mosaic under a fixed active-memory budget. We further introduce RememBench, a …","authors":["Yiwen Zhang","Haocheng Xi","Michael Tian-Yue Liu","Alexei A. Efros","Hadar Averbuch-Elor","Qianqian Wang","Haiwen Feng"],"primaryCategory":"cs.CV","categories":["cs.CV","cs.GR"],"publishedAt":"2026-10-01T17:51:05Z","updatedAt":"2026-10-01T17:51:05Z","htmlUrl":"https://arxiv.org/abs/2610.02153v1","pdfUrl":"https://arxiv.org/pdf/2610.02153v1","doi":null},{"arxivId":"2610.02150","version":"v1","title":"From Knowledge Access to Source Learning: Developing Source-Specific Competence","abstract":"Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propos…","authors":["Lucheng Fu","Kejing Xia","Yiyang Wang","Yiqiao Jin","Jinjin He","Xiyuan Yang","Haoxin Liu","Ye Yu"],"primaryCategory":"cs.CL","categories":["cs.CL","cs.AI","cs.LG"],"publishedAt":"2026-10-01T17:50:16Z","updatedAt":"2026-10-01T17:50:16Z","htmlUrl":"https://arxiv.org/abs/2610.02150v1","pdfUrl":"https://arxiv.org/pdf/2610.02150v1","doi":null},{"arxivId":"2610.02148","version":"v1","title":"Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation","abstract":"Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate embedding model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone's own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Tr…","authors":["Mohammed Irfan Kurpath","Jaseel Muhammad Kaithakkodan","Sahal Shaji Mullappilly","Ivan Laptev","Hisham Cholakkal"],"primaryCategory":"cs.CV","categories":["cs.CV"],"publishedAt":"2026-10-01T17:49:40Z","updatedAt":"2026-10-01T17:49:40Z","htmlUrl":"https://arxiv.org/abs/2610.02148v1","pdfUrl":"https://arxiv.org/pdf/2610.02148v1","doi":null},{"arxivId":"2610.02144","version":"v1","title":"Faynt: Scaling and Optimizing Policies for Competitive Melee","abstract":"We introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee, each controlling all 26 characters with a single checkpoint. After reinforcement learning (RL), the 10M wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases on their supported rosters, with a winning record against every release. These opponents retain 21- or 24-frame action delays; Faynt uses no added delay, and we have not isolated the effect of this difference. In a separate evaluation against a privately supplied zero-delay Slippi-AI model, the 10M wins all 68 games across two conditioning settings. We study architecture, optimization, scaling, and hyperparameter transfer to guide pretraining on approximately 840,000 human replays. P…","authors":["Ali Janati","Nikita Kuzmin","Rohit Swamy","Charles Niu"],"primaryCategory":"cs.LG","categories":["cs.LG"],"publishedAt":"2026-10-01T17:48:38Z","updatedAt":"2026-10-01T17:48:38Z","htmlUrl":"https://arxiv.org/abs/2610.02144v1","pdfUrl":"https://arxiv.org/pdf/2610.02144v1","doi":null},{"arxivId":"2610.02142","version":"v1","title":"Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models","abstract":"Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a 1,109M model (web-heavy multi-phase curriculum; 6B-token tool-SFT) share decoder, tokenizer, and special tokens, scoring almost identically on lenient tool-use metrics (B4: 0.660 vs. 0.650). Verbatim-reproduction checks on training examples separate them completely: the 600M emits valid tool calls with generalized arguments on 6/6 examples; the 1B does so on 0/6 across checkpoints. A first-token probe localizes the 1B's failure to a missing prior (prob. $10^{-4}$--$10^{-5}$ …","authors":["Juan S. Santillana"],"primaryCategory":"cs.CL","categories":["cs.CL"],"publishedAt":"2026-10-01T17:46:38Z","updatedAt":"2026-10-01T17:46:38Z","htmlUrl":"https://arxiv.org/abs/2610.02142v1","pdfUrl":"https://arxiv.org/pdf/2610.02142v1","doi":null},{"arxivId":"2610.02140","version":"v1","title":"Finetuning with Sampling: SFT Learns Better Than You Think","abstract":"Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution t…","authors":["Aayush Karan","Sitan Chen","Yilun Du"],"primaryCategory":"cs.LG","categories":["cs.LG","cs.AI","cs.CL"],"publishedAt":"2026-10-01T17:45:07Z","updatedAt":"2026-10-01T17:45:07Z","htmlUrl":"https://arxiv.org/abs/2610.02140v1","pdfUrl":"https://arxiv.org/pdf/2610.02140v1","doi":null},{"arxivId":"2610.02136","version":"v1","title":"MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI","abstract":"Unsupervised anomaly detection (UAD) methods for brain MRI are ranked by a single score, yet that score rests on choices that are rarely reported: how each anomaly map is aligned with the reference, how and on which data the threshold is set, and which false-positive budget, metric, aggregation and lesion definition are used. We present MIRTO, an evaluation protocol that makes these choices explicit and measures their effect. It gates the geometry of every comparison with a registration check and label-free diagnostics of known power, sets thresholds on validation data alone and reports the false-positive volume actually realised on test, repeats each comparison over 15,552 defensible evaluation pipelines, and attaches paired subject-bootstrap intervals with multiplicity control. Applied …","authors":["Negin Kafee Hernashki","Soumick Chatterjee"],"primaryCategory":"cs.CV","categories":["cs.CV","cs.AI","eess.IV","physics.med-ph"],"publishedAt":"2026-10-01T17:44:22Z","updatedAt":"2026-10-01T17:44:22Z","htmlUrl":"https://arxiv.org/abs/2610.02136v1","pdfUrl":"https://arxiv.org/pdf/2610.02136v1","doi":null},{"arxivId":"2610.02131","version":"v1","title":"Linear Programming Representations and Strongly Polynomial Algorithms for Robust Markov Decision Processes","abstract":"We study linear programming (LP) representations and strongly polynomial algorithms for robust Markov decision processes (RMDPs) with rational polyhedral state-action rectangular uncertainty in rewards and transitions. By encoding a finite sequence of robust policy-iteration steps, we construct a single LP whose optimal solutions recover the robust optimal value and all optimal stationary randomized policies. At fixed discount, the LP has polynomial dimension and encoding length and can be constructed in strongly polynomial time. We also develop a general complexity analysis of robust policy iteration that combines the cost of minimizing over uncertainty sets with the number of iterations needed to evaluate a policy. For a fixed discount factor, we use this analysis to improve the known c…","authors":["Han Zhong","Yinyu Ye"],"primaryCategory":"cs.LG","categories":["cs.LG","cs.DS","math.OC"],"publishedAt":"2026-10-01T17:41:29Z","updatedAt":"2026-10-01T17:41:29Z","htmlUrl":"https://arxiv.org/abs/2610.02131v1","pdfUrl":"https://arxiv.org/pdf/2610.02131v1","doi":null},{"arxivId":"2610.02128","version":"v1","title":"Sample complexity bounds for categorical Markov random fields via Discrete Diffusions","abstract":"Many applications in statistics, economics, and physics require sampling from high-dimensional categorical distributions with local dependence structures. Examples include finite memory language models, Ising and Potts systems in statistical physics and protein folding, etc. In modern machine learning, discrete diffusions have emerged as a flexible approach for sampling such data, with strong empirical performance. Motivated by this, we develop learning methods with end-to-end sample complexity bounds for discrete diffusion with uniform noising under local dependence, which we model through low order Markov random fields (MRFs). Our main technical insight is a new \\emph{pinning decomposition} of the discrete score. It shows that unlike in continuous diffusions, the score decomposes into c…","authors":["Shivam Kumar","Nabarun Deb"],"primaryCategory":"math.ST","categories":["math.ST","cs.LG","stat.ML"],"publishedAt":"2026-10-01T17:40:21Z","updatedAt":"2026-10-01T17:40:21Z","htmlUrl":"https://arxiv.org/abs/2610.02128v1","pdfUrl":"https://arxiv.org/pdf/2610.02128v1","doi":null},{"arxivId":"2610.02126","version":"v1","title":"Local Support Learning","abstract":"We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge…","authors":["Assaf Ben-Kish","Akarsh Kumar","James Glass","Raja Giryes"],"primaryCategory":"cs.LG","categories":["cs.LG","cs.AI"],"publishedAt":"2026-10-01T17:39:03Z","updatedAt":"2026-10-01T17:39:03Z","htmlUrl":"https://arxiv.org/abs/2610.02126v1","pdfUrl":"https://arxiv.org/pdf/2610.02126v1","doi":null},{"arxivId":"2610.02123","version":"v1","title":"Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation","abstract":"Mixture-of-Experts (MoE) architectures scale model capacity through sparse computation, routing each token through only a small subset of experts. In this work, we explore whether this sparsity gives rise to emergent intrinsic organization in multimodal MoEs. We find that experts develop strong semantic specialization across modalities and domains despite not being explicitly trained for modularity. Building on this structure, we introduce ExpertLens, a data-free method that identifies domain-specialized experts directly from pretrained model weights by decoding router weights into semantically meaningful vocabulary tokens. We leverage this specialization for efficient multimodal adaptation by selectively fine-tuning experts relevant to a target domain. Across math, medical, and remote se…","authors":["Damiano Marsili","Raphi Kang","Aditya Mehta","Pietro Perona","Georgia Gkioxari"],"primaryCategory":"cs.CV","categories":["cs.CV"],"publishedAt":"2026-10-01T17:37:13Z","updatedAt":"2026-10-01T17:37:13Z","htmlUrl":"https://arxiv.org/abs/2610.02123v1","pdfUrl":"https://arxiv.org/pdf/2610.02123v1","doi":null},{"arxivId":"2610.02122","version":"v1","title":"Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows","abstract":"Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We expo…","authors":["Gabriel Tomitsuka","Arman Raayatsanati","Emma Xing","Duke Gand","Joseph J Ma"],"primaryCategory":"cs.CL","categories":["cs.CL","cs.AI","cs.DB"],"publishedAt":"2026-10-01T17:35:44Z","updatedAt":"2026-10-01T17:35:44Z","htmlUrl":"https://arxiv.org/abs/2610.02122v1","pdfUrl":"https://arxiv.org/pdf/2610.02122v1","doi":null},{"arxivId":"2610.02117","version":"v1","title":"Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes","abstract":"On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query.…","authors":["Sophia Sirko-Galouchenko","Monika Wysoczanska","Andrei Bursuc","Nicolas Thome","Spyros Gidaris"],"primaryCategory":"cs.CV","categories":["cs.CV","cs.AI","cs.CL","cs.LG"],"publishedAt":"2026-10-01T17:34:10Z","updatedAt":"2026-10-01T17:34:10Z","htmlUrl":"https://arxiv.org/abs/2610.02117v1","pdfUrl":"https://arxiv.org/pdf/2610.02117v1","doi":null},{"arxivId":"2610.02116","version":"v1","title":"A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification","abstract":"A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreeme…","authors":["Javier Diaz Esteban-Herreros","David Muñoz-Valero","Raquel Martínez-España","Jose M. Juarez","Juan Moreno-Garcia"],"primaryCategory":"cs.AI","categories":["cs.AI","cs.CL"],"publishedAt":"2026-10-01T17:34:02Z","updatedAt":"2026-10-01T17:34:02Z","htmlUrl":"https://arxiv.org/abs/2610.02116v1","pdfUrl":"https://arxiv.org/pdf/2610.02116v1","doi":null},{"arxivId":"2610.02114","version":"v1","title":"Surface-volume self-supervised representation learning of brain MRI for genetic discovery","abstract":"Existing genome-wide association studies (GWAS) of brain imaging provide predefined or deep-learning-derived imaging phenotypes, yet these phenotypes come from either volumetric scans or cortical surface meshes, so each captures only part of the heritable variation in brain anatomy. Here we introduce MEVA (Mesh-Enhanced Volumetric Autoencoder), a self-supervised framework that encodes voxel-level image intensity together with cortical mesh geometry, including curvature and cortical thickness at each surface vertex, into one shared set of imaging features. Combining the mesh and volumetric inputs in MEVA yields modest performance gains in age and sex prediction over models that use either input alone. When these features serve as phenotypes for GWAS in the UK Biobank, they reveal more geno…","authors":["Tian Xia","Nuo Chen","Zihao Zhu","Huiwen Han","Ziqian Xie","Zhiwen Fan","Degui Zhi"],"primaryCategory":"cs.CV","categories":["cs.CV"],"publishedAt":"2026-10-01T17:32:52Z","updatedAt":"2026-10-01T17:32:52Z","htmlUrl":"https://arxiv.org/abs/2610.02114v1","pdfUrl":"https://arxiv.org/pdf/2610.02114v1","doi":null}],"summary":{"by_primary_category":{"cs.CV":18,"cs.CL":6,"cs.RO":4,"cs.AI":3,"cs.LG":18,"math.ST":1},"top_authors":[{"author":"Ivan Laptev","count":2},{"author":"Bo Liu","count":2},{"author":"Cristian McGee","count":2},{"author":"El Houcine Bergou","count":2},{"author":"Aritra Dutta","count":2},{"author":"Sophia Feldman","count":1},{"author":"Assaf Shocher","count":1},{"author":"Ramazan Fazylov","count":1},{"author":"Stamatis Lefkimmiatis","count":1},{"author":"Kaiyu Yue","count":1}]}}}