Features

11 Papers Accepted at NeurIPS 2026

Held on 6 – 12 December in Sydney, Australia, the Fortieth Annual Conference on Neural Information Processing Systems (NeurIPS 2026) is a premier global event for researchers, engineers, and students in artificial intelligence, machine learning, and computational neuroscience.

Congratulations to the following researchers from A*STAR Centre for Frontier AI Research (A*STAR CFAR) on having their papers accepted at NeurIPS 2026:

  • Prof Ivor Tsang, Director, A*STAR CFAR
  • Prof Ong Yew Soon, Chief Artificial Intelligence (AI) Scientist and Advisor
  • Dr Atsushi Nitanda, Group Lead, Optimisation in AI (OAI)
  • Dr Nancy Chen, Group Lead, Ethical & Trust AI (ETAI)
  • Mr Liu Zhengyuan, Team Lead, Ethical AI, ETAI
  • Dr Yin Haiyan, Team Lead, Agentic AI, Agentic Super Intelligence (ASI)
  • Dr Zhen Liangli, Team Lead, Trustworthy AI, ETAI
  • Dr Lyu Yueming, Senior Scientist
  • Dr Xu Xun, Senior Scientist
  • Dr Qu Bohao, Scientist
  • Dr Tan Zhi Xuan, Scientist
  • Dr Tanya Veeravalli, Scientist
  • Dr Yu Xingrui, Scientist

List of accepted papers:

  1. NesyProAct: Proactive Neural-Symbolic Control for Web Agents
    Tianyi Tang, Keyi Xiang*, Jie-Jing Shao***, Yueming Lyu, Ivor Tsang, Yew-Soon Ong, Haiyan Yin

    NesyProAct introduces proactive neural-symbolic control that integrates grounding and execution verification directly into web-agent decision making. It improves robustness and generalisation by detecting and repairing execution errors during interaction.

  2. An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
    Mingzhong Sun, Teresa Yeo, Armando Solar-Lezama, Tan Zhi-Xuan

    Large reasoning models (LRMs) are trained to produce reasoning, but are they good at evaluating reasoning? We find frontier LRMs display surprising and substantial failures at evaluating flawed reasoning on simple math problems, dropping to as low as 50% accuracy. We also find evidence that this is driven by the presence of a valid answer in the solution despite the reasoning being invalid.

  3. Slowly Annealed Langevin Dynamics: Theory and Applications to Training-Free Guided Generation
    Atsushi Nitanda, Dake Bu*, Yueming Lyu, Tanya Veeravalli

    This work proposes VA-SALD, a new theoretically grounded framework for training-free guided generation with pretrained generative models.

  4. FlowLeak: Coverage-Guided Extraction of Dynamic Workflows in LLM-Based Multi-Agent Systems
    Zhiyao Ren, Siyuan Liang, Yibing Zhan, Jun Long Tan, Xiaobing Sun, Liangli Zhen, Baosheng Yu, Dacheng Tao

    LLM-based multi-agent systems (MAS) increasingly rely on dynamic workflows whose internal structure is hidden from users. We present FlowLeak, a black-box extraction attack that, using only query access, recovers a target system's hidden agents, their private configurations, tool descriptions, and communication topology.

  5. Rethinking Bayesian Optimisation for Co-Optimising LLM Training Configurations
    Zhiliang Chen*, Alfred Wei Lun Leong, Shao Yong Ong, Apivich Hemachandra, Gregory Kang Ruey Lau, Chuan-Sheng Foo, Zhengyuan Liu, Nancy F. Chen, Bryan Kian Hsiang Low

    This work introduces JoBS, a BO-based approach that co-optimises data and model configurations without requiring a full training run at every iteration. JoBS allocates an initial portion of the optimisation budget to learn a scaling-law-inspired performance predictor that estimates fully trained LLM performance from only a small number of training steps. The remaining budget is then used to run BO with this predictor, eliminating the need for full training runs and enabling JoBS to explore significantly more data and model configurations within the same budget.

  6. Denoising Time Matters: Diverse Generation in Diffusion Language Models
    Jingxuan Wu, Zhenglin Wan*, Yuzhe Yang, Yiqiao Huang, Chubin Zhang, Xingrui Yu, Ivor Tsang, Yang You

    This paper introduces TAPS, a training-free sampling strategy that encourages early semantic branching through time-aware perturbations, improving diffusion language models’ semantic and lexical diversity while preserving generation quality and reasoning ability with negligible overhead.

  7. Calibration Is Not Control: Intervention Advantage for LLM-Agent Oversight
    Chubin Zhang, Zhenglin Wan*, Xingrui Yu, Jingxuan Wu, Qi Wen, Pengfei Zhou, Wangbo Zhao, Ivor Tsang

    This paper reframes LLM-agent oversight around intervention advantage—the expected utility gain from intervening—and demonstrates regime-dependent reductions in control regret that risk-score recalibration alone cannot achieve.

  8. Resilient Latent Readouts for Long-Context Question Answering
    Jingyi Liao, Wenhao Sun, Yiting Li, Zhao Jin, Khin Mi Mi Aung, Xun Xu, Zhuoyi Lin, Rong-Cheng Tu, Dacheng Tao

    We propose LIRA, a latent-memory reader that exposes full-prefix contextualised states rather than re-encoded text. It amortises one full-prefix pass into a reusable latent store; for each query, localises candidate positions with calibrated endogenous attention, repairs incomplete support with coverage-aware planning, and packs selected states into a position-consistent readout for autoregressive generation.

  9. Flow-Direct: Feedback-Efficient and Reusable Guidance for Flow Models via Non-Parametric Guidance Field
    Kim Yong Tan***, Yueming Lyu, Ivor Tsang, Yew-Soon Ong

    This paper proposes Flow-Direct, a training-free framework that guides pre-trained flow models using a non-parametric guidance field. This field is analytically derived from the log-density ratio between the base and reward-weighted target distributions and is implemented as a non-parametric estimator, providing query-efficient and reusable guidance without requiring repeated queries.

  10. SeqDiCO: Sequence-oriented Diffusion for Scale-Generalisable Neural Combinatorial Optimisation
    Yu Wang, Qiaolin Lu, Yang Wu, Bohao Qu, Shirui Pan, Yi Chang, Chengqi Zhang

    This paper proposes SeqDiCO, a sequence-oriented diffusion framework for scale-generalisable neural combinatorial optimisation. It learns a local denoising operator over fixed-endpoint sequence subproblems, enabling reusable local reconstruction across problem sizes and allowing the same model to construct and refine large-scale routing solutions beyond the instance sizes seen during training.

  11. A Primal-dual Approach for Semi-Infinitely Constrained Reinforcement Learning
    Di Wang, Liangyu Zhang, Haishan Ye, Guang Dai, Ivor Tsang

    SI-PDPO addresses reinforcement learning with a continuum of constraints using infinite-dimensional primal-dual optimisation. It establishes strong duality, regularises dual variables, and applies first-order policy updates, avoiding costly inner-loop optimisation while reducing conservatism and consistently outperforming existing primal-type methods.

Learn more about NeurIPS 2026.

* denotes former CFAR student
** denotes former CFAR researcher
*** denotes current CFAR student
(accurate at time of posting)