1. X
  2. Xiao Yu @ ICLR2026
Log inSign up
Xiao Yu @ ICLR2026
41 posts
user avatar
Xiao Yu @ ICLR2026
@xy2437
PhD student at @Columbia, collaborating with @MSFTResearch | Previsouly undergrad at @CUSEAS
jasonyux.com
Joined January 2023
128
Following
88
Followers
RepliesRepliesMediaMedia
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Don't miss what's happening
People on X are the first to know.
Log inSign up

New to X?

Sign up now to get your own personalized timeline!

Create account

By signing up, you agree to the Terms of Service and Privacy Policy, including Cookie Use.

  • user avatar
    Xiao Yu @ ICLR2026
    @xy2437
    Feb 9
    For AI agents to scale beyond narrow tasks, they need to learn how the world works from their own interactions — in a self-supervised way, without relying on expert data or rewards. We introduce Reinforcement World Model Learning (RWML), a self-supervised method that trains
    119
  • user avatar
    Xiao Yu @ ICLR2026
    @xy2437
    Jan 26
    Accepted by #ICLR2026 🎉See ya in Brazil!
    user avatar
    Xiao Yu @ ICLR2026
    @xy2437
    Oct 15, 2025
    Why can (V)LMs agents ace coding and math, yet struggle so badly in more complex environments like computer or phone use? 🤔 We find that one key factor lies in models' ability to understand and *simulate* the environment’s dynamics — and propose **Dyna-Mind** to address this!
    423
  • user avatar
    Xiao Yu @ ICLR2026
    @xy2437
    Oct 15, 2025
    Why can (V)LMs agents ace coding and math, yet struggle so badly in more complex environments like computer or phone use? 🤔 We find that one key factor lies in models' ability to understand and *simulate* the environment’s dynamics — and propose **Dyna-Mind** to address this!
    2.7K
  • user avatar
    Xiao Yu @ ICLR2026
    @xy2437
    Feb 12, 2025
    Thanks @MSFTResearch for all the support! Excited to share that our work is also accepted by #ICLR2025! - Paper: arxiv.org/abs/2410.02052 - Code: github.com/microsoft/ExACT - Website: agent-e3.github.io/ExACT/
    user avatar
    Microsoft Research
    Microsoft
    @MSFTResearch
    Feb 11, 2025
    ExACT combines Reflective-MCTS and Exploratory Learning to improve AI agents' decision-making, enabling test-time compute scaling. Learn how these methods help agents refine strategies for state-of-the-art performance and improved computational efficiency: msft.it/6017Um7P3
    00:00
    2.5K
  • user avatar
    Xiao Yu @ ICLR2026
    @xy2437
    Oct 22, 2024
    Website link was broken… it’s now fixed!
    user avatar
    Xiao Yu @ ICLR2026
    @xy2437
    Oct 14, 2024
    To effectively solve modern computer tasks, AI agents need to be able to strategically explore the environment and efficiently learn from past interactions. We present R-MCTS and Exploratory Learning for building o1-like models for agentic applications. Our GPT-4o powered R-MCTS
    00:00
    212