1. X
  2. Yuxiang Wei
Log inSign up
Yuxiang Wei
173 posts
user avatar
Yuxiang Wei
@YuxiangWei9
AI Research @Meta MSL. PhD @siebelschool
Mountain View, CA
yuxiang.cs.illinois.edu
Joined October 2021
372
Following
2,563
Followers
RepliesRepliesMediaMedia
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Don't miss what's happening
People on X are the first to know.
Log inSign up

New to X?

Sign up now to get your own personalized timeline!

Create account

By signing up, you agree to the Terms of Service and Privacy Policy, including Cookie Use.

  • Pinned
    user avatar
    Yuxiang Wei
    @YuxiangWei9
    Dec 23, 2025
    Software agents can self-improve via self-play RL Introducing Self-play SWE-RL (SSR): training a single LLM agent to self-play between bug-injection and bug-repair, grounded in real-world repositories, no human-labeled issues or tests. 🧵
    528K
  • user avatar
    Yuxiang Wei
    @YuxiangWei9
    Jul 9
    let’s push the horizon longer 🚀
    user avatar
    Alexandr Wang
    Meta
    @alexandr_wang
    Jul 9
    1/ muse spark 1.1 is an industry-competitive agentic and coding model. across many agentic evals it rivals gpt-5.5 and opus-4.8. available now through the new meta model api and in meta ai. 🧵
    376
  • user avatar
    Yuxiang Wei
    @YuxiangWei9
    Apr 30
    Accepted to ICML 2026! Big thanks to all the collaborators 🎉
    user avatar
    Yuxiang Wei
    @YuxiangWei9
    Dec 23, 2025
    Software agents can self-improve via self-play RL Introducing Self-play SWE-RL (SSR): training a single LLM agent to self-play between bug-injection and bug-repair, grounded in real-world repositories, no human-labeled issues or tests. 🧵
    4.8K
  • user avatar
    Yuxiang Wei
    @YuxiangWei9
    Apr 16
    > Aggressive cheating attempts ...attempt to write their own verifiers to check if their cheating attempts would pass an anti-cheating verifier Super interesting work. Also wild that the agents are relentlessly "creative" at gaming the eval.
    user avatar
    Justus Mattern
    Proximal
    @MatternJustus
    Apr 16
    Introducing FrontierSWE, an ultra-long horizon coding benchmark. We test agents on some of the hardest technical tasks like optimizing a video rendering library or training a model to predict the quantum properties of molecules. Despite having 20 hours, they rarely succeed
    1.1K
  • user avatar
    Yuxiang Wei
    @YuxiangWei9
    Jan 26
    Wow this is smart: per-token feedback, no manual reward design, applicable to partial trajectories in principle, all pain points of outcome RL. so curious to see it being applied at scale or combined with existing RL algos
    user avatar
    Siyan Zhao
    @siyan_zhao
    Jan 22
    Introducing 💡On-Policy Self-Distillation💡, a simple method that enables LLM to teach itself with dense per-token feedback on its own on-policy generations—achieving 4-8x more token efficiency vs. GRPO and outperforming both GRPO and SFT/Off-Policy Distillation. Key insight:
    27K