1. X
  2. Gedas Bertasius
Log inSign up
Gedas Bertasius
594 posts
user avatar

Gedas Bertasius

@gberta227
Assistant Professor at UNC, previously a postdoc at Meta AI, PhD from UPenn, video understanding, multimodal AI, a basketball enthusiast.
Chapel Hill, NC
gedasbertasius.com
Joined June 2020
1,409
Following
1,628
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Gedas Bertasius
    @gberta227
    Jun 17
    Sharing my CVPR 2026 talk from the Vision for Intelligent Task Assistants workshop: "From Perception to Agency: The Cognitive Stack for Video Task Assistants." It covers our SVI-Bench project (svi-bench.github.io) plus a video+robotics project we'll release soon. w/
    00:00
  • user avatar
    Gedas Bertasius
    @gberta227
    Jul 23
    Our proposed VideoTreeSearch achieves state-of-the-art results across all grounded long-video QA benchmarks while also being more efficient than prior agentic video methods. We also show that we can achieve strong temporal grounding results even if we don't use any manually
    user avatar
    Ce Zhang
    @cezhhh
    Jul 23
    Excited to share our new work, “Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA.” Grounded LVQA requires a model to answer a question about a long video and localize the short interval supporting its answer. We introduce VideoTreeSearch (VTS), which
  • user avatar
    Gedas Bertasius
    @gberta227
    Jul 1
    ICML 2026 is coming up, which means it's been 5 years since I presented TimeSformer at ICML 2021 (arxiv.org/pdf/2102.05095). It turned out to be the most cited paper I've published, and it's still the only ICML paper I have. A bit of the story behind it. 🧵👇
  • user avatar
    Gedas Bertasius
    @gberta227
    Jun 26
    We just released WatchAct, a benchmark for behavior-grounded robot manipulation (covered in the talk below). The robot has to watch a human video, infer what was done, make a plan, and then execute it. The best VLM (Gemini-3.1-Pro) only reaches 36.8% planning success rate, so
    user avatar
    Gedas Bertasius
    @gberta227
    Jun 17
    Sharing my CVPR 2026 talk from the Vision for Intelligent Task Assistants workshop: "From Perception to Agency: The Cognitive Stack for Video Task Assistants." It covers our SVI-Bench project (svi-bench.github.io) plus a video+robotics project we'll release soon. w/
    00:00
  • user avatar
    Gedas Bertasius
    @gberta227
    Jun 3
    Excited to give an invited talk at the VITA workshop @ CVPR 2026 (Vision for Intelligent Task Assistants) tomorrow: "From Perception to Agency: The Cognitive Stack for Video Task Assistants" — what assistants need beyond perception, and how far we are from it. Wed, June 3, Room

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.