PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

Programmatic Video Prediction Using Large Language Models

View Full Paper
JTJing TangStatistics FinlandKEK. V. EllisRutherford Appleton LaboratorySLSuhas LohitMitsubishi Electric (United States)

Key Points

  • ProgGen demonstrates superior video frame prediction, outperforming existing techniques in multiple environments.
  • The method utilizes large language models to create human-interpretable states for predicting future video frames.
  • Empirical evaluations in PhyWorld and Cart Pole show significant enhancements in video generation tasks.
  • Counter-factual reasoning features enable further interpretability and application versatility in video-related tasks.

Abstract

The task of estimating the world model describing the dynamics of a real world process assumes immense importance for anticipating and preparing for future outcomes. For applications such as video surveillance, robotics applications, autonomous driving, etc. this objective entails synthesizing plausible visual futures, given a few frames of a video to set the visual context. Towards this end, we propose ProgGen, which undertakes the task of video frame prediction by representing the dynamics of the video using a set of neuro-symbolic, human-interpretable set of states (one per frame) by leveraging the inductive biases of Large (Vision) Language Models (LLM/VLM). In particular, ProgGen utilizes LLM/VLM to synthesize programs: (i) to estimate the states of the video, given the visual context (i.e. the frames); (ii) to predict the states corresponding to future time steps by estimating the transition dynamics; (iii) to render the predicted states as visual RGB-frames. Empirical evaluations reveal that our proposed method outperforms competing techniques at the task of video frame prediction in two challenging environments: (i) PhyWorld (ii) Cart Pole. Additionally, ProgGen permits counter-factual reasoning and interpretable video generation attesting to its effectiveness and generalizability for video generation tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tang et al. (2025) studied this question.

synapsesocial.com/papers/68f5a78aab63786de5b46136https://doi.org/10.48550/arxiv.2505.14948
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation2025
  2. 2Video-Language Models as Flexible Social and Physical Reasoners2024
  3. 3Video as the New Language for Real-World Decision Making2024 · 2 citations
  4. 4Can World Models Benefit VLMs for World Dynamics?2025
  5. 5ViPro: Enabling and Controlling Video Prediction for Complex Dynamical Scenarios using Procedural Knowledge2024