Skip to content

Latest commit

 

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

World Models, Three Ways: Koopman Operators, Selective SSMs & Learned Video Dynamics

Section 1

This repository contains implementations and experiments for Koopman-based Robust Model Predictive Control (RMPC) applied to a 2R2C building thermal model. The core idea is to approximate nonlinear thermal dynamics with a linear predictor in a lifted (Koopman) space, enabling the use of linear MPC techniques while retaining nonlinear system fidelity.

The repository includes two complementary implementations:

  • A MATLAB implementation based on classical Extended Dynamic Mode Decomposition (EDMD) with handcrafted basis functions.
  • A Python implementation using Deep Koopman learning, where the lifting functions are learned via neural networks.

Both implementations follow the same conceptual pipeline and are intended to be compared side-by-side.


Problem Overview

The system of interest is a 2R2C thermal building model, commonly used to describe indoor temperature dynamics under heating/cooling inputs and environmental disturbances. The dynamics are nonlinear due to bilinear and quadratic terms arising from heat transfer and control interactions.

The control objective is to:

  • Accurately predict thermal states over a finite horizon
  • Enable robust MPC design using a linear predictor in lifted space

Koopman operator theory provides the bridge between nonlinear dynamics and linear control.


Koopman Modeling Approach

Both MATLAB and Python solutions follow the same logic:

  1. Nonlinear system dynamics ( x_{k+1} = f(x_k, u_k) )

  2. Data collection Simulate multiple trajectories under random excitation to collect ( (x_k, u_k, x_{k+1}) ) samples.

  3. Lifting to Koopman space Map the original state to a higher-dimensional space: [ z_k = \psi(x_k) ]

  4. Linear dynamics in lifted space [ z_{k+1} = A z_k + B u_k ]

  5. State reconstruction [ x_k \approx C z_k ]

The key difference lies in how the lifting function ( \psi(\cdot) ) is constructed.


MATLAB Implementation (EDMD)

The MATLAB implementation uses explicit basis functions:

  • Radial Basis Functions (RBFs)
  • Original state components

Key characteristics:

  • Deterministic lifting map
  • Linear regression to identify (A), (B), and (C)
  • Inspired by Koopman MPC formulations by Korda & Mezic

This approach is transparent and interpretable, but requires manual design of basis functions.


Python Implementation (Deep Koopman)

The Python implementation replaces handcrafted bases with learned lifting functions:

  • A neural network encoder learns ( z = \psi_\theta(x) )
  • Linear layers represent the Koopman operator ((A, B))
  • A decoder maps lifted states back to physical space

Training objectives include:

  • One-step prediction loss in state space
  • Koopman consistency loss in lifted space

This formulation preserves the Koopman structure while increasing expressive power and reducing manual feature engineering.


Results

Prediction Performance

The Deep Koopman model achieves:

  • Very low one-step prediction error
  • Stable short-horizon rollouts
  • Gradual divergence over long horizons, consistent with Koopman eigenvalue sensitivity

A representative rollout comparison between the true nonlinear system and the Koopman predictor is shown below.

Figure: Time-domain comparison of true states and Koopman-predicted states for both state variables.


Key Observations

  • Both EDMD (MATLAB) and Deep Koopman (Python) successfully linearize the dynamics in lifted space.
  • Deep Koopman provides greater flexibility by learning task-specific basis functions.
  • Long-horizon prediction error is primarily driven by small Koopman eigenvalue inaccuracies, a known theoretical limitation.

Section 2: SSMs

Selective-SSM ("Mamba-ish") + Koopman latent linearity + RL (A2C) on POMDP CartPole


Section 3: Grid-World & Maze Video World Models

A third, complementary line of work moves away from Koopman/SSM latent dynamics and instead learns pixel-space world models: a CNN maps (current image, action) → next image directly, with no explicit state representation or dynamics equations. Planning is then done by rolling the learned model forward and greedily minimizing predicted distance-to-goal.

Both experiments share the same recipe:

  • Render the environment as a multi-channel semantic image (one channel each for obstacles/walls, start, goal, agent).
  • Train a small CNN encoder–decoder (TinyCNNWorldModel / TinyMazeWorldModel) on randomly sampled (state, action, next_state) transitions, using a positive-pixel-weighted BCE loss (agent/goal pixels are rare, so under-weighting them collapses training).
  • At test time, decode the predicted agent position from the model's imagined next frame, and greedily pick the action whose imagined outcome is closest (Manhattan distance) to the goal, with a penalty for revisiting cells.

Grid World (WM/Video_Models/GridWM.py): 9×9 grid with block obstacles, start (0,0), goal (6,6).

Figure: Rollout of the learned CNN world model driving greedy planning — the agent reaches the goal in 12 steps, navigating around obstacles using only image-conditioned next-state predictions (no ground-truth dynamics).

Maze (WM/Video_Models/Video_GridWM.py): 7×7 maze with explicit interior walls, start (6,0), goal (0,6). Includes a BFS oracle (bfs_oracle_path) for comparison against the true shortest path.

Figure: The learned model correctly predicts the first couple of moves, but the greedy planner then stalls at (4,0) and never reaches the goal. Because the policy is a memoryless one-step-lookahead over the Manhattan-distance heuristic, it has no way to backtrack out of a local optimum once a wall blocks the locally-best action — unlike the open grid-world case, tight maze corridors expose this limitation of greedy planning over a learned model.

A related, unexecuted script, WM/Video_Models/VM_Planning.py, sketches a larger video-world-model pipeline (VQ-VAE tokenizer + SSM latent dynamics + MPC) for MetaWorld robotic manipulation, following iVideoGPT; it depends on external packages (metaworld, mujoco) not vendored here and has no results yet.


Repository Structure

.
├── results/                        # Experimental results and figures
│ ├── Koopman_plot1.png             # True vs Koopman rollout comparison
│ ├── Grid_World_Model.png          # Grid-world CNN world-model rollout
│ └── Maze_WM_Video.png             # Maze CNN world-model rollout
│
├── src/                             # Koopman / SSM source code
│ ├── 2R2C_Lifted_EDMD.m            # MATLAB EDMD-based Koopman identification (RBF lifting)
│ ├── deep_koopman_2RRC.py          # Python Deep Koopman implementation (neural lifting)
│ ├── deep_koopman_cartpole.py      # Deep Koopman on CartPole
│ ├── ssm_world_model_koopman_rl.py # Selective-SSM + Koopman + A2C on POMDP CartPole
│ └── simple_WM.py
│
├── WM/Video_Models/                 # Pixel-space (image/video) world models
│ ├── GridWM.py                     # Grid-world CNN world model + greedy planning
│ ├── Video_GridWM.py               # Maze CNN world model + greedy planning + BFS oracle
│ └── VM_Planning.py                # VQ-VAE + SSM + MPC sketch for robotic manipulation
│
├── README.md

References

  • M. Korda, I. Mezic, Linear predictors for nonlinear dynamical systems: Koopman operator meets model predictive control, Automatica, 2018.
  • S. Lusch, J. N. Kutz, S. L. Brunton, Deep learning for universal linear embeddings of nonlinear dynamics, Nature Communications, 2018.

Notes

This repository is intended for research and educational purposes, particularly for studying Koopman-based modeling and control in building energy systems and related thermal processes.

📌 For the latest papers, benchmarks, datasets, and state-of-the-art research on World Models, particularly Spatial and 3D World Models, please visit our companion repository: Awesome Spatial and 3D World Models.

Releases

Packages

Used by

Contributors

Languages