aqDeveloper environment for foundation models

World Models

A collection of Aquin work on world models: learned dynamics, checkable environments, and interventions you can score against ground truth. Training runs use aq.

01 H3 Avatar

A talking-avatar LoRA on MiniMax-H3 Ref2VA: one character image plus dialogue becomes a clip with model-generated speech, lip sync, and light motion. The recipe is locked behind a local UI so users pick a reference, a line, direction, and a background — not checkpoints or audio modes.

H3 Avatar demo frame: a character portrait used as the single-image reference for talking-avatar generation
Demo reference frame from the locked audio-v2 LoRA. A short video clip will replace this when available.

Base model is MiniMax-H3 Ref2VA with a rank-16 / alpha-32 adapter on to_qkv, to_out.0, linear_1, and linear_2 (~10.7M trainable params). The locked checkpoint is h3-avatar-audio-v2-step1000: 481 accepted transcript-conditioned clips, 385 train, 1,000 optimizer steps at 384×384 with a 56-frame cache — about 2.6 passes over the train set.

Why transcript conditioning

An earlier audio pass improved lip timing but not reliable English: the data only carried a language tag, not the spoken words. Audio-v2 runs ASR, keeps short English phrases with usable confidence, and puts the transcript into the prompt so the model is conditioned on what the character should say. Step 1000 was locked before further joint video/audio training could drift already-good visuals.

BaseMiniMax-H3 Ref2VA
AdapterLoRA r16 / α32 · 10.7M trainable
Train set385 of 481 transcript-conditioned clips
Cache384×384 · 56 frames · video + audio + text
Locked step1000 · audio loss scale 2.0 · lr 5e−5

Training and eval live in the repo with an AQ recipe handoff. Serving keeps the H3 base hot; the adapter itself is small. A Modal warm endpoint is the reference path; production should use a dedicated inference host.

02 Mario World Model

An action-conditioned JEPA world model for Super Mario Bros, with LeJEPA/SIGReg in latent space and probe gates that catch planning failures early. Trained with aq as an interpretability-in-the-loop run: a model that predicts perfectly can still fail to plan, and the gate stops the train before that waste lands.

Super Mario Bros World 1-1 episode with live x position and macro action overlay
Live episode strip with RAM-backed x and macro action labels. Ground truth grades the frozen representation; it is never a training input.

A 9.8M-parameter JEPA on 8,000 episodes of World 1-1 (4.88M observations), predicting five steps ahead in a 192-dimensional latent. Mario is the vehicle; the reusable gate harness is the product. Prior work (LeMario) only found that height was missing after planning failed (y-probe R² = 0.188). Here the same failure is caught automatically at epoch 2 when any probe falls under threshold.

What SIGReg needs in this setting

SIGReg transfers out of image self-supervised learning into an action-conditioned video world model, but not unmodified. Report T / null(n), not raw SIGReg: under the Gaussian null the statistic floors at 0.51914 / n, so λ is not comparable across batch sizes without that ratio. Collapse-detection power scales with batch size, while video memory pushes the batch down. A trailing BatchNorm cuts the collapse signal about 25×. Temporally correlated windows break the iid null unless the statistic is scored per frame position. And in this setting the stop-gradient was still required: collapse remains the prediction objective's global optimum unless the target path cannot be chased.

Anti-collapse is not usable state

Preventing collapse did not preserve the state a planner needs. A SIGReg-only model stayed non-degenerate and still could not report Mario's height. An isotropic-Gaussian embedding is a floor, not a sufficient condition, which is why the run gates on a probe of control-relevant state per checkpoint.

ProbeSIGReg only+ aux heads
Height (y)−0.3830.796
Position (x)0.4510.870
Camera (scroll)0.4640.878
Death within 5 (AUC)0.6280.927

With aux heads, in-distribution height probe R² reaches 0.943 against LeMario's 0.188 on about 4% of the planned schedule. Numbers above are from a 4,000-step run on one level; the aux column is trained to make height decodable, so it shows a cheap head fixes the problem, not that the self-supervised objective alone learned a better representation. Full write-up and limitations live in the repo docs.

03 Pixel-Art-GWM

A playable 21 MB action-conditioned world model of a 2D pixel-art character, trained from scratch on self-generated data with aq, and a measured failure mode when the agent is about 1% of the frame. 4.4M parameters, 17.7 MB on disk, about 100 fps on a laptop GPU.

Pixel-art platformer: imagined rollout strip above, ground-truth versus reconstruction grid below, showing a small orange agent across platforms
Top: action-conditioned imagined rollout. Bottom: ground truth versus reconstruction over eight frames. The agent is a few pixels; whole-frame metrics can look fine while it vanishes.

The finding

Reconstruction objectives averaged over a frame are not neutral about what they preserve. When the controllable agent occupies 1.3% of the observation, such an objective can be reduced by discarding it, and standard metrics improve while it happens. A reconstruction containing no agent at all still scores 0.987 whole-frame accuracy.

TokenizerWhole-frameAgent-only
Adapted general-purpose, step 1,0000.42450.0151
Adapted general-purpose, step 3,0000.76560.0010 ↓
This work0.99980.9994

Two changes fix it: spatially structured latents (the agent occupies specific cells and cannot be marginalised) and per-pixel classification over the known 21-colour palette (cross-entropy cannot blur; class weights can express relevance).

Results

Tokenizer2,942,173 params · 0.9994 agent accuracy · 24× compression
Dynamics1,459,080 params · 0.885 rollout agent accuracy
Drift over 8 imagined frames0.999 → 0.995 (none measurable)
Action fidelityleft −16 px, right +35 px, idle −0.00 px
Speed101 fps M2 GPU · 26 fps M2 CPU · 93 fps A100
Training compute2.3 GPU-hours on aq

Dataset: 2,025,385 frames across 35,000 episodes, generated rather than scraped. Released with the model and a full report.

Not sure if Aquin is right for you?