Fully continuous multimodal pretraining

Multimodal Flow

Unified Flow Modeling of Language and Vision in Embedding Spaces

Language and vision generated by one continuous vector field — no visual quantization, no modality-dependent objective, no second sampler.

Authors

Hongyuan Tao1

Xinggang Wang1,✉

Lianghui Zhu1

Yongkang Li1

Yunchao Wei2

 

Bin Feng1

Shaoyu Chen3

Qian Zhang3

Chang Huang3

Kai Yu3

Affiliations

1Huazhong University of
Science and Technology

2Beijing Jiaotong University

3Horizon Robotics

Resources

arXiv:2609.40362

Paper (PDF)

Code repository

Model weights

Contact

hongyuantao@hust.edu.cn

xgwang@hust.edu.cn

TL;DR Multimodal Flow removes discrete tokens from the generative process entirely. Text blocks and images both become continuous hyperchunks, and a single Flow Matching vector field generates each chunk conditioned on the chunks before it. Same objective, same sampler, same backbone — for language and for vision. Tasks differ only in the order and modality of the chunks, never in the model — which is also why the formulation does not stop at two modalities.

The compromise we wanted to avoid

Unified multimodal models have to answer an awkward question: language is an ordered sequence of discrete symbols, images are dense continuous grids. If one model is going to generate both, something has to give.

Today's systems give in one of two places. Fully discrete models quantize images into visual tokens so both modalities live in one categorical sequence — elegant, but now visual fidelity is capped by the tokenizer, and detail it discards is simply unavailable to the model downstream. Hybrid discrete–continuous models keep visual states continuous and bolt diffusion or flow onto an autoregressive language model — fidelity is preserved, but the backbone is now serving two different objectives and two different sampling procedures that happen to share parameters.

Neither gives you all three of the properties you would actually want. Pick a paradigm:

Property Fully discrete Hybrid Multimodal Flow
Continuous visual statesno quantization bottleneck ✗ ✓ ✓
Shared generative objectiveone loss for both modalities ✓ ✗ ✓
Shared sampling procedureone decoding mechanism ✓ ✗ ✓

Click a paradigm to isolate its column. Fully discrete models buy a shared process by paying with a quantization bottleneck; hybrids buy fidelity by paying with two objectives.

Three multimodal modeling paradigms: fully discrete, hybrid discrete–continuous, and Multimodal Flow.
Three ways to build a unified model. Blue denotes discrete states, green continuous ones. Only the fully continuous formulation keeps continuous visual states while sharing a generative objective and a sampling procedure across modalities.

So the question this work asks is simple: can language and vision share one generative process without being forced into one homogeneous representation? Our answer has four moving parts — a generative unit, an ordering, an objective, and one split inside the backbone — plus two consequences that fall out of them for free. The rest of this page walks through them in that order.

1 The generative unit: hyperchunks

A token is the wrong unit for this. Generating an image one patch at a time throws away the reason continuous visual models work; generating a whole paragraph in one shot throws away the sequential structure of language. We want something in between, and we want the same kind of thing for both modalities.

A hyperchunk (or just chunk) is one atomic act of generation. Text is cut into contiguous blocks of B = 8 tokens; each block is one chunk. An image is one chunk — all 256 positions of it. Inside a chunk, positions are modeled jointly; between chunks, the model is causal. That one asymmetry is what lets the same model be sequential at the level of events and parallel inside them.

How a chunk is built
raw text
↓ cut into blocks of 8 tokens
text chunks
↓ frozen T5-small encoder, then normalize
continuous state
cbℓ  ∈  ℝ8 × 512 token order preserved inside the chunk
raw image
Example image: a nebula. 224 × 224
↓ frozen SigLIP2-so400m, patch 14
one visual chunk

a 16 × 16 grid of patch embeddings — kept together as a single chunk

↓ normalize per position & channel
continuous state
cv  ∈  ℝ256 × 1152 spatial grid preserved inside the chunk

Both encoders — and both decoders — stay frozen throughout pretraining and finetuning. They are the interface to text and pixels; the only thing that learns is the flow backbone operating on continuous states. When we say the model is fully continuous, this is precisely what we mean: the trainable generative process never touches a discrete symbol.

2 Chunk-causal factorization

Once everything is a chunk, a multimodal sequence is just an ordered list 𝒞 = (c1, …, cK), and the joint distribution factorizes the obvious way:

pθ(𝒞)  =  ∏k=1…K   pθ(ck | c<k)

Note what this is not: it is not token-level autoregression, and it is not a joint denoiser over the whole sequence. Attention is causal across chunk boundaries and bidirectional within a chunk. Chunk order and target modality come from the task, so the same backbone runs under any multimodal context you care to define.

Chunk-causal attention mask
hover or tap a row
target ↓  /  context →
  • clean past chunks — full context
  • current noisy target — bidirectional inside
  • future chunks — masked

The sequence here is text A (c1–c3), an image (c4), text B (c5–c6) and another image (c7) — exactly the interleaved case. Hover a row to see what that target may look at.

Because the hidden states of completed chunks never depend on future chunks, this mask is also what makes KV caching work at inference: keys and values of the clean prefix are computed once and reused across every sampling step of the current chunk.

3 One objective, both modalities

Here is the part that makes the whole thing cohere. A target chunk — text or image, it does not matter — is perturbed along the same linear probability path between noise and data:

zt(k)  =  t·x(k)  +  (1 − t)·ε(k) ε(k) ∼ 𝒩(0, I),   t = 0 is noise,  t = 1 is data

The backbone sees the noisy target and the clean prefix, and predicts the clean endpoint x̂θ, which is converted into a velocity and matched against the true one. One loss, averaged over every valid position of every target chunk in the batch:

v̂θ = (x̂θ − zt) / (1 − t) · v = x − ε · ℒflow = 𝔼 ‖ v̂θ − v ‖2

Drag the slider below. The left panel is a text chunk, the right panel is a visual chunk, and they are on the same path at the same t. Nothing about the objective knows which is which.

The same path, both modalities
0.00
noise  t = 0t = 1  data
Text chunk cℓ — decoded by the frozen text decoder

8 token positions, resolving out of Gaussian noise in T5 embedding space.

Visual chunk cv — decoded by the frozen image decoder
Example generated image: an astronaut.

256 positions on a 16 × 16 grid, on the same probability path.

Illustrative: the real trajectory lives in embedding space and is rendered only at the end by the frozen decoders. The point is structural — one path, one objective, one t.

Two practical details worth naming. Timesteps are sampled independently per chunk from a shifted logit-normal distribution (α = 8 for image targets, α = 6 for text), so a single sequence trains many noise levels at once. And the velocity conversion clamps the denominator at max(1 − t, 0.05), which is what keeps the clean endpoint numerically well behaved.

A pleasant consequence of sharing the formulation: classifier-free guidance works in both directions. The same guidance equation that sharpens text-to-image generation also sharpens image-conditioned text generation — you simply drop the conditioning chunks and extrapolate. We know of no equivalent knob in a discrete language decoder.

4 Share interaction, not computation

A shared objective does not mean a uniformly shared network. Text embeddings and visual embeddings have genuinely different statistics, and squashing them through identical machinery costs you something. The design rule we settled on is narrow and, we think, transferable: share the parts that mix modalities, specialize the parts that process them.

Multimodal Flow architecture: frozen encoders produce hyperchunks, a shared chunk-causal backbone with joint attention and modality-specific FFNs generates them, and frozen decoders map them back to text or images.
The full dataflow. Frozen encoders map text blocks and images into hyperchunks. A shared chunk-causal backbone applies joint attention across all chunks and routes each modality through its own feed-forward network. Frozen decoders map the generated states back to text or pixels. Positions are encoded with MRoPE, which carries both chunk order and within-chunk structure.

Attention is where cross-modal information actually moves, so it is joint and its projections are shared. The feed-forward networks are where per-modality representation statistics get absorbed, so they are separate. Four parameter-matched configurations, 50B pretraining tokens each:

Attn. projectionFFN Text PPL ↓GenEval ↑DPG ↑CIDEr ↑CLIPScore ↑
SpecificSpecific28.090.25571.0148.140.820
SpecificShared29.670.23370.8247.310.820
SharedSpecific27.430.22671.0447.910.820
SharedShared28.350.23769.7042.020.788

Sharing the attention projections is roughly free — the top and bottom halves trade wins. Sharing the FFNs is not: the last row loses 6 CIDEr and 0.03 CLIPScore, and the damage lands squarely on image-conditioned text generation. Highlighted row is the configuration MF-1 uses.

5 Tasks are just chunk orders

This is the payoff of the previous four sections, and the part we would most like a reader to take away. Once a task is nothing but a sequence of chunks with some of them marked as targets, the distinction between pretraining tasks and downstream tasks mostly evaporates. Text-only language modeling, image generation, captioning, VQA — same backbone, same objective, same sampler. Only the strip changes.

One formulation, many tasks
pick a task
Mixed pretraining
Downstream finetuning
clean context flow target text vision

Mixed pretraining is then not a special recipe, it is just sampling a task τ from a mixture and letting the same flow loss apply to whatever chunk sequence it produces. The model picks up unimodal distributions and bidirectional cross-modal conditionals from the start, rather than acquiring vision as a later bolt-on.

Finetuning changes the data and the sequence shape. It does not change the interface, the factorization, or the loss. Which raises the question §7 takes up: if a task is only a chunk sequence, what else could a chunk be?

6 Parallel training, sequential generation

A causal factorization normally forces you to choose between training efficiency and a meaningful generation order. The chunk-causal mask lets us keep both. During training, every chunk in a sequence can be a target simultaneously: each gets its own noisy view and its own sampled timestep, and a perturbed target attends to the clean prefix but never to its own clean counterpart. All of it resolves in one forward pass.

Training vs. inference

Parallel training with a chunk-causal mask, mixed multimodal pretraining, and sequential chunk-by-chunk inference.
(a) Parallel training. Clean chunks form the context, noisy copies form the targets, and the chunk-causal mask keeps each target honest. Every target contributes a Flow Matching loss in the same pass. (b) Sequential inference. Chunks are generated one at a time; a completed chunk is appended to the cached context before the next one begins.

For throughput, independent examples are packed into physical sequences of up to 32,768 positions, with sequence identifiers in the mask preventing any attention across example boundaries. At inference the picture inverts: chunks are produced one at a time, each by integrating the vector field from t = 0 to t = 1, and each completed chunk joins the KV cache before the next begins.

7 What the formulation admits

Notice what the backbone never learns. It does not learn "text" and it does not learn "images". It learns a vector field over chunks — ordered groups of continuous positions, generated jointly on the inside and causally across. Language and vision are the two instances we happened to train. Neither is privileged by the formulation.

Which invites an obvious test: what would it actually cost to add a third?

What you add, per modality
  • Encoder frozen, into any continuous space
  • Chunking rule what counts as one chunk
  • Modality FFN one block per layer
  • Decoder frozen, back out to the raw signal

Local, swappable, small.

What stays byte-for-byte the same
  • Chunk-causal factorization & mask
  • Joint attention across all chunks
  • The Flow Matching objective
  • Velocity conversion & timestep sampling
  • Sequential sampler, CFG, KV cache
  • Sequence packing

Everything carrying generative semantics.

And what you never need: a second objective to balance · a different sampler · a vocabulary to extend · a tokenizer to train · a codebook to size

The one genuine design decision is the chunking rule, and it is a question worth asking of any signal: what is one coherent act of generation here? Language answers "a short run of tokens." Images answer "the whole grid at once." Most other modalities have an answer just as natural:

trained in MF-1  ·  admitted by the formulation, not trained — these are structural arguments, not results.

Compose a sequence

Here is the claim made testable. Build any sequence you like out of the chunk types above and mark which chunks are targets. The attention mask underneath is derived from whatever you build — it is not a set of hand-drawn examples. Every sequence you can construct is a valid training example for the same loss.

The badge keeps us honest, and distinguishes three things that are easy to conflate: an objective MF-1 was actually trained on; an arrangement of familiar modalities that we simply never ran; and a sequence involving a modality the model has never seen. Only the first is evidence.

Sequence composer
Try:

Click a chunk to flip it between context and target; click its × to remove it.

derived attention mask
clean context noisy target masked

Where this stops being speculation

We want to be precise about the status of this section, because it is the easiest place on this page to overclaim. Everything above is an argument about structure: the formulation does not forbid these sequences, and adding a modality does not perturb the generative core. That is not the same as evidence that they work. Four places where we expect the argument to meet friction:

  • The codec ceiling applies to every new modality. A chunk type is only as good as the frozen encoder–decoder pair behind it. Modalities without a strong pretrained continuous codec inherit that weakness directly, and the flow cannot repair it.
  • Long sequences are untested. Joint attention is quadratic in total positions, and a single image already costs 256. Minutes of video, or a long interleaved document, is a different regime from anything we trained — likely needing sparsity or hierarchy that we have not studied.
  • Low-dimensional targets may need a different noise schedule. The shifted logit-normal timestep distribution was tuned for 512- and 1152-dimensional embeddings. A short control horizon is a far smaller, differently shaped target; we would expect α to need retuning, and possibly more than that.
  • Nothing explicitly ties neighbouring chunks together. Coherence across chunks comes only through causal context. For text that has been enough; for temporally continuous signals like video or audio, it is an open question whether it is.

None of this is settled. But the reason we find the formulation interesting is that these are questions about data and tuning, not about architecture. Adding a modality here does not require inventing a way for it to coexist with the others — that part is already done.

What we observe

A note on scale. MF-1 is pretrained from scratch on 150B tokens. That is small — one to two orders of magnitude below the budgets behind the models it gets compared against, and far below any serious LLM or VLM. We are not claiming a leaderboard position, and you should not read these numbers as one. What we are checking is narrower and, for a new paradigm, more useful: does a fully continuous chunk-based flow behave like a sound pretraining objective?

The objective behaves

Under one mixed-pretraining protocol at 0.6B, 1.2B and 1.6B, all four quantities move in the right direction together and keep moving: the flow loss falls, language perplexity falls, and both text-to-image and image-understanding scores rise. Capacity helps, and it helps more as training continues — the curves are not crossing or flattening into each other.

Four training curves at 0.6B, 1.2B and 1.6B: flow loss, GPT-2-large perplexity, GenEval, and CLIPScore, all improving with steps and with scale.
Mixed multimodal pretraining across capacities. Flow objective, language perplexity, text-to-image GenEval and captioning CLIPScore, tracked to 180K steps. A single continuous objective develops language modeling, image-conditioned text generation and text-to-image generation at the same time.

The comparison we actually trust

Cross-paper leaderboards confound architecture with data, schedule and initialization. So the measurement we weight most is the controlled one: three paradigms, matched data, matched optimization schedule, matched trainable-parameter budget. The Transfusion-style hybrid is the sharpest reference point here, since it shares MF-1's visual pathway and backbone design and differs essentially in using autoregressive text modeling.

ArchitectureTypeGenEval ↑GQA ↑VQAv2 ↑MMBench ↑SEEDB ↑
Multimodal FlowFully continuous0.71355.6069.0346.7451.85
Transfusion-styleDiscrete–continuous hybrid0.66952.8168.4933.6831.31
Chameleon-styleFully discrete0.37445.8356.3738.4042.40

Matched data, optimization schedule and trainable parameters. The gap on the understanding benchmarks is wider than on generation, which is the opposite of what we expected going in.

Pretraining transfers

The other question a new paradigm has to answer is whether its pretraining actually buys anything downstream, or whether the finetuning data is doing all the work. Holding the 1.6B architecture fixed and matching the baseline's downstream tokens against the pretrained model's combined pretraining-plus-finetuning budget, mixed pretraining is worth a great deal — SEED-Bench moves from 31.6 → 62.4 and MMBench from 36.0 → 67.2 against a randomly initialized flow backbone. Same tokens, very different model.

Two design findings

Comparison of visual representation spaces: DINOv2, SigLIP2, FLUX.2 VAE, SD-VAE and raw pixels.
Semantic embeddings beat reconstruction latents. DINOv2 and SigLIP2 representations outperform FLUX.2 and SD VAE latents — and raw pixels — on both generation and captioning. DINOv2 leads on generation; SigLIP2 gives the better balance, which is why MF-1 uses it.
Performance as a function of classifier-free guidance scale, for image captioning and text-to-image generation.
Guidance works in both directions. Sweeping a single guidance scale on a fixed checkpoint: captioning peaks around 3, text-to-image around 5. One mechanism, two output modalities, different optimal strengths.
Headline benchmark numbers, for completeness

The 1.6B MF-1 checkpoint scores 0.821 on GenEval, 83.44 on DPG-Bench, and averages 75.3 across VQAv2, MMBench and POPE. Among unified models it is competitive on generation and holds up on understanding despite training from scratch on 150B tokens, consistently beating comparable from-scratch unified models such as Muddit and D-DiT. We list these because reviewers and readers reasonably want them — not because they are the argument. The argument is the formulation above.

Limitations & outlook

  • Training scale. 150B tokens is a proof of concept, not a converged model. Every trend on this page is measured inside that budget, and we cannot tell you where the curves go at 10×.
  • The frozen codecs are a ceiling. The flow can only be as faithful as the image decoder can reconstruct and the text decoder can detokenize. Anything SigLIP2 or T5-small does not encode is simply not reachable, no matter how good the backbone gets.
  • Fixed visual geometry. One resolution, 224 × 224, one chunk of 256 positions per image. Variable resolution, multi-crop and higher-detail regimes are unexplored.
  • Short interleaving. We trained short text/image sequences — not long interleaved documents, and not video, audio, or control. §7 argues that the formulation admits them and spells out where we expect that argument to run into trouble; none of it is demonstrated here.

The claim we want to defend is modest and specific: continuous, chunk-based embedding flow is a workable third option for unified multimodal modeling — one that keeps continuous visual states without giving up a shared generative process. Whether it is the right option at scale is an open question, and an expensive one. We have released the code and the model weights so that it does not have to stay open.

BibTeX

@article{tao2026multimodalflow,
  title   = {Multimodal Flow: Unified Flow Modeling of Language and
             Vision in Embedding Spaces},
  author  = {Tao, Hongyuan and Wang, Xinggang and Zhu, Lianghui and
             Li, Yongkang and Wei, Yunchao and Feng, Bin and
             Chen, Shaoyu and Zhang, Qian and Huang, Chang and Yu, Kai},
  journal = {arXiv preprint arXiv:2609.40362},
  year    = {2026}
}

Acknowledgments. We thank Lunbin Zeng and Shuai Zhang for valuable discussions and insightful feedback that contributed to this work.