Blog/world-models/JEPA vs LLMs: what LeCun's world models actually change for engineers
JEPA vs LLMs: what LeCun's world models actually change for engineers
Analysis: JEPA predicts embeddings, not tokens. My text-chunk stand-in: token-space prediction collapsed 99.4% with 10x less data, embedding-space only 45.8%.
What this post covers
How this was tested
Analysis post. Own experiment: Ridge regression and multi-label LogisticRegression (scikit-learn 1.9.1), Python 3.12.14, on 2,041 train / 1,187 test context-to-next-chunk pairs built from the FastAPI English docs corpus (151 files, split by file). This is a small text-only stand-in inspired by the same idea VL-JEPA applies to vision-language -- predict a target's embedding instead of its tokens -- built because no JEPA model was run. Not tested: I-JEPA, V-JEPA, V-JEPA 2, V-JEPA 2.1, VL-JEPA, LeJEPA or LeWorldModel. No GPU and no torch/timm/decord were available in this environment, and V-JEPA 2's checkpoints alone run 300MB-8GB, so no official JEPA code or weights were downloaded or run here.
Does LeCun’s “world model” architecture change anything for someone shipping RAG or agents today? One-line answer: not directly yet — the released JEPA models are vision and video models, not something you drop into a text pipeline. But the core idea behind them — train a model to predict a representation of a target instead of the target’s exact pixels or tokens — already has a text/vision-language version (VL-JEPA), and it’s worth understanding because it’s the same design choice you already make every time you pick an embedding model over a generative one for retrieval. I couldn’t run any official JEPA model on this machine (no GPU, no torch), so I built a small text-only experiment that tests the same shaped claim — predicting embeddings needs less paired data than predicting exact tokens — on our own docs corpus.
What JEPA actually is
JEPA (Joint-Embedding Predictive Architecture) is Yann LeCun’s alternative to two dominant paradigms: autoregressive next-token prediction (LLMs) and generative pixel reconstruction (diffusion/autoencoder image models). The architecture, first shown in I-JEPA, has three pieces (Meta AI, “I-JEPA: The first AI model based on Yann LeCun’s vision for more human-like AI”):
- A context encoder that processes the visible/known part of the input.
- A target encoder — a slowly-updated (exponential moving average) copy of the context encoder — that produces the ground-truth representation of the masked/future part.
- A predictor that takes the context encoder’s output and predicts the target encoder’s representation, not the raw input.
The loss is prediction error in embedding space, not pixel or token reconstruction error. Meta’s own framing of why: generative models try “to fill-in every bit of missing information, even though the world is inherently unpredictable,” which is why they burn capacity on unpredictable detail (Meta’s example: rendering exact finger positions) instead of the semantics that actually matter (same source). V-JEPA 2 applies the same idea to video: the predictor “takes in a video embedding and additional context about what to predict and outputs predicted embeddings,” not pixels (Meta AI, “Introducing the V-JEPA 2 world model and new benchmarks for physical reasoning”).
What’s actually released, as of 2026-09-28
| Model | What it predicts | Released | Code / weights |
|---|---|---|---|
| I-JEPA (2023) | image patch embeddings | yes | training code + checkpoints, facebookresearch/ijepa |
| V-JEPA (2024) | video clip embeddings | yes | facebookresearch/jepa |
| V-JEPA 2 (Jun 2025) | video embeddings + zero-shot robot planning; trained on 1M+ hours of video, 1M images, 62 hours of robot data | yes | ViT-L/H/g checkpoints (300M-1B params), facebookresearch/vjepa2 |
| V-JEPA 2.1 (Mar 2026) | same, “more temporally consistent dense features” | yes | ViT-B through ViT-G checkpoints (80M-2B params), same repo |
| LeJEPA (Nov 2025) | theory + SIGReg regularizer for any JEPA | yes | ~50-line regularizer, rbalestr-lab/lejepa, paper arXiv:2511.08544 |
| LeWorldModel (Mar 2026) | pixel-to-pixel-free, stable end-to-end JEPA from raw pixels, ~15M params | paper only, code not confirmed released as of this writing | arXiv:2603.19312 |
| VL-JEPA (ICLR 2026) | text embeddings from vision-language input, decoded to tokens only when text output is needed | paper; code/weights not stated in the abstract | arXiv:2512.10942 |
VL-JEPA is the closest thing to “JEPA for the RAG/agents audience”: instead of generating an answer token-by-token, it predicts the target text’s embedding and only invokes “a lightweight text decoder… when needed to translate VL-JEPA predicted embeddings into text,” which the paper reports cuts decoding operations 2.85x versus uniform decoding while using half the trainable parameters of comparable vision-language models (arXiv:2512.10942). None of that is something you can pip install today — the abstract doesn’t state a code or weight release, and I didn’t find one.
Worth knowing for context: Yann LeCun left Meta in November 2025 after 12 years as Chief AI Scientist specifically to build world models outside a company focused on LLMs, co-founding AMI (Advanced Machine Intelligence) Labs, which raised a $1.03B seed round announced March 10, 2026 (CNBC) (MIT Technology Review, “Yann LeCun’s new venture is a contrarian bet against large language models”). LeWorldModel and V-JEPA 2.1 both list LeCun as a co-author and postdate that departure, so this research direction is active, not abandoned.
Why this matters now, and why it doesn’t (yet)
None of the released JEPA models take text-only input the way an embedding model in a RAG pipeline does — they’re vision/video encoders. So there’s nothing here to swap in for text-embedding-3-small or your current retriever today. What is directly relevant is the design principle: predicting a representation of the target, not the target’s exact surface form, is a real, separate choice from next-token prediction — and RAG builders already made a version of this choice by using embedding models (representation prediction, trained with contrastive loss) instead of asking an LLM to regenerate a document to compare against a query (surface-form prediction). VL-JEPA is evidence that applying JEPA’s specific training recipe to that same idea, for vision-language, beats CLIP-style contrastive embeddings and matches instruction-tuned VLMs on some benchmarks (arXiv:2512.10942) — that’s a vendor-adjacent, third-party-unverified paper result, not something I measured.
Since I can’t run any of these models, I tested the underlying claim that’s most exportable to a text pipeline: does predicting a target’s embedding need less paired training data than predicting the target’s exact discrete content? That’s the sample-efficiency argument JEPA is built around, tested here at toy scale on our own text corpus.
My experiment: predicting embeddings vs. predicting tokens, on text
Task: given one chunk of FastAPI documentation (the “context”), predict the chunk immediately following it in the same file (the “target”). Both chunks are represented with a shared TF-IDF vectorizer (max_features=1500, stop_words="english", fit only on training text).
- Task A, embedding-space (JEPA-style): a
Ridgeregression maps the context’s TF-IDF vector to the target’s TF-IDF vector. Scored by nearest-neighbor retrieval: does the predicted embedding’s closest match (by cosine similarity), among all test-set target embeddings, equal the true next chunk? (hit@1, hit@5, MRR.) The prediction only has to land close to the right representation, not reproduce it exactly. - Task B, token-space (next-token-style): multi-label
LogisticRegression(150 classifiers, one per common vocabulary word) predicts which of the 150 most frequent target-vocabulary words will appear in the next chunk, from the same context TF-IDF input. Scored by micro-F1 against the true word set. The prediction has to get exact discrete symbols right — the same shape as next-token prediction, at word-presence granularity.
Both tasks share the same input features and are retrained at four training-set sizes (10%, 25%, 50%, 100% of 2,041 train pairs) to compare how each degrades as paired data shrinks. Files are split 120 train / 31 test, so no chunk from a test file appears in training. This is not JEPA — no masking, no target encoder, no EMA, no video — it’s the smallest honest test of “embedding-space prediction vs. token-space prediction, same data, same input features” I could run without a GPU.
Results
| training data | task A hit@1 | task A hit@5 | task A MRR | task B micro-F1 |
|---|---|---|---|---|
| 10% (204 pairs) | 0.011 | 0.026 | 0.025 | 0.001 |
| 25% (510 pairs) | 0.012 | 0.051 | 0.039 | 0.105 |
| 50% (1,020 pairs) | 0.018 | 0.061 | 0.049 | 0.120 |
| 100% (2,041 pairs) | 0.020 | 0.083 | 0.062 | 0.230 |
Zero-shot baseline (no training at all — retrieve using the raw context TF-IDF vector directly against the target pool): hit@1=0.005, hit@5=0.127, MRR=0.068. Task B’s majority-class baseline (predict each word present if it appears in over 50% of training targets): micro-F1=0.000. Full output, including the per-run timings, is in results.txt.
What the numbers mean
Token-space prediction collapses far harder than embedding-space prediction when data is scarce. Going from 100% to 10% of training pairs (a 10x reduction), task B’s micro-F1 drops 99.4% relative (0.230 to 0.001, computed in bench_embed_vs_token.py) — it effectively stops working. Task A’s hit@1 drops 45.8% relative (0.020 to 0.011) over the same data cut — worse, but nowhere near collapse. That’s a 2.17x steeper relative drop for the token-space task. This is the direction JEPA’s sample-efficiency argument predicts: forcing a model to get exact discrete symbols right, word by word, needs more paired examples than a model that only has to land near the right region of embedding space.
The surprising result: an untrained baseline beat the trained embedding predictor on two of three retrieval metrics. The zero-shot baseline — using the raw context TF-IDF vector as the query, no Ridge model at all — scored hit@5=0.127 and MRR=0.068, both higher than the fully-trained task A model’s hit@5=0.083 and MRR=0.062 (zero-shot is 1.54x and 10.2% ahead respectively). The trained model only won on hit@1 (0.020 vs 0.005, a 4x edge). The likely reason: adjacent paragraphs in the same documentation file share a lot of vocabulary already (same feature, same code example, same section), so raw lexical overlap is a strong retrieval signal on its own. A single global linear map (Ridge, L2-regularized) trained on ~2,000 examples doesn’t have much room to improve on that shared-vocabulary prior — and its regularization pulls predictions toward the average of all training targets, which can blur out the specific overlap that made the zero-shot version work. This doesn’t contradict the sample-efficiency finding above; it says the architecture of the predictor matters as much as the training objective, and a linear map is a weak predictor to test JEPA’s design against. Real JEPA implementations use a Vision Transformer predictor and a separately-trained (EMA) target encoder, not a shared TF-IDF space with a linear regressor.
Honest caveats
- This is not JEPA. No masking strategy, no EMA target encoder, no vision/video data, no official code. It’s a same-shape stand-in built to test one claim (data efficiency of representation-prediction vs. token-prediction) on text, because no JEPA model could run here.
- Task B is a rough analog of next-token prediction, not a real one. Predicting which of 150 words appear anywhere in the next chunk (multi-label, order-free) is easier in some ways and harder in others than predicting the exact next token in sequence. Don’t read the 0.230 micro-F1 as comparable to any published LLM benchmark.
- Absolute retrieval numbers are low across the board (best hit@1 is 0.020, 2% of test cases) because the candidate pool is the full 1,187-chunk test set and the task — “guess the next paragraph in a technical doc, exactly” — is genuinely hard. The relative comparison between task A and task B, not the absolute numbers, is the finding.
- One corpus, one task, one seed. The FastAPI docs corpus has unusually high adjacent-paragraph lexical overlap because of how its tutorial/advanced pages are structured; a corpus with less local topic continuity might close the zero-shot-vs-trained gap differently.
Ridgeis a linear model, chosen for CPU speed, not because it’s what real JEPA predictors use. A small MLP predictor might change the trained-vs-zero-shot result in task A.
The pattern, copyable
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import Ridge
from sklearn.metrics.pairwise import cosine_similarity
# context_texts[i] pairs with target_texts[i]; candidate_texts is the retrieval pool at inference
vec = TfidfVectorizer(max_features=1500, stop_words="english", min_df=2)
vec.fit(context_texts + target_texts)
X = vec.transform(context_texts).toarray()
Y = vec.transform(target_texts).toarray()
predictor = Ridge(alpha=5.0).fit(X, Y) # predicts a representation, not exact text
def retrieve(query_text, candidate_texts, candidate_embeddings):
pred = predictor.predict(vec.transform([query_text]).toarray())
sims = cosine_similarity(pred, candidate_embeddings)[0]
return candidate_texts[sims.argmax()]
This is the shape of the idea, not a recommendation to use it as-is: check whether a trained predictor beats a zero-shot embedding-similarity baseline on your data before adding the extra model, the same way this experiment found it didn’t here.
Rule of thumb
Don’t wait for a text version of V-JEPA — none exists as a released model today, and the closest thing (VL-JEPA) is vision-language, unreleased, and reported only by its own paper. What you can use now is the design question underneath it: for any pipeline step that compares or ranks content (retrieval, deduplication, routing), ask whether you’re predicting exact content or a representation of it, and pick the objective that matches the amount of paired training data you actually have — representation-prediction degrades more gracefully when that data is thin, as it did here even in a small, non-JEPA test. And before trusting a learned predictor over a naive similarity baseline in a retrieval step, measure the zero-shot baseline first; this experiment’s own trained model lost to it on two of three metrics.
For the retrieval side of this same pipeline, see this site’s hybrid search post for a retriever that’s a better default than either of the toy predictors here, and the chunking benchmark for how chunk boundaries change what “the next chunk” even means. For another case of a specialized architecture pitched as an LLM alternative, see Jev vs LLMs — a decision model, not a world model, but the same “narrower architecture, narrower job, check the vendor’s numbers yourself” pattern applies.
Full experiment script
bench_embed_vs_token.py and results.txt are in experiments/jepa-vs-llm-world-models/ — download them from the Experiment files box below, run against the corpus fetched by experiments/fetch_corpus.sh. It builds the file-split context/target pairs, trains both tasks at four data sizes, and prints the zero-shot baseline, both result tables, and the sample-efficiency ratios shown above.
Experiment files
Want the production version of this?
A production RAG + MCP starter kit for FastAPI: hybrid search, validated citations, evals in CI, Docker, 32 tests. It runs offline in about 60 seconds.
See ShipRAG or get the free RAG checklist first.