Skip to content

Accuracy — image k-NN on Imagenette (PyTorch vs jepa.cpp, per dtype)

Raw measurement report — the curated view is Benchmarks → Accuracy.

docs/parity.md shows that the jepa.cpp graph reproduces the PyTorch encoders to 6 decimals on eight stored reference tensors. This page answers the question that follows from it: does that hold up on a real dataset, from the real input path, and does quantisation cost accuracy? The task is Imagenette (10 ImageNet classes) with frozen features and a k-NN vote — no training of any kind, the encoders are used exactly as shipped and the "classifier" is a similarity vote over gallery features.

Box: AMD Ryzen Threadripper PRO 7995WX (96 cores / 192 threads, AVX-512), gcc 13.3.0, ggml @ 36da5713, -O3 -march=native, GGML_LLAMAFILE=ON; torch 2.13.0+cpu / transformers 5.16.1, float32, no autocast. 32 threads everywhere, on an otherwise idle box — the PyTorch and jepa.cpp passes were alternated inside one sweep, and every pass records how much CPU time on the machine was not its own (see "Throughput"). Total sweep: 87 min wall, 86.7 of them inside the encoders — the splits, the graph-timing stage and the scoring together account for 25 s of it.

Machine-readable copy of every number below: [tests/results/accuracy-image.json(https://github.com/aselimc/jepa.cpp/blob/main/tests/results/accuracy-image.json); the split definition: [tests/results/accuracy-image-splits.json(https://github.com/aselimc/jepa.cpp/blob/main/tests/results/accuracy-image-splits.json).

Protocol

Implemented once in [scripts/knn_eval.py(https://github.com/aselimc/jepa.cpp/blob/main/scripts/knn_eval.py) and applied unchanged to every backend, so a number can only move because the features moved.

  1. L2-normalise every feature vector; similarity = cosine (a dot product after the normalisation).
  2. kNN top-1: k = 20 nearest gallery vectors, DINO-style weighted vote — each neighbour adds exp(sim / 0.07) to its class, arg-max wins.
  3. centroid top-1: nearest class centroid (mean of the L2-normalised gallery vectors of a class, re-normalised). Parameter-free, so it cannot be tuned into looking good.
  4. agreement: fraction of query items whose kNN prediction equals PyTorch's. This is not "both right" — a shared mistake counts as agreement.
  5. feature cosine: per-item cosine between the jepa.cpp and the PyTorch feature of the same image; the table reports the mean and the single worst item of the query set.
  6. NN margin (used in "Where the predictions differ"): for a query item, the cosine to its best gallery neighbour minus the cosine to the best neighbour of any other class, measured on the PyTorch features. Near zero = the top two classes are tied and any perturbation flips the vote.

Everything is deterministic: fixed seeds, no random tie-breaks, no sampling inside the vote.

Data

data/imagenette/imagenette2-160 — 10 classes, 9 469 train + 3 925 val images.

split used as size
train2000 gallery, 200 images/class, numpy.default_rng(0).choice(...) per class 2 000
train_full gallery (LeJEPA only — it is cheap enough to encode the whole train split) 9 469
val query (I-JEPA, LeJEPA) 3 925
val1000 query (LeWM sanity row), 100/class, default_rng(1) 1 000

The split file stores the per-class indices into the sorted *.JPEG list plus the rng recipe and a sha256 of the resolved file names, so it reproduces exactly on any checkout without committing 6 000 paths.

Feature per model

Each model's own convention, matching tests/fixtures/ref/<model>/manifest.json:

model feature PyTorch jepa.cpp
I-JEPA ViT-H/14 mean of the 256 patch tokens (no CLS) IJepaModel(...).last_hidden_state.mean(1) jepa-embed --pool mean
LeJEPA ViT-S/16 cls = CLS after the final LN, and mean = mean of the 196 patch tokens ViTv2 latent / patch_latent.mean(1) jepa-embed --pool none, then row 0 / mean of rows 1…196
LeWM ViT-Ti/14 emb = projector(CLS), the world-model state re-implementation of dump_reference.build_lewm jepa-embed --pool lewm

PyTorch uses each model's exact reference preprocessing (I-JEPA: squash 224 bilinear, mean = std = 0.5; LeJEPA: short side 256 bicubic + centre crop 224, ImageNet mean/std; LeWM: squash 224 bilinear, ImageNet mean/std) on PIL-decoded pixels. jepa.cpp is given the JPEG file and uses its own stb_image decoder and its own preprocessor — the real-world path, not a fixture replay.

Results

I-JEPA ViT-H/14 (ijepa_vith14_1k) — feature = mean of the 256 patch tokens

backend dtype feature gallery kNN top-1 % centroid top-1 % agreement % mean feat cos worst feat cos img/s
pytorch f32 mean 2 000 (200/class) 95.36 92.25 1 1 5.5
jepa.cpp f16 mean 2 000 (200/class) 95.31 92.25 99.82 0.999944 0.998962 6.2
jepa.cpp q8_0 mean 2 000 (200/class) 95.31 92.31 99.75 0.999774 0.997379 7.0
jepa.cpp q4_k mean 2 000 (200/class) 95.08 92.25 99.11 0.995049 0.979088 4.7

LeJEPA ViT-S/16 (lejepa-vits16-pretrain-in1k)

backend dtype feature gallery kNN top-1 % centroid top-1 % agreement % mean feat cos worst feat cos img/s
pytorch f32 cls 2 000 (200/class) 92.84 87.36 1 1 89.1
jepa.cpp f32 cls 2 000 (200/class) 92.76 87.34 99.85 0.999983 0.999702 66.9
jepa.cpp f16 cls 2 000 (200/class) 92.76 87.36 99.85 0.999983 0.999700 67.3
jepa.cpp q8_0 cls 2 000 (200/class) 92.71 87.41 99.77 0.999843 0.991077 75.0
jepa.cpp q4_k cls 2 000 (200/class) 92.74 87.11 99.06 0.988725 0.885579 67.2
pytorch f32 cls 9 469 (full train) 94.45 87.59 1 1 89.1
jepa.cpp f32 cls 9 469 (full train) 94.52 87.64 99.85 0.999983 0.999702 66.9
jepa.cpp f16 cls 9 469 (full train) 94.55 87.64 99.82 0.999983 0.999700 67.3
jepa.cpp q8_0 cls 9 469 (full train) 94.50 87.64 99.87 0.999843 0.991077 75.0
jepa.cpp q4_k cls 9 469 (full train) 94.22 87.67 99.08 0.988725 0.885579 67.2
pytorch f32 mean 2 000 (200/class) 90.50 86.60 1 1 89.1
jepa.cpp f32 mean 2 000 (200/class) 90.55 86.57 99.77 0.999987 0.999784 66.9
jepa.cpp f16 mean 2 000 (200/class) 90.55 86.57 99.77 0.999987 0.999781 67.3
jepa.cpp q8_0 mean 2 000 (200/class) 90.52 86.75 99.44 0.999929 0.988898 75.0
jepa.cpp q4_k mean 2 000 (200/class) 90.27 86.62 97.38 0.993192 0.968564 67.2
pytorch f32 mean 9 469 (full train) 92.43 87.34 1 1 89.1
jepa.cpp f32 mean 9 469 (full train) 92.33 87.41 99.82 0.999987 0.999784 66.9
jepa.cpp f16 mean 9 469 (full train) 92.33 87.41 99.82 0.999987 0.999781 67.3
jepa.cpp q8_0 mean 9 469 (full train) 92.36 87.44 99.69 0.999929 0.988898 75.0
jepa.cpp q4_k mean 9 469 (full train) 92.41 87.18 97.89 0.993192 0.968564 67.2

LeWorldModel encoder (lewm-pusht) — feature = emb = enc.proj(CLS), sanity row

backend dtype feature gallery kNN top-1 % centroid top-1 % agreement % mean feat cos worst feat cos img/s
pytorch f32 emb 2 000 (200/class) 27.00 24.10 1 1 189.7
jepa.cpp f32 emb 2 000 (200/class) 26.70 24.00 97.40 0.999994 0.999916 95.3
jepa.cpp f16 emb 2 000 (200/class) 26.60 24.00 97.40 0.999994 0.999916 95.2
jepa.cpp q8_0 emb 2 000 (200/class) 26.70 24.00 94.40 0.999929 0.999692 101.0
jepa.cpp q4_k emb 2 000 (200/class) 28.20 24.60 83.40 0.994496 0.980771 85.2
model feature dtype items flipped vs PyTorch PyTorch right jepa.cpp right both wrong median NN margin of the flipped items median NN margin, all items
ijepa-vith14-1k mean f16 7 / 3925 2 0 5 +0.0318 +0.2901
ijepa-vith14-1k mean q8_0 10 / 3925 4 2 4 +0.0247 +0.2901
ijepa-vith14-1k mean q4_k 35 / 3925 17 6 12 +0.0236 +0.2901
lejepa-vits16 cls f32 6 / 3925 1 4 1 +0.0242 +0.1357
lejepa-vits16 cls f16 7 / 3925 1 5 1 +0.0209 +0.1357
lejepa-vits16 cls q8_0 5 / 3925 1 3 1 +0.0134 +0.1357
lejepa-vits16 cls q4_k 36 / 3925 20 11 5 +0.0271 +0.1357
lejepa-vits16 mean f32 7 / 3925 5 1 1 +0.0062 +0.0475
lejepa-vits16 mean f16 7 / 3925 5 1 1 +0.0062 +0.0475
lejepa-vits16 mean q8_0 12 / 3925 6 3 3 +0.0142 +0.0475
lejepa-vits16 mean q4_k 83 / 3925 34 33 16 +0.0090 +0.0475
lewm-pusht emb f32 26 / 1000 9 6 11 +0.0079 +0.0122
lewm-pusht emb f16 26 / 1000 9 5 12 +0.0080 +0.0122
lewm-pusht emb q8_0 56 / 1000 13 10 33 +0.0093 +0.0122
lewm-pusht emb q4_k 166 / 1000 25 37 104 +0.0093 +0.0122

Throughput (32 threads, end-to-end: JPEG decode + preprocess + encode)

backend model dtype split images wall s img/s
PyTorch, batch 32 ijepa-vith14-1k f32 train2000 2000 363.1 5.51
PyTorch, batch 32 ijepa-vith14-1k f32 val 3925 719.5 5.45
jepa-embed, batch 1 ijepa-vith14-1k f16 train2000 2000 324.9 6.16
jepa-embed, batch 1 ijepa-vith14-1k f16 val 3925 637.2 6.16
jepa-embed, batch 1 ijepa-vith14-1k q8_0 train2000 2000 282.4 7.08
jepa-embed, batch 1 ijepa-vith14-1k q8_0 val 3925 557.4 7.04
jepa-embed, batch 1 ijepa-vith14-1k q4_k train2000 2000 419.5 4.77
jepa-embed, batch 1 ijepa-vith14-1k q4_k val 3925 826.8 4.75
PyTorch, batch 32 lejepa-vits16 f32 train_full 9469 109.6 86.40
PyTorch, batch 32 lejepa-vits16 f32 val 3925 44.0 89.15
jepa-embed, batch 1 lejepa-vits16 f32 train_full 9469 141.3 66.99
jepa-embed, batch 1 lejepa-vits16 f32 val 3925 58.7 66.90
jepa-embed, batch 1 lejepa-vits16 f16 train_full 9469 140.6 67.34
jepa-embed, batch 1 lejepa-vits16 f16 val 3925 58.3 67.32
jepa-embed, batch 1 lejepa-vits16 q8_0 train_full 9469 126.5 74.83
jepa-embed, batch 1 lejepa-vits16 q8_0 val 3925 52.3 75.01
jepa-embed, batch 1 lejepa-vits16 q4_k train_full 9469 141.0 67.15
jepa-embed, batch 1 lejepa-vits16 q4_k val 3925 58.4 67.16
PyTorch, batch 32 lewm-pusht f32 train2000 2000 9.7 206.16
PyTorch, batch 32 lewm-pusht f32 val1000 1000 5.3 189.66
jepa-embed, batch 1 lewm-pusht f32 train2000 2000 21.0 95.33
jepa-embed, batch 1 lewm-pusht f32 val1000 1000 10.5 95.26
jepa-embed, batch 1 lewm-pusht f16 train2000 2000 21.0 95.36
jepa-embed, batch 1 lewm-pusht f16 val1000 1000 10.5 95.21
jepa-embed, batch 1 lewm-pusht q8_0 train2000 2000 20.1 99.67
jepa-embed, batch 1 lewm-pusht q8_0 val1000 1000 9.9 101.05
jepa-embed, batch 1 lewm-pusht q4_k train2000 2000 23.1 86.43
jepa-embed, batch 1 lewm-pusht q4_k val1000 1000 11.7 85.16

Measured back-to-back in one sweep on an otherwise idle box: over the 28 timed passes above the machine spent 2673 CPU-minutes out of idle, 2651 of them this benchmark's own process trees, leaving 22.9 CPU-minutes — an average of 0.26 of one core out of 96 — for everything else on the machine (occupancy per row in the JSON: /proc/stat non-idle minus os.times() self+children).

Graph compute alone (jepa-embed --time --repeat 2, 8 query images, 32 threads)

model dtype GGUF ms / image (min–max) median
ijepa-vith14-1k f16 ijepa_vith14_1k-f16.gguf 152.2–160.9 156.8
ijepa-vith14-1k q8_0 ijepa_vith14_1k-q8_0.gguf 130.9–139.2 137.2
ijepa-vith14-1k q4_k ijepa_vith14_1k-q4_k.gguf 202.9–208.4 205.9
lejepa-vits16 f32 lejepa-vits16-pretrain-in1k-f32.gguf 13.0–16.0 13.1
lejepa-vits16 f16 lejepa-vits16-pretrain-in1k-f16.gguf 12.6–15.7 12.8
lejepa-vits16 q8_0 lejepa-vits16-pretrain-in1k-q8_0.gguf 11.2–13.6 11.3
lejepa-vits16 q4_k lejepa-vits16-pretrain-in1k-q4_k.gguf 12.7–16.4 12.7
lewm-pusht f32 lewm-pusht-f32.gguf 9.1–11.1 9.2
lewm-pusht f16 lewm-pusht-f16.gguf 9.3–11.3 9.5
lewm-pusht q8_0 lewm-pusht-q8_0.gguf 8.7–10.2 8.8
lewm-pusht q4_k lewm-pusht-q4_k.gguf 10.2–11.7 10.3

The same already-preprocessed image encoded twice per row, with the JPEG decode, the preprocessing and the model load all outside the number — the throughput table above contains all three, this table contains none of them.

What the numbers say

The headline. PyTorch I-JEPA ViT-H/14 reaches 95.36 % kNN top-1 (92.25 % nearest-centroid) on Imagenette with a 2 000-image gallery — comfortably the sanity bar for this dataset. Every jepa.cpp dtype of that model lands within 0.28 pp of it, and f16 and q8_0 within 0.05 pp. The same holds for LeJEPA ViT-S/16: 94.45 % PyTorch vs 94.22–94.55 % for jepa.cpp f32/f16/q8_0/q4_k on the full-train gallery. Over those two models and both LeJEPA features — ten PyTorch/jepa.cpp comparisons in all — no jepa.cpp row is further than 0.28 pp from its PyTorch row, and none further than 0.13 pp at f32, f16 or q8_0. For I-JEPA and LeJEPA the engine does not cost accuracy at any dtype we ship for images.

That claim is deliberately scoped to those two. The third model, LeWM, is a sanity row on a 1 000-image query set where the encoder sits just above chance, and its top-1 wanders −0.4 / +1.2 pp across dtypes — the q4_k row is 1.2 pp better than PyTorch, which is noise, not a result (see "LeWM is the sanity row" below).

Quantisation costs feature cosine, not decisions. The pooled-feature cosines measured here on 3 925 real images reproduce the numbers docs/quantization.md obtained on eight COCO images by dequantising the GGUF in Python and running the numpy graph — that is, the weight error alone, with no C++ kernel involved:

model / feature q8_0 pooled cos (docs/quantization.md) q8_0 measured here q4_k pooled cos (docs/quantization.md) q4_k measured here
I-JEPA pooled_mean 0.999974 0.999774 0.995850 0.995049
LeJEPA cls 0.999960 0.999843 0.988700 0.988725
LeJEPA pooled_mean 0.999973 0.999929 0.993358 0.993192
LeWM emb 0.999961 0.999929 0.993206 0.994496

Agreement to three or four decimals on every row: the eight-image quantisation study predicts the whole-dataset behaviour. It does not predict it exactly, and the residual is worth reading, because it is where the input path shows up. Measured as distance from 1, the q8_0 rows sit 1.8× to 8.7× further from 1 than the weight-only prediction — I-JEPA 2.26e-4 measured against 2.60e-5 predicted (8.7×), LeJEPA cls 1.57e-4 against 4.00e-5 (3.9×), LeJEPA pooled_mean 7.14e-5 against 2.70e-5 (2.6×), LeWM emb 7.07e-5 against 3.90e-5 (1.8×) — while at q4_k, where the weight error is two orders of magnitude larger, measured and predicted agree to within 0.8–1.2×. That is the signature of a fixed floor that only becomes visible once the weight error is small enough: the JPEG decoder, which the f32 rows measure on its own at 1−cos = 6.3e-6 (LeWM) to 1.7e-5 (LeJEPA cls) with no quantisation involved at all (next section). So the honest statement is not "the C++ kernels add nothing measurable" — at q8_0 something does add roughly as much again as the weights do — but that what it adds is the decoder, is bounded by the f32 rows, and is two orders of magnitude below the q4_k weight error.

What the weight error buys in accuracy terms is essentially nothing: dropping LeJEPA's CLS cosine from 0.99984 (q8_0) to 0.98873 (q4_k) costs 0.28 pp of kNN top-1 on the full gallery (94.50 → 94.22) and gains 0.03 pp on the 2 000-image one (92.71 → 92.74). The reason is in the flips table: the 36 items that q4_k decides differently from PyTorch have a median NN margin of +0.027 against +0.136 for the query set as a whole — they are the items where the two top classes were already tied, and the flips split roughly evenly between "PyTorch was right" and "jepa.cpp was right" (20 vs 11 there, 34 vs 33 for LeJEPA mean q4_k, 17 vs 6 for I-JEPA q4_k). A k-NN decision is a rank comparison; a 1 % cosine perturbation only changes it where the ranking was a coin flip anyway.

So for retrieval / k-NN style use, q4_k is a legitimate configuration for these image encoders — which is not the conclusion the per-token numbers support: docs/quantization.md measures I-JEPA q4_k worst-token cosine 0.526, and test-parity puts every file below 8 bits per weight in its advisory tier for exactly that reason (docs/parity.md, "Thresholds"). Pooling a couple of hundred tokens averages that noise away; anything that consumes tokens individually (the V-JEPA 2 predictor, an attentive pooler) does not get that benefit.

The f32 row is the input-path floor, not a graph error. jepa.cpp f32 does not agree with PyTorch 100 % of the time (LeJEPA: 99.85 %, feature cosine 0.999983) even though docs/parity.md measures cosine 1.000000 for the same file on the stored reference input. The whole difference is the JPEG decoder: docs/parity.md §"Preprocessing" shows resize + crop + normalisation are bit-exact against the HF torchvision backend, while stb_image and PIL/libjpeg disagree by ±1–2 levels on ~2 % of pixels. That is the floor for any comparison that starts from a JPEG, and it is worth 0.05–0.10 pp of top-1 in either direction. Read the f16 and q8_0 rows against it: they are indistinguishable from the f32 row, so at those dtypes the engine's contribution to the difference is smaller than the decoder's. (Feed --rgb-dir-style pixels, as test-parity does, if you need bit-exact agreement.)

Gallery size matters more than dtype. Going from 2 000 to 9 469 gallery images buys LeJEPA +1.6 pp on the CLS feature (92.84 → 94.45 PyTorch) and +1.9 pp on the mean feature — an order of magnitude more than the 0.03–0.23 pp spread across all four dtypes. The nearest-centroid number barely moves (87.36 → 87.59) because a class mean over 200 images is already converged; the k-NN number keeps improving because more neighbours means a better local decision boundary. If accuracy matters, spend the budget on gallery coverage, not on file precision.

CLS beats mean-pooling for LeJEPA by 2.0–2.3 pp on k-NN (94.45 vs 92.43 on the full-train gallery, 92.84 vs 90.50 on the 2 000-image one), which is expected for a model trained with a CLS-level objective — and the two features are affected differently by quantisation (q4_k agreement 99.08 % on CLS, 97.89 % on mean), because the patch-mean features are more tightly clustered and sit closer to the class boundary (median NN margin +0.048 against +0.136 for CLS), so the same perturbation crosses more of them.

LeWM is the sanity row and behaves like one. The Push-T world-model encoder is a ViT-Ti trained on synthetic 224×224 renders of a planar pushing task; on Imagenette it manages 27.0 % kNN top-1 against a 10 % chance floor — above chance (the features are not noise) and far below a general-purpose image encoder, exactly as it should be. Its median NN margin is +0.012, i.e. almost every query item is a near-tie, which is why its agreement column is the most sensitive in the table (97.4 % at f32, 83.4 % at q4_k) while its accuracy wanders by −0.4 / +1.2 pp. Do not read the LeWM agreement number as a jepa.cpp defect: its feature cosine is 0.999994 at f32/f16, the best in the whole table.

Throughput

Two different shapes are compared and the asymmetry is deliberate:

  • PyTorch runs batch 32 with the model loaded once; the reported time excludes the load.
  • jepa-embed ran one image per call when the table above was taken (--batch 1, which was then the only mode), inside a process that is handed a chunk of 512 images and reloads the GGUF once per chunk; that load is in the number.

The table above predates encoder batching. jepa_encode now folds up to 32 image items into one graph (jepa-embed --batch, default 32 in this script) and the features are bit-identical, so every accuracy number on this page is unchanged — verified by re-running the LeJEPA and LeWM f32 passes at --batch 1 and --batch 32 and diffing the JSON (identical payload, tests/results/batching.json). Only the throughput columns move, and they move a lot on the small models: over 561 Imagenette val JPEGs at 32 threads, LeJEPA f16 goes 67.0 → 94.5 img/s and LeWM f16 95.4 → 159.0, while I-JEPA f16 goes 6.16 → 6.92 and q8_0 6.94 → 7.60 (performance.md has the per-batch table). LeJEPA therefore now passes PyTorch's 86–89; LeWM is still behind at 0.77–0.84×. The rows below were not re-measured: re-running the 28-pass sweep would change 28 numbers with no accuracy consequence, and the batch-1 rows remain the honest record of what the tool cost before the change.

On the idle box the picture is much narrower than an earlier, contended version of this page claimed. On the big model jepa.cpp is ahead at q8_0 and f16 and behind at q4_k:

I-JEPA ViT-H/14, end to end PyTorch batch 32 jepa.cpp f16 jepa.cpp q8_0 jepa.cpp q4_k
img/s (train2000 / val) 5.51 / 5.45 6.16 / 6.16 7.08 / 7.04 4.77 / 4.75
vs PyTorch, same split 1.12–1.13× 1.29× 0.87×

At batch 1 PyTorch won the two small models outright: LeJEPA 67–75 img/s against 86–89 (0.75–0.87×) and LeWM 85–101 against 190–206 (0.42–0.53×), where per-image fixed cost and PyTorch's batching dominate. So the range across everything measured in the table is 4.8–7.1 img/s for I-JEPA against PyTorch's 5.5, and the conclusion docs/parity.md draws holds: jepa.cpp is competitive where the matmuls dominate and lost where launch overhead did. That last half is what batching fixed — it is exactly the fixed cost that amortises, which is why it is worth 1.4× on LeJEPA end to end and 1.1× on I-JEPA (see the note above and performance.md).

The load conditions are measured, not asserted. Across the 28 timed passes the machine spent 2 673 CPU-minutes out of idle and this benchmark's own process trees account for 2 651 of them, leaving 22.9 CPU-minutes — 0.26 of one core out of 96 — for everything else. (uptime's load average cannot establish that: a back-to-back sweep of 32-thread passes pins it near 32 however idle the machine is. The per-pass occupancy block in the JSON subtracts os.times() self+children from /proc/stat's non-idle time, which does establish it.) The previous version of this table was taken while a second 32-thread agent shared the box, which cost PyTorch I-JEPA 5.5 → 3.7 img/s and made LeJEPA f16 (41.7 img/s) look slower than the same work at f32 (53.8) — at 67.3 vs 67.0 img/s here, that gap was contention and nothing else.

The dtype comparison needs the graph-compute table, not these numbers. End to end, JPEG decode, preprocessing and one GGUF load per 512-image chunk are all inside the img/s figure. Held out (jepa-embed --time --repeat 2, above):

I-JEPA ViT-H/14, graph compute f16 q8_0 q4_k
ms / image, 32 threads (median) 156.8 137.2 205.9
relative to f16 1.00× 0.87× 1.31×

q4_k is the slowest of the three despite being the smallest file (344 MiB vs 644 q8_0 vs 1206 f16). That is consistent with docs/ggml-notes.md §"llamafile": the accelerated sgemm paths measured there cover F32/F16/Q8_0, and the K-quants fall back to ggml's generic vec-dot. So q4_k on this model is a memory win (3.5× smaller than f16), not a speed win — pick it when the model has to fit, q8_0 when it has to be fast. The penalty is a property of the wide ViT-H matmuls, not of q4_k as such: on LeJEPA ViT-S/16 and LeWM ViT-Ti/14 the same table puts q4_k at 0.99× and 1.08× of f16, i.e. inside the noise, because those graphs are small enough that the vec-dot path is not what the clock is spent on.

Reproduce

Datasets first: scripts/download_datasets.sh fetches Imagenette-160 and the UCF101 subset into data/ (~400 MB, git-ignored).

cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release
cmake --build build -j 32 --target jepa-embed

# everything: splits -> PyTorch features -> jepa.cpp features (all dtypes) -> graph timing -> scoring
.venv/bin/python scripts/bench_accuracy_image.py --stage all --threads 32

# or one stage / one model / one dtype at a time (features are cached under tmp/feat/, so a
# re-run only does the missing work)
.venv/bin/python scripts/bench_accuracy_image.py --stage torch --models ijepa-vith14-1k
.venv/bin/python scripts/bench_accuracy_image.py --stage cpp   --models ijepa-vith14-1k --dtypes q4_k
.venv/bin/python scripts/bench_accuracy_image.py --stage graphtime     # the graph-compute table
.venv/bin/python scripts/bench_accuracy_image.py --stage eval          # re-score cached features

# --batch sets the images per encoder graph in the cpp stage (default 32, PyTorch's batch).  The
# features do not depend on it, so this only moves the throughput columns; --stage graphtime always
# uses --batch 1 because it reports the graph cost of ONE image.
.venv/bin/python scripts/bench_accuracy_image.py --stage cpp --batch 1 --models lejepa-vits16

# 1-minute smoke test.  --limit refuses to write the default --out, and caches its features under
# their own <split>-limit60 names, so it cannot be mistaken for or mixed into a full run.
.venv/bin/python scripts/bench_accuracy_image.py --stage all --limit 60 --dtypes f16 \
                                                 --out tmp/smoke.json

# the tables on this page, straight into it (between the BEGIN/END markers under "Results")
.venv/bin/python scripts/render_accuracy_md.py --write docs/accuracy-image.md

# the protocol on its own, for any two feature files
.venv/bin/python scripts/knn_eval.py --gallery g.npy --gallery-labels gl.json \
                                     --query q.npy --query-labels ql.json --ref-query torch_q.npy

The throughput numbers are only meaningful on an unloaded machine, so every timed pass records what else the box was doing: the 1/5/15-minute load average, and an occupancy block holding the CPU-seconds the whole machine spent out of idle (/proc/stat) against the CPU-seconds this process tree spent (os.times()). Load average alone cannot tell "my own 32 threads" from a second 32-thread job; the leftover foreign CPU time can. Check foreign_cores in tests/results/accuracy-image.json before believing any img/s number, including the ones here.

The features themselves are one jepa-embed call per chunk, e.g.

build/jepa-embed -m models/gguf/ijepa_vith14_1k-q8_0.gguf -i a.JPEG -i b.JPEG ... \
                 --pool mean -o feats.npy -t 32 --print-n 0

jepa-embed emits one pooling mode per call, so the LeJEPA pass that needs both CLS and patch-mean asks for --pool none (all 197 tokens) and reduces them in bench_accuracy_image.py — one encode, two features. A --pool cls,mean (or a multi-output mode) would save the 197×-larger intermediate .npy, which is the only thing missing from the tool for this benchmark.

Wall times on the box above, 32 threads, idle, one pass at a time: PyTorch features 21 min (I-JEPA 18 of them), jepa.cpp features 66 min (I-JEPA 51), the graph-compute stage 11 s, scoring ~1 min — 87 min for the whole sweep.

Not measured

  • I-JEPA f32 in jepa.cpp — skipped for budget. docs/parity.md measures cosine 1.000000 and rel_max 7.9e-5 against PyTorch on the stored input for ijepa_vith14_1k-f32.gguf, and the f16 row here is already within 0.05 pp of PyTorch, so an f32 row would only re-measure the JPEG-decoder floor described above, out of a 2.4 GB file (3.7× q8_0) and at least the 16 min the f16 pass over the same 5 925 images took.
  • q4_0, q5_k, q6_k — the shipped GGUFs exist, but docs/quantization.md already orders them by pooled-feature cosine, and this page shows that ordering does not translate into a k-NN gap worth another 1–2 h of encoding.
  • Video k-NN (UCF-101) — a separate benchmark, not this page.