Model Lab

Empirical model testing, published openly.

Real measurements from real generation runs: convergence per step, quality per resolution and aspect ratio, cost, reproducibility, failure rate. Published when the method is reproducible by a stranger with the same open-source model and config. Client-derived tests and proprietary pipeline internals stay private, on principle, not as a limitation, see the disclosure note below.

Measured, not decorated

What the model actually did, step by step.

Real per-step data from actual generation runs, rendered, not prompted, not fabricated. One well-instrumented run, shown through six different lenses — this is what pipeline instrumentation looks like, and it's the same per-step visibility every pipeline I ship has to offer.

Real run · moon-01 · AI-generated, GENERAITR Engine · three lenses01–04 · scroll →

The real generation Rose, Violet and Ochre below are derived from, same run, same 9 real steps, three different lenses on one real event.

Source · AI-generated, real per-step capture (moon-01)

Lattice · real step lattice (from moon-01)

PCA scatter · real PCA scatter (from moon-01)

Polar plot · real polar plot (from moon-01)

A different real pipeline, still fully AI-generated

Not a generation step like the runs above: a Gaussian splat reconstruction, a single AI-generated image resolved into an orbitable 3D point cloud. Both the source image and the reconstruction are AI-generated, a real double-AI pipeline, not a photographed object turned into a splat.

AI-generated source image, teddy bear

Input: one AI-generated still (Flux 2 Pro), not a photograph

3D from one single image · Gaussian splat, teddy bear

AI-generated source image, vintage Vespa

Input: one AI-generated still (Flux 2 Pro), not a photograph

3D from one single image · Gaussian splat, vintage Vespa

Loading real frames…
Height-fieldReal height-field, step 1/24
Loading real frames…
Delta fieldReal inter-frame delta, step 1/24
Loading real frames…
ScanlineReal per-cell crossfade, step 1/24

Media wall: every modality, every stage

real tiles: solid border · fake tiles: dashed red + tag
Real denoising frame, step 1

Moon-01 · step 1/24 · real frame

Real denoising frame, step 2

Moon-01 · step 2/24 · real frame

Real denoising frame, step 4

Moon-01 · step 4/24 · real frame

Real denoising frame, step 6

Moon-01 · step 6/24 · real frame

Real denoising frame, step 9

Moon-01 · step 9/24 · real frame

Real denoising frame, step 12

Moon-01 · step 12/24 · real frame

Real denoising frame, step 16

Moon-01 · step 16/24 · real frame

Real denoising frame, step 20

Moon-01 · step 20/24 · real frame

Real denoising frame, step 24

Moon-01 · step 24/24 · real frame

Real denoising frame, step 1Real denoising frame, step 2Real denoising frame, step 4Real denoising frame, step 6Real denoising frame, step 9Real denoising frame, step 12Real denoising frame, step 16Real denoising frame, step 20Real denoising frame, step 24step 1

Moon-01 · full sweep, real capture · synced to shared clock

Loading real splat data…
Real capture, freely orbitableConventional capture (RealityCapture + Postshot)

Flowers · gaussian splat · conventional photogrammetry from MANY photos, the opposite pipeline to the single-image AI splats above

Real product-shot generation frame, step 1

Product shot · step 1/24 · real frame · Qwen-Image-Edit-2511

Real product-shot generation frame, step 2

Product shot · step 2/24 · real frame · Qwen-Image-Edit-2511

Real product-shot generation frame, step 4

Product shot · step 4/24 · real frame · Qwen-Image-Edit-2511

Real product-shot generation frame, step 6

Product shot · step 6/24 · real frame · Qwen-Image-Edit-2511

Real product-shot generation frame, step 9

Product shot · step 9/24 · real frame · Qwen-Image-Edit-2511

Real product-shot generation frame, step 12

Product shot · step 12/24 · real frame · Qwen-Image-Edit-2511

Real product-shot generation frame, step 16

Product shot · step 16/24 · real frame · Qwen-Image-Edit-2511

Real product-shot generation frame, step 20

Product shot · step 20/24 · real frame · Qwen-Image-Edit-2511

Real product-shot generation frame, step 24

Product shot · step 24/24 · real frame · Qwen-Image-Edit-2511

Fake

Splat reprojection · 0° · not run yet

Fake
90°

Splat reprojection · 90° · not run yet

Fake
180°

Splat reprojection · 180° · not run yet

Fake

Audio · per-step sonification · no real capture yet

Fake
10%

Splat growth · 10% densified · optimizer not built yet

Fake
40%

Splat growth · 40% densified · optimizer not built yet

Fake
75%

Splat growth · 75% densified · optimizer not built yet

Fake
100%

Splat growth · 100% densified · optimizer not built yet

Real subjects, compared · every recipe

real final frames, real metadata

Moon-01

ModelGENERAITR Engine

Real text-to-image run, 9 real captured denoising steps, real Laplacian detail-energy measured per step (see the chart above).

Moon-01, real final frame

Product shot

ModelQwen-Image-Edit-2511

Real edit-model run, 9 real captured steps: starts from a real reference image, not noise, detail falls as the edit converges.

Product shot, real final frame

Mouse shot

ModelQwen-Image-2512

Real edit-model run, 9 real captured steps, the first batch to meet reveal-content-capture-spec.md's requirements in full.

Mouse shot, real final frame

Parameter optimization · bending the model

A LoRA is a small, fast add-on trained to reshape a base model's behavior toward a specific brand, subject, or look, without retraining the whole model from scratch.

Want this method run on your own model or brand? Get in touch →

Finding the sweet spot

"Sweet spot" isn't a subjective call here. For any real independent variable (step count, resolution, aspect ratio) a model is swept across, a real quality or convergence signal is measured at each point, and the sweet spot is found geometrically: the point of maximum curvature on the normalized curve (the "kneedle" elbow-detection family). Marked automatically on every chart, not eyeballed.

Sweep axes

Step count: real per-step measured detail/convergence, same technique proven on the homepage's moon-01 pieces.

Resolution / aspect ratio: real per-configuration quality measurement across a model's supported range.

Spatial & temporal consistency: how much a model drifts across a batch (same prompt/style, different seeds) or across frames in a video run, real perceptual-hash distance between real outputs. This answers "will asset #1000 still be on-brand" for a production batch. Not measured yet, planned alongside the resolution sweep.

Step-count convergence

1 real run
Detail energy vs. stepLaplacian magnitude
Denoising stepReal measured detail energy

Real per-step Laplacian-filter magnitude from the moon-01 generation's 9 real captured steps. Convergence within one run's trajectory, a proxy for the sweet-spot question, not yet a true cross-run sweep. Elbow found in today's data: step 12.

Step count, in numbers

Steps captured9samples
Range1–24steps
Sweet spot12steps
Detail gain past sweet spot19%

Resolution & aspect-ratio sweet spot

Preview · fake data
Quality vs. resolutionper model, real sweep pending
Resolution / aspect ratioReal measured quality signal

Invented numbers, layout preview only. No resolution/AR sweep has run yet.

n=15 real runs

Cost & latency: LTX-2.3

Cost per asset0.0174USD
Median latency30.1s
P95 latency39.4s
Failure rate0%
n=13 real runs

Cost & latency: LTX-2.5

Cost per asset0.0180USD
Median latency31.2s
P95 latency34.7s
Failure rate0%

Real Verda serverless runs, RTX PRO 6000 96GB, distilled i2v (image-conditioned video, 768×512-class resolution, 121 frames, 24fps). Cost/asset = median wall-clock time × the RTX PRO 6000 serverless hourly rate ($2.079/hr). Zero failures in either sample after the pipeline stabilized — not a claim about failure rate under load, just what this batch measured.

Disclosure policy

Published if a stranger could reproduce it with the same open-source model and published seed/config. GENERAITR-specific pipeline architecture and any client-derived test stays private, results only where shown at all.

Start with a pipeline audit

Compare models

Preview · fake data
SpecModel A (fake)Model B (fake)Model C (fake)
Steps (sweet spot)202816
Resolution1024x10241024x10241280x1280
Cost / asset$0.012$0.019$0.021
Median latency6.4s9.1s10.8s
VRAM12 GB16 GB24 GB
LicenseApache 2.0OpenRAIL-MApache 2.0

Invented specs, layout preview only. No real head-to-head has run yet.

SageAttention kernel A/B

SpecTrellis (3D, image-to-splat)Wan 2.2 (video, t2v)
Result18.3% SLOWER1.29x FASTER
Baseline (xformers)5233ms median
SageAttention6193ms median55.9s mean
GPUL40SH100
QualityIntact (within run-to-run noise floor)Intact (no artifacting, no drift)
Whyhead_dim=64, ~4200 small calls/gen, CUDA-locked to slower v1 kernelLarger attention chunks, fewer calls per generation, the profile SageAttention is built for
Baseline (bf16 SDPA)72.3s mean

Real A/B benchmark, not a spec sheet: same GPU, same input, same seed, interleaved runs so thermal/scheduler drift can't bias either arm. Quality checked both directions, not just speed. Trellis's loss predicted Wan 2.2's win — same mechanical explanation, opposite outcome, same day.

LTX-2.3 vs LTX-2.5, real generations

SAM3 segmentation control

What didn't work

real attempts, honest results

A section that only ever shows wins is marketing copy with extra steps. These are real attempts with a real result, published with the same rigor as everything above, including the parts that didn't fully land.

TAOS, live · not a simulation

polls the real public stats endpoint

A live particle field driven by TAOS, the agent OS that runs GENERAITR's own operations. Amber nodes are the real agents that did real work in the last 7 days; pulses spawn from the real event-throughput delta between polls, and stop when it is zero. The motion is designed; what drives it is measured.

Live connection to TAOS unavailable
moon-01 1/24