Model Lab
Empirical model testing, published openly.
Real measurements from real generation runs: convergence per step, quality per resolution and aspect ratio, cost, reproducibility, failure rate. Published when the method is reproducible by a stranger with the same open-source model and config. Client-derived tests and proprietary pipeline internals stay private, on principle, not as a limitation, see the disclosure note below.
What the model actually did, step by step.
Real per-step data from actual generation runs, rendered, not prompted, not fabricated. One well-instrumented run, shown through six different lenses — this is what pipeline instrumentation looks like, and it's the same per-step visibility every pipeline I ship has to offer.
Real run · moon-01 · AI-generated, GENERAITR Engine · three lenses01–04 · scroll →
The real generation Rose, Violet and Ochre below are derived from, same run, same 9 real steps, three different lenses on one real event.
Source · AI-generated, real per-step capture (moon-01)
Lattice · real step lattice (from moon-01)
PCA scatter · real PCA scatter (from moon-01)
Polar plot · real polar plot (from moon-01)
A different real pipeline, still fully AI-generated
Not a generation step like the runs above: a Gaussian splat reconstruction, a single AI-generated image resolved into an orbitable 3D point cloud. Both the source image and the reconstruction are AI-generated, a real double-AI pipeline, not a photographed object turned into a splat.
Input: one AI-generated still (Flux 2 Pro), not a photograph
3D from one single image · Gaussian splat, teddy bear
Input: one AI-generated still (Flux 2 Pro), not a photograph
3D from one single image · Gaussian splat, vintage Vespa
Media wall: every modality, every stage
real tiles: solid border · fake tiles: dashed red + tag
Moon-01 · step 1/24 · real frame

Moon-01 · step 2/24 · real frame

Moon-01 · step 4/24 · real frame

Moon-01 · step 6/24 · real frame

Moon-01 · step 9/24 · real frame

Moon-01 · step 12/24 · real frame

Moon-01 · step 16/24 · real frame

Moon-01 · step 20/24 · real frame

Moon-01 · step 24/24 · real frame








step 1Moon-01 · full sweep, real capture · synced to shared clock
Flowers · gaussian splat · conventional photogrammetry from MANY photos, the opposite pipeline to the single-image AI splats above

Product shot · step 1/24 · real frame · Qwen-Image-Edit-2511

Product shot · step 2/24 · real frame · Qwen-Image-Edit-2511

Product shot · step 4/24 · real frame · Qwen-Image-Edit-2511

Product shot · step 6/24 · real frame · Qwen-Image-Edit-2511

Product shot · step 9/24 · real frame · Qwen-Image-Edit-2511

Product shot · step 12/24 · real frame · Qwen-Image-Edit-2511

Product shot · step 16/24 · real frame · Qwen-Image-Edit-2511

Product shot · step 20/24 · real frame · Qwen-Image-Edit-2511

Product shot · step 24/24 · real frame · Qwen-Image-Edit-2511
Splat reprojection · 0° · not run yet
Splat reprojection · 90° · not run yet
Splat reprojection · 180° · not run yet
Audio · per-step sonification · no real capture yet
Splat growth · 10% densified · optimizer not built yet
Splat growth · 40% densified · optimizer not built yet
Splat growth · 75% densified · optimizer not built yet
Splat growth · 100% densified · optimizer not built yet
Real subjects, compared · every recipe
real final frames, real metadataMoon-01
Real text-to-image run, 9 real captured denoising steps, real Laplacian detail-energy measured per step (see the chart above).

Product shot
Real edit-model run, 9 real captured steps: starts from a real reference image, not noise, detail falls as the edit converges.

Mouse shot
Real edit-model run, 9 real captured steps, the first batch to meet reveal-content-capture-spec.md's requirements in full.

Parameter optimization · bending the model
A LoRA is a small, fast add-on trained to reshape a base model's behavior toward a specific brand, subject, or look, without retraining the whole model from scratch.
Want this method run on your own model or brand? Get in touch →
Finding the sweet spot
"Sweet spot" isn't a subjective call here. For any real independent variable (step count, resolution, aspect ratio) a model is swept across, a real quality or convergence signal is measured at each point, and the sweet spot is found geometrically: the point of maximum curvature on the normalized curve (the "kneedle" elbow-detection family). Marked automatically on every chart, not eyeballed.
Sweep axes
Step count: real per-step measured detail/convergence, same technique proven on the homepage's moon-01 pieces.
Resolution / aspect ratio: real per-configuration quality measurement across a model's supported range.
Spatial & temporal consistency: how much a model drifts across a batch (same prompt/style, different seeds) or across frames in a video run, real perceptual-hash distance between real outputs. This answers "will asset #1000 still be on-brand" for a production batch. Not measured yet, planned alongside the resolution sweep.
Step-count convergence
1 real runReal per-step Laplacian-filter magnitude from the moon-01 generation's 9 real captured steps. Convergence within one run's trajectory, a proxy for the sweet-spot question, not yet a true cross-run sweep. Elbow found in today's data: step 12.
Step count, in numbers
Resolution & aspect-ratio sweet spot
Preview · fake dataInvented numbers, layout preview only. No resolution/AR sweep has run yet.
Cost & latency: LTX-2.3
Cost & latency: LTX-2.5
Real Verda serverless runs, RTX PRO 6000 96GB, distilled i2v (image-conditioned video, 768×512-class resolution, 121 frames, 24fps). Cost/asset = median wall-clock time × the RTX PRO 6000 serverless hourly rate ($2.079/hr). Zero failures in either sample after the pipeline stabilized — not a claim about failure rate under load, just what this batch measured.
Disclosure policy
Published if a stranger could reproduce it with the same open-source model and published seed/config. GENERAITR-specific pipeline architecture and any client-derived test stays private, results only where shown at all.
Start with a pipeline auditCompare models
Preview · fake data| Spec | Model A (fake) | Model B (fake) | Model C (fake) |
|---|---|---|---|
| Steps (sweet spot) | 20 | 28 | 16 |
| Resolution | 1024x1024 | 1024x1024 | 1280x1280 |
| Cost / asset | $0.012 | $0.019 | $0.021 |
| Median latency | 6.4s | 9.1s | 10.8s |
| VRAM | 12 GB | 16 GB | 24 GB |
| License | Apache 2.0 | OpenRAIL-M | Apache 2.0 |
Invented specs, layout preview only. No real head-to-head has run yet.
SageAttention kernel A/B
| Spec | Trellis (3D, image-to-splat) | Wan 2.2 (video, t2v) |
|---|---|---|
| Result | 18.3% SLOWER | 1.29x FASTER |
| Baseline (xformers) | 5233ms median | – |
| SageAttention | 6193ms median | 55.9s mean |
| GPU | L40S | H100 |
| Quality | Intact (within run-to-run noise floor) | Intact (no artifacting, no drift) |
| Why | head_dim=64, ~4200 small calls/gen, CUDA-locked to slower v1 kernel | Larger attention chunks, fewer calls per generation, the profile SageAttention is built for |
| Baseline (bf16 SDPA) | – | 72.3s mean |
Real A/B benchmark, not a spec sheet: same GPU, same input, same seed, interleaved runs so thermal/scheduler drift can't bias either arm. Quality checked both directions, not just speed. Trellis's loss predicted Wan 2.2's win — same mechanical explanation, opposite outcome, same day.
LTX-2.3 vs LTX-2.5, real generations
SAM3 segmentation control
What didn't work
real attempts, honest resultsA section that only ever shows wins is marketing copy with extra steps. These are real attempts with a real result, published with the same rigor as everything above, including the parts that didn't fully land.
TAOS, live · not a simulation
polls the real public stats endpointA live particle field driven by TAOS, the agent OS that runs GENERAITR's own operations. Amber nodes are the real agents that did real work in the last 7 days; pulses spawn from the real event-throughput delta between polls, and stop when it is zero. The motion is designed; what drives it is measured.