# Reproduce one Cosmos shot

This small request helper makes four takes of your own keyframe. It uses the full Cosmos image-to-video settings from our studio recipe. It is not the studio engine, critic, take selector or film editor. It does not deploy anything or publish a film.

The helper has passed syntax and offline mocked-client checks. We have not run this public helper on a GPU, and it does not attest which weights a server actually loaded. The archived studio measurements in the article are separate evidence.

## Prepare your server and keyframe

1. Follow NVIDIA's [official Cosmos3 full-model vLLM-Omni serving recipe](https://huggingface.co/nvidia/Cosmos3-Super-Image2Video#vllm-omni). Use the full model, not the 4-step checkpoint. Complete the official model/guardrail setup and keep guardrails enabled. This client sends `guardrails: true`; a server without the required setup may reject the request.
2. For the recorded recipe, the expected model revision is `580f3f28e33ba93c8d464768876da8c322619aad`, and the vLLM-Omni image is `vllm/vllm-omni@sha256:970dee6658ea223f615b2438ce41e47f1d5322225482546e6e6bc5d8134f757c`. Obtain the weights through your own authorized access, following the model's terms. Record the revision and container you actually use; changing them changes the experiment.
3. If you provision the server through Modal, use your own account and resources. Modal's [model-weight guide](https://modal.com/docs/guide/model-weights) explains persistent weights; [lifecycle functions](https://modal.com/docs/guide/lifecycle-functions) explain loading a model once per container. Our full-model worker uses one B200, four CPU cores and 128 GiB of memory per container, with a 90-second idle scale-down window. This helper does not configure those resources. Loading, generation, idle time and storage can incur charges; our grant coverage does not transfer to another account.
4. Make a 720 x 1280 PNG containing every figure and prop needed for the action. The helper checks its PNG header and dimensions but does not decode or resize it. The example prompt assumes one generic paper figure and one blue block already exist in your image. Replace the example to match your actual keyframe.

The default endpoint is the server on the same machine as the client: `http://127.0.0.1:8000`. A server listening only on a Modal container's loopback address is not reachable from your laptop's loopback address. Run this helper where the server is reachable, or use your own correctly configured endpoint. This kit does not expose a public service or manage authentication.

## Make four takes

Save these three files together. Python 3.10 or newer is sufficient; the client has no third-party Python dependencies.

```bash
python3 generate_takes.py \
  --keyframe keyframe.png \
  --prompt prompt.json \
  --out takes-001
```

To use another already running server, add `--endpoint http://your-server:8000`. The endpoint must be a server origin with no embedded credentials, query or path. The helper appends `/v1/videos/sync` and refuses redirects.

The JSON file contains exactly `prompt` and `negative_prompt`. `prompt` can be a string or a temporal-caption object, as in the example. The negative prompt is an editable string. The example is illustrative; no generated output or quality result is claimed for it.

Each request sends a PNG as multipart `input_reference`, requests MP4 bytes and fixes:

| Setting | Value |
| --- | --- |
| Output size | 720 x 1280 |
| Frames / frame rate | 121 / 24 fps (about 5.04 seconds) |
| Inference steps | 50 |
| Guidance scale | 6.0 |
| Flow shift | 5.0 |
| Seeds, sequentially | 7, 21, 42, 99 |
| Resolution/duration templates | Disabled; explicit settings above |
| Guardrails | Enabled |

Four seeds are four generation requests. There are no retries. The first failure stops the remaining requests. An existing output directory or take path is rejected. An interrupted or timed-out client does not prove the remote GPU operation stopped.

## Inspect what came back

The new directory contains `run.json`, four `take-s<seed>.mp4` files, a JSON receipt for each attempted take and `summary.json`. Receipts include request fields, UTC times, elapsed request time, byte count and SHA-256. Elapsed request time is not GPU inference time, billed time or an invoice. The manifest records expected server pins, not a verified runtime identity.

The client rejects non-200 responses, a MIME type other than `video/mp4`, oversized responses and bytes without an MP4 file-type header. That is transport validation, not a full decode or a quality verdict. Open every take, verify that the video decodes with the requested dimensions/frame count, and inspect the complete action before choosing it. A high-quality still does not ensure the video preserves its figure, prop or intended movement.

Keep the receipts and record any further trims, retiming, captions or audio separately. Stop resources you no longer need through your own hosting workflow. This helper does not stop or delete a server.
