2026-07-15

NVIDIA TAO Agent Skills: One Day, 780 Scheduled A100-Hours

A source-pinned review turns NVIDIA's one-day Cosmos 3 result into a compute ledger: 4 scheduled A100-hours for one LoRA run versus a 780-hour AutoML envelope.

NVIDIA TAO Agent Skills: One Day, 780 Scheduled A100-Hours cover illustration

Source-pinned compute review

NVIDIA's “one day” label describes elapsed time; a scheduler sees forty A100s. The Cosmos 3 example reports a 43-trial AutoML sweep finishing in 19.5 hours across five parallel 8×A100 nodes. Reserving that shape for the stated interval creates a 780 A100-hour capacity envelope. The reported single LoRA run occupies about four.

A fresh source check on 24 July 2026 did not weaken that comparison. The TAO skill-bank head advanced to 5b4e574e9ecc48974c7e33d90550ea990bba80bd, yet a Git diff found no change in the Cosmos Reason skill subtree or version registry since the article's earlier pinned snapshot. All six current primary-source checks returned HTTP 200. The published validation score still moves from 87.14% after one LoRA run to 93.35% after the sweep: another 6.21 percentage points for a radically larger envelope.

The scheduler, not the stopwatch, owns this decision.

Retained result: a source-pinned launch-ledger probe matched all 22 expected decisions. One complete packet reached launch review; 21 mutations were blocked for missing source hashes, data separation, objective binding, baseline, test seal, compute budget, approval, or retained outputs. Fixture SHA-256: d9df01945b494a372a482d44561f2b796bcbeaa7229c0e7ffc9df961ddfead6d.

Nineteen and a half hours can reserve forty GPUs

The official NVIDIA post is unusually candid once it reaches the hardware section. A zero-shot Cosmos 3 Nano baseline scored 54.41% exact-match accuracy on a four-choice video question-answering task. One LoRA run reached 87.14% after about 30 minutes on eight A100 80 GB GPUs. The larger AutoML experiment processed 43 trials in 19.5 hours using five parallel nodes, one 8×A100 node per search strategy, and reported 93.35% validation accuracy.

Those are different resource stories hidden inside the same phrase, “in one day.” The small run is one node for half an hour. The sweep is a forty-GPU cluster operating across most of a day. Wall time is useful for planning a deadline; scheduled capacity is what procurement, quotas, and failure recovery must absorb.

4scheduled A100-hours for the reported single LoRA run
780scheduled A100-hour envelope for five nodes across 19.5 hours
+6.21 ppvalidation gain from the LoRA result to the reported AutoML best

The 780 figure is capacity arithmetic, not a claim that every GPU was busy for every second. NVIDIA does not publish per-trial utilization, queue idle time, energy, failure retries, or cloud price in the post.

Requirements for an honest capacity ledger

Reported runWall time and shapeScheduled-capacity arithmeticPublished result
Zero-shot baselineEvaluation environment not fully costed in the postNot computable from the published dimensions54.41% validation accuracy
Single LoRA0.5 hours × 8 A1004 A100-hours87.14%
AutoML envelope19.5 hours × 5 nodes × 8 A100780 A100-hours if fully reserved93.35%
Full-parameter SFT reference3 hours 34 minutes × 8 H10028.53 H100-hoursDifferent accelerator; not a direct cost comparison
A thin blue wall-clock band crossing a much larger field of amber scheduled compute blocks
The same deadline can hide very different capacity commitments. The larger block field is an envelope, not measured utilization.

The H100 reference belongs in a separate column because an H100-hour and an A100-hour are not interchangeable units. Hardware generation, software stack, utilization, batch shape, and price all differ. A tidy ratio between the two would manufacture precision the source does not provide.

That unit mismatch matters.

The repository moved; the Cosmos contract did not

We rechecked the TAO skill bank at current commit 5b4e574e9ecc48974c7e33d90550ea990bba80bd, not an unpinned installer. The repository head changed after the earlier snapshot at c6b8f24e4a433a073a3da93a4cbb29865a217b57, but the Cosmos Reason subtree and versions.yaml did not. The skill still defaults to nvidia/Cosmos3-Nano, declares AutoML support, names the Hugging Face credential, and recommends an 8×A100 or H100 80 GB node.

For an accuracy objective, the current source still says to optimize an evaluation metric rather than val/avg_loss. It requires a baseline evaluation before recommendation jobs. Before launch, it asks the agent to surface recommendation count, search parameters, ranges, dataset subset, runtime per recommendation, and total runtime. The current TAO reference adds a blunt backend rule: run one AutoML job at a time per backend. Natural language may request a search; it does not define the queue policy or budget.

Important version boundary: this review binds the skill contract to commit 5b4e574e9ecc48974c7e33d90550ea990bba80bd, checked on 24 July 2026. Skill behavior, container compatibility, schemas, and conversion requirements can change after that point. Record both the repository commit and resolved container versions before launch.

The expensive question begins after 87.14%

The first LoRA run accounts for most of the published lift: 54.41% to 87.14%, or 32.73 percentage points. AutoML adds 6.21 points beyond that result. Whether the second jump is worth the larger envelope depends on the error distribution, business loss, serving constraints, and repeatability—not on the fact that the sweep fits inside one date on a calendar.

A useful approval separates exploration from promotion. Approve one bounded LoRA run and its baseline comparison first. Read the errors, not only the aggregate. If the remaining failures matter, approve a defined search budget with an early-stop rule. Do not let an agent infer that “improve the result” authorizes 43 trials, five nodes, a new external LLM endpoint, or an open-ended search space.

The search itself can also overfit the validation set. Forty-three recommendations create forty-three opportunities to select noise. The best validation checkpoint is still a candidate until a sealed test set opens once, after the search is over.

Troubleshooting the packet before the scheduler does

Our local probe encoded one minimum packet and then removed or altered one control at a time. The valid packet reached READY_FOR_LAUNCH_REVIEW. Every other case stopped before a GPU or credential was touched.

Source identityExact skill-bank commit, hashes for the model skill, AutoML policy, structured metadata, train schema, and version registry.
Data boundaryDistinct train, validation, and sealed-test locations; content hashes; license acceptance; no silent annotation repair.
ObjectiveAccuracy, maximize direction, retained zero-shot baseline, and one agreed scoring function used across recommendations.
Capacity envelopeNode type, GPUs per node, wall-clock ceiling, trial count, parallelism, stop condition, and who approved the budget.
Search contractNamed parameters and ranges. Fixed epochs or warmup stay fixed unless the reviewer explicitly expands them.
Retained outputsBest config, weights, validation results, sealed-test result, trial ledger, failures, and cost/capacity ledger.
A sealed blue launch packet receiving verified evidence routes while incomplete red routes stop at a compute gate
A natural-language prompt can request the run. It should not replace the packet that binds source, data, metric, capacity, approval, and outputs.

The sealed set gets the final vote

NVIDIA labels the reported 93.35% as validation accuracy. That is the correct label to retain. A reproducible promotion decision needs an untouched test set or a later prospective evaluation that the search process never saw. Reusing validation for both selection and final claims turns the winning trial into its own examiner.

The Woven Traffic Safety dataset and the Cosmos 3 model collection establish useful source context, but they do not supply a team's private acceptance threshold, licensing record, data lineage, or cost ceiling. Those belong in the local packet.

Promotion should therefore compare three retained objects: the zero-shot baseline, the chosen LoRA or AutoML checkpoint on the same validation contract, and one sealed test result. If the test regresses, the launch ledger should say so even when the validation chart looks impressive.

What we did not pretend to reproduce

The retained artifact is static source analysis plus deterministic JavaScript. It did not log into NVIDIA, NGC, Hugging Face, or a cloud provider. It did not download the WTS dataset, pull a container, obtain model weights, reserve a GPU, run Cosmos 3, reproduce any NVIDIA accuracy number, execute AutoML, observe utilization, or calculate a bill.

The capacity ledger verifies arithmetic against NVIDIA's published dimensions. The 22-case admission probe verifies that a proposed packet fails closed under the rules encoded here. Neither is a training benchmark. The next honest step is a separately authorized baseline and one bounded LoRA run with real telemetry—not a claim that the source probe reproduced 93.35%.

Decision: keep the agent skills, but approve the compute envelope in stages. One LoRA result earns the right to request a sweep; it does not silently authorize the sweep.

Sources and retained checks

  • NVIDIA developer blog: Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills — reported accuracy, trial count, runtime, node shape, and hardware comparison.
  • NVIDIA TAO skill bank — current source check at commit 5b4e574e9ecc48974c7e33d90550ea990bba80bd; the relevant Cosmos subtree and version registry were unchanged from c6b8f24e4a433a073a3da93a4cbb29865a217b57.
  • NVIDIA TAO AutoML documentation — current algorithms, prerequisites, objectives, launch fields, and one-run-per-backend note.
  • Retained launch-ledger fixture and current source check — 22 expected admission decisions with zero mismatches, six live primary sources with HTTP 200, an unchanged relevant repository contract, and recomputed 4-versus-780 scheduled-capacity arithmetic. Fixture SHA-256: d9df01945b494a372a482d44561f2b796bcbeaa7229c0e7ffc9df961ddfead6d.