I Tested AI Agents Against a Real AI Artist
Gives AI filmmakers an objective look at current frontier agent capabilities (Astra, Luma, Runway connectors), benchmarking generation costs, pipeline autonomy, and creative quality against human-directed workflows.
What this lesson covers
Curious Refuge compares human-directed creative workflows against automated AI agents (Luma's agent and ChatGPT Astra with a Runway connector) to produce a 30-second commercial. The test evaluates the execution speed, compute cost, brand accuracy, and emotional resonance of autonomous agent pipelines against human artists.
Key takeaways from the creator
AI-extracted notes, not independently verified product claims. Timestamp links let you check each point in the original video.
- 01:18 ↗
Application-based AI agents and frontier model agents can autonomously process a creative brief into keyframes, video clips, and edited video files.
- 03:35 ↗
Luma's automated creative agent generated a 30-second commercial in 15 minutes for $14, but exhibited brand inaccuracies and visual artifacts.
- 05:23 ↗
ChatGPT Astra connected to Runway completed a full ad generation and edit in 17 minutes and 22 seconds for $20 ($12 Runway credits, $8 Astra compute).
- 06:12 ↗
Frontier agents can operate desktop video editors like Adobe Premiere Pro and DaVinci Resolve or execute automated video editing pipelines using JavaScript.
- 07:23 ↗
Autonomous AI agents currently excel at rapid mood reels and baseline concepts, but human creative direction remains superior in conceptual depth and emotional resonance.
Workflow outlined in the video
- Prepare a standardized creative brief detailing target audience, visual style, and reference product imagery.
- Upload the creative brief PDF and supporting assets into an AI agent environment like Luma or ChatGPT.
- Configure ChatGPT in the Work tab using ChatGPT Astra on extra high effort with external connectors such as Runway.
- Execute autonomous prompts to let the agent create shot keyframes, render video assets, and edit the final sequence.
- Benchmark the resulting generation times, compute costs, and asset consistency against human-crafted workflows.
Before you use this workflow
These notes describe the source video at its publication date. Model access, pricing, connectors and interfaces may have changed. Check the original source and the provider’s current documentation before installing an add-on, connecting an account or spending credits.
ReelStack has not independently tested this workflow. Preview one representative shot and check motion, continuity and output quality before applying it to a full production. No result, cost saving or model capability is guaranteed.
Explore the tools
How to interpret the numbers
6,775 views captured 6 October 2026. ReelStack recommendation score: 36/100. Formula: radar-v5. This is ReelStack’s calculation, not a YouTube rating or a measure of factual accuracy.
Show the calculation
- 65% performance: 0.29× current views vs the channel's recent median; same-age history is not yet available. View-evidence factor 87% (views / (views + 1,000))
- 25% freshness: 76/100 with a 14-day half-life
- Momentum pending: collecting daily snapshots; its 20% weight goes to observed performance, not free points
- 10% engagement: 40/100; likes + 4× comments, smoothed with a 500-view neutral prior