From Script to Final Cut - How I Made a Product Video with AI

By Feng Qiu
July 23, 2026
Artificial Intelligence (AI)
💡
TL;DR
• AI did not generate a finished video for me in one shot. What it really reduced was the cost of iteration: I could identify a problem, describe the gap, and see a revised version quickly. • The production method should follow the product. For a real app, a stable interface and controllable animation matter more than generated people, hands, or elaborate scenes. • A script is not merely voiceover copy. It is a set of decisions about what the product is and what deserves attention. In a 40-second video, choosing what to leave out matters as much as choosing what to include. • A product video is not an App Store screenshot with motion added. Full-screen product views establish context, close-ups explain value, and titles and animation must keep contributing new information. • The available assets should be allowed to change the script. If a feature cannot be captured truthfully, invented visuals should not be used to fill the gap. • The easier revision becomes, the easier it is to iterate forever. I eventually relied on a few stable stopping criteria: the visuals had to be authentic, the claims accurate, every shot purposeful, and the overall tone consistent with the product.

When DailyTrace 1.7 was nearly ready for release, I decided to make an English-language product video for it. The initial brief sounded simple: about a minute long, vertical, and able to explain what the app is, how it works, and how it differs from an ordinary productivity tool. Once I considered the pace of social media, I shortened it to 40 seconds, set the main output to 1080 × 1920, and planned a second version at 1242 × 2688.

I had no professional experience in video editing or motion design. While building DailyTrace, however, I had already grown accustomed to using Codex to clarify requirements, write code, review design decisions, and prepare App Store materials. That led to an obvious question: if someone who did not know Swift could use AI to build an app, could I use a similar approach to make its product video?

The answer was yes, but the process looked nothing like “generating a video from a single prompt.” I initially expected animation, transitions, and export to be the difficult parts. Once production began, making things move turned out to be relatively easy. The harder work was deciding what to show, what each shot should help the viewer understand, and whether something that was technically complete actually worked as a viewing experience.

The path from goal and script to final cut had clear stages, but it was never a straight line. The availability of authentic footage, the clarity of a keyframe, or the mood of the music could send the project back to an earlier stage.

Article image

1. Choose a Production Method That Fits the Product

I began by researching different approaches to AI video production. The most obvious option was text-to-video: generate a scene of someone taking out a phone in a coffee shop, logging their work, and checking the day's statistics later on. It would look more like a conventional advertisement, but it was a poor fit for this project.

DailyTrace is a real product. The text, controls, and data on screen need to remain stable, while people, hands, and phone interactions are exactly the elements that generative video struggles to keep consistent. A scene can have a wonderful atmosphere, but once the app interface begins to mutate, the demonstration loses the credibility that matters most.

I eventually chose HeyGen's HyperFrames as the core production environment. It is not a tool that turns one description into an entire video. It is closer to building motion with web technologies and code: text, images, and real product interfaces are placed on a fixed canvas, then animated and transitioned on a timeline.

DailyTrace already contains timelines, colored activity blocks, and data visualizations, which made it a much better fit for this approach than for a performance-driven commercial. I captured real screens from the iOS Simulator, used HyperFrames for layout and animation, generated the English voiceover with HeyGen, and relied on Codex to help organize the storyboard, revise scenes, and prepare the final exports.

The main advantage was not novelty. It was control. If a title sat too low, I could move it. If a screen was cropped too aggressively, I could restore the full interface. If a shot lingered for too long, I could change its duration directly.

That gave me a more useful way to evaluate production tools: do not begin by asking how complex a scene a tool can generate. Ask whether it can present the product's most important details reliably.

2. Lock the Script Before Starting Production

The first voiceover followed a familiar product-demo structure. It opened with a problem, demonstrated Start, Switch, and Stop, then moved through Day, Week, Month, Past Year, Widget, Live Activity, Dynamic Island, and Apple Watch. It ended with a privacy message and a download prompt.

As a feature list, it was comprehensive. But it failed to answer the more important question: after 40 seconds, what should a viewer remember about DailyTrace? If every feature received a second or two, the video could show almost everything while leaving no clear impression of the product.

Before formal production, I therefore introduced a script approval gate. The script package included more than the English voiceover. For every line, it recorded the intended meaning in Chinese, its narrative purpose, estimated duration, corresponding shot, and supporting product evidence. Until I explicitly approved the script, there would be no final voiceover, no footage capture, and no animation production.

This extra step looked like overhead, but it prevented much larger revisions later. Once a scene has been built around a line of narration, changing the line can also change the shot length, animation timing, captions, and musical cues.

The approval process also forced me to examine every product claim. “No guessed gaps” means that DailyTrace does not infer what happened during unrecorded time; it does not mean the app automatically understands a user's life. “Your records stay on your device” accurately describes the current data model, but it cannot be expanded into unsupported promises about encryption, cloud backup, or synchronization.

A script needs to sound good, but that is only the first test. More importantly, it must define the product accurately, and every claim must be demonstrable with real visuals.

3. Select the Features That Best Represent the Product

Approving the script did not end the decision-making. Once the first keyframes were ready, I realized that Frame 3 was weak. It focused on Start, Switch, and Stop. These actions matter—they are the basis of using DailyTrace—but they do not distinguish it clearly from a conventional timer.

A feature can be important without deserving one of the few major shots in a 40-second product video.

I reconsidered which views best expressed what DailyTrace is, without limiting the choice to features introduced in version 1.7. I settled on three main visuals: When It Happens, Focus Flow, and Activity Detail.

When It Happens reveals when different activities tend to occur during the day. Focus Flow shows how recorded activity changes across a year. Activity Detail lets someone explore the recent trend for a single activity, its accumulated time, and what usually happens before or after it. Together, these screens answer one central question: how does time that a person records deliberately become a picture they can revisit and understand?

This changed the video's narrative. Instead of distributing attention evenly across the feature set, it began with “Where did your time go?” It defined DailyTrace as a time journal built from deliberate records, then moved through daily patterns, year-long change, and individual activity details.

The change taught me that a product video is neither a release note nor a contest to display the largest number of features. A shot should be chosen because it represents the product, not because it is easy to make or happens to be new in the current release.

4. Capture Authentic Assets—and Accept Their Limits

With the narrative in place, I began capturing real product screens from the iOS Simulator: Home, Month, Past Year, and Activity Detail. The original plan also called for the Widget, Live Activity, Dynamic Island, and Apple Watch.

“Authentic” meant more than making a screen look like DailyTrace. The footage could not include the Internal build name, test utilities, a mouse pointer, simulator chrome, the wrong language, or any other trace of development.

In practice, the Apple Watch simulator produced only a disconnected recovery screen. I could not capture clean, isolated English views of Live Activity or Dynamic Island consistently either. The tempting solution was to redraw them or ask a generative tool for something similar. But this was a product video. A beautifully fabricated system surface was still less valuable than a smaller set of screens that genuinely existed.

I removed the system surfaces I could not capture reliably from both the script and the storyboard, leaving only assets that I had verified.

That changed how I understood the relationship between script and footage. The script determines which assets are needed, but the assets that can be captured truthfully must also be allowed to change the script. Product-video production is not a one-way path from words to finished film. It is a continuing negotiation among product facts, available material, and the communication goal.

Authentic assets constrain the creative process, but that constraint prevents a video from gradually inventing features simply to appear complete.

5. Establish the Full Interface Before Moving Into Detail

The first visual draft revealed an obvious problem: the app interface was never shown in full. To make the charts larger, I had cropped the source screenshots immediately. Viewers could see a detail but had no idea which screen it belonged to or how someone would reach it. The result felt like an enlarged image, not a real product.

I added a consistent phone frame to every full-app view, then divided the longer Frames 4 and 5 into two parts: full screen and close-up. The first part showed the entire Past Year or Activity Detail screen, including its navigation, title, and the chart's location. The second moved through the phone display and brought Focus Flow, Before & After, or Activity Trend into a full-frame view.

The full interface establishes the authenticity and context of the product. The close-up makes its most important information legible. A detail does not have to remain trapped inside a phone frame, but the preceding shot must first establish how it relates to the complete product.

I initially tried red outlines to call attention to important charts. They were unambiguous, but they made the video feel like a tutorial. I removed them and used scale, screen movement, title changes, and short pauses to guide attention instead.

Here, motion was not decoration or a way to avoid stillness. It explained where a feature lived and why the viewer needed to see it more closely.

6. Make Titles and Layouts Carry Information

Titles were another recurring source of revision. Early versions included phrases such as “The Next Question,” “Month View,” “The Whole View,” and “Look Closer.” None was grammatically wrong, but none added much. “Month View” merely named the screen already visible, while “Look Closer” asked for attention without explaining why the next detail deserved it.

I returned to DailyTrace's App Store screenshots. Their headlines do not repeat screen names. They first describe what someone can gain from the view, then let the interface serve as evidence.

The video titles gradually became:

  • “YOUR DAY, AS IT HAPPENS.”
  • “THE MONTH TAKES SHAPE.”
  • “A YEAR, IN FULL COLOR.”
  • “YOUR FOCUS, OVER TIME.”
  • “PATTERNS, MADE VISIBLE.”

They remained short, but now they contributed to the story. Instead of naming a page, each title helped the viewer understand why the page might matter to them.

Once the titles improved, the layout problems became easier to see. Frames 1 and 2 were too similar, while Frames 3, 4, and 5 repeated nearly the same composition. Individually, each shot looked acceptable. In sequence, they felt like the same poster appearing again and again.

I reviewed the typographic hierarchy, safe areas, alignment, and visual balance against Apple's design guidance, while varying the direction and composition from scene to scene. Video design cannot be judged one frame at a time. The contrast between adjacent shots is part of the design too.

7. Let the Visuals Advance With the Voiceover

Frames 4 and 5 carried more information and therefore stayed on screen longer. In the first draft, however, almost nothing changed after the page appeared. The static keyframe looked composed in isolation, but it became tedious during playback. The problem was not that the image was badly designed. The narration had moved to the next idea while the visual remained stuck on the previous one.

After splitting the scenes, visual changes began to correspond to changes in meaning. When the viewer needed time to read, the image held still. When the narration moved to another layer of information, the camera moved with it.

This taught me that pacing does not mean keeping every element in constant motion. Effective pacing means introducing a meaningful visual change when the voiceover adds new information—and allowing the frame to pause when the audience needs to understand what is already there.

8. Choose a Voice and Music That Belong to the Product

I generated the voiceover with HeyGen and auditioned Camden, Marcia, and Skylar before choosing Marcia. Her voice was clear and carried the story forward, but it did not sound like a conventional commercial announcer. Nor did it turn “No account. No scores. No guessed gaps.” into a security warning.

The generated narration included word-level timestamps. That meant the animation no longer had to be designed around an approximate five-second block. It could follow the actual delivery. By the time the voice said “past year,” the Past Year screen needed to be readable; the move into Activity Trend should begin only when she said “recent trend.”

Music was harder than I expected. The first version was too energetic and competed with the narration. After I reduced the intensity, the next version became too slow and carried an unintended hint of horror-score tension. I eventually converged on something bright, restrained, and moderate in tempo. Clicks, switches, and the completion of the logo animation received a few sound effects, but I did not add a sound to every movement.

Those revisions showed me that “slower,” “lighter,” or “more energetic” are not useful enough as music directions. A brief should also say whether the track competes with the voice, whether its mood leans too dark, whether the percussion is too forceful, and how the sound relates to the product's character.

9. Review the Exported Video From Beginning to End

HyperFrames Studio makes it easy to inspect a shot, scrub through the timeline, and review keyframes and transitions. It is useful for pausing at an exact moment and checking the layout. But something that looks correct in Studio is not guaranteed to work in the exported MP4.

Many problems only appear during continuous viewing: a shot stays on screen too long, three consecutive scenes use overly similar compositions, a crop feels abrupt while it moves, or the emotional tone of the music changes unexpectedly.

Late in production, I also discovered a strip of empty space along the bottom of every framed phone screen. The screenshots themselves were fine. The content dimensions of the phone frame did not precisely match the images' aspect ratio. It was easy to miss in the scene code and preview, but it damaged the illusion of a real device in the final export.

After fixing it, I extracted several phone keyframes from the final MP4 and checked that each screenshot reached the bottom edge of the frame. I then reviewed the captions, safe areas, duration, resolution, frame rate, and audio track.

Automated checks helped. They could reveal runtime errors, out-of-bounds elements, and certain layout problems. They could not tell me whether a title was meaningful or why a piece of music felt wrong. The final video still had to be watched from beginning to end and paused for frame-by-frame inspection at key moments.

As with an app, “it runs” is only the beginning of acceptance testing.

The Video Was Not Generated. It Was Revised Into Existence.

What AI changed most was the cost of revision. In a conventional workflow, deciding that one shot needed to become two might mean rearranging the edit, animation, and sound. Here, I could identify the specific problem, ask Codex to adjust the relevant scene and timeline, and see another version quickly.

If a title carried no information, I rewrote it. If the relationship between a detail and the product was unclear, I added a full-screen view first. If the footage could not support a line of narration, I returned to the script. If the music felt emotionally wrong, I auditioned another direction.

AI did not let me skip any of these problems. It shortened the distance between “something is wrong here” and “now I can see another way to do it.”

That made the production process different from what I had imagined. I did not first conceive a complete, correct video and then ask AI to execute it. Many judgments became possible only after I saw a concrete result. Making and evaluating were not separate phases; they alternated throughout the project.

Easier revision creates its own problem: a video can always be adjusted again. The title could change, another track could be tested, and a transition could always have one more variation. Without stable criteria, faster iteration becomes nothing more than faster indecision.

What eventually allowed me to stop were the standards that had remained constant throughout production: the visuals had to come from the real product; every shot had to add information; every line of narration had to be supported by product facts; and the finished video could not drift away from the character of DailyTrace.

After this project, I no longer think of “AI video production” as writing a prompt and waiting for a finished film. For a product video, at least, it is a different way of working: turn the script, authentic assets, and storyboard into something watchable as early as possible, then revise against a concrete result.

In the past, not knowing editing or animation might have stopped me at the idea stage. This time, every “not quite right” could become another version.

The 40-second video is merely the final file that remains. The deeper result was completing the entire journey from understanding a product to expressing it through image and sound. When I build an app, I decide what the product should become. When I make its video, I decide what within that product is most worth seeing.

AI made it possible for one person to keep moving through all of those stages. What I ultimately learned was how to compress a product into 40 seconds without inventing it—or losing what made it distinctive in the first place.

Share this article

© 2026 Feng Qiu. All rights reserved.