How voice to video works in Planetanium
Most AI video generators start from a text prompt and hand you a clip you didn’t direct. Planetanium starts from your voice and keeps you in charge the whole way:
- Upload a recording — MP3, WAV, M4A, or OGG. A voice memo is enough. No recording? Type your script and generate narration with selectable AI voices.
- The AI transcribes it — automatically, with word-level editing and speaker detection, so multi-voice sketches and dialogue work too.
- An agent storyboards it — your narration is split into timed scenes and shots on a waveform timeline. Every shot is synced to the exact moment in your audio where it belongs.
- Images and motion are generated — with consistent characters, locations, and a project-wide visual style you choose: anime, watercolor, cinematic, claymation, and more.
- Export the film — MP4 or WebM up to 1080p, square, widescreen, or vertical, with optional animated captions — or hand the whole timeline to DaVinci Resolve or Final Cut Pro as FCPXML.
Stay the director
The agent does the technical work; you keep the intention. Regenerate or restyle any shot, nudge timing on the timeline, correct a word in the transcription, or just tell the project-aware assistant what to change. The result stays synced to your performance — because the performance is the spine of the film.
What people make with it
Narrated stories and audio sketches, YouTube videos and Shorts, animatics for pitching, visualized podcasts, storyboards with a soundtrack. If it starts as a voice, Planetanium can make it watchable.
There’s no subscription — pricing is pay-as-you-go for the AI you actually use. See the full feature list or the FAQ for details.