Your First AI Video: Text-to-Video Basics
- Write a text-to-video prompt with all four required components
- Generate a clip on Kling or Pika using the free tier
- Identify and avoid the four most common beginner prompt mistakes
- Use iteration (multiple variations) to improve results
What Text-to-Video Actually Means
Text-to-video is the simplest entry point to AI video generation: you type a description and the AI generates a video from scratch — no input image, no reference footage. It sounds simple, and it is the easiest way to start. It's also the hardest to control, because the AI has to invent everything from nothing.
The key insight: a text-to-video prompt is not the same as an image prompt. You are writing a brief for a cinematographer, not describing a painting. The AI needs to know what to film, how to film it, and what the world looks and feels like.
The Anatomy of a Working Video Prompt
Every reliable text-to-video prompt includes four components:
- Subject and action — Who or what is in the frame, and what are they doing? Be specific: "a woman laughing" is better than "a person smiling."
- Camera behavior — How is the camera moving? "Static shot," "slow zoom in," "camera pans left to reveal" are all valid.
- Environment — Where is this happening? Interior or exterior? What is the light source?
- Style and mood — Cinematic, documentary, handheld, animated? Name the feeling.
A young chef in a modern kitchen slices vegetables in slow motion, overhead shot looking straight down, warm tungsten lighting, cinematic depth of field, professional food photography style.
Every element of that prompt answers one of the four questions. That's why it works.
Generating Your First Clip on Kling
Kling (kling.ai) has the best free tier for text-to-video. Here's the workflow:
- Create a free account at kling.ai
- Click AI Video → Text to Video
- Paste your prompt into the description field
- Choose aspect ratio: 16:9 for landscape, 9:16 for vertical and social media
- Select duration: 5 seconds for quick tests, 10 seconds for more developed scenes
- Click Generate and wait 60–120 seconds
Use 5-second clips for experimentation — they cost fewer credits and let you test more prompts. Save 10-second generation for prompts that have already worked at 5 seconds.
Common Beginner Mistakes
These mistakes produce consistently bad results:
- Writing image prompts — "A beautiful mountain at sunset with snow and trees" describes a painting, not a scene. Add action and camera movement.
- Too many subjects — One or two subjects with clear actions work best. Five characters usually produce chaos.
- No camera instruction — Without camera guidance, the AI defaults to static or random movement. Always specify.
- Expecting perfect first results — Text-to-video requires iteration. Generate 3–5 variations before deciding the concept doesn't work.
Try This Now
Generate these two prompts and compare the results:
A campfire burning at night in a forest.
A campfire burns in a dark forest clearing at night, camera slowly pulling back to reveal the surrounding trees, orange flickering light, handheld, cinematic, peaceful and mysterious mood.
The second prompt will almost always produce a better clip. Keep both — you'll use them to build your prompt instincts as you work through this course.
- Text-to-video needs subject+action, camera behavior, environment, and style to work reliably
- Write for a cinematographer, not a photographer
- Always specify camera behavior or the AI will default to static or random movement
- 5-second test clips save credits during iteration