Krust
Free guide · copy-paste everything

How to Make AI Video Look Real

The video below was generated from one paragraph of text, first take, on the cheapest preview tier. This page is the full method behind it: the realism checklist, the vocabulary, and the exact prompts.

“The fountain dare” · Seedance 2.5, 480p, 15s, one take, no reference footage. The full prompt is further down this page.

To make AI video look real, remove polish instead of adding quality. Prompt for one continuous handheld phone take with visible sensor noise, mixed practical light sources, one clumsy unflattering beat in the action, and people reacting in the background; direct the audio explicitly (voice character, room tone, natural dialogue); and never use cinematic vocabulary or platform names. Realism is subtraction plus specificity.

Why most AI video looks fake

Video models default to polish: cinematic lighting, smooth camera moves, perfect people mid-performance. Real phone footage has none of that. The fix is not asking for more quality; it is deliberately prompting the imperfections that phones and real moments produce. Every ingredient below is concrete and promptable.

The 6 realism ingredients

1

One continuous handheld take

No cuts. Whip pans, tilts, framing that keeps getting adjusted. Every cut is an edit, and edits mean production.

2

Two light sources fighting

Warm streetlamps against a cool glowing pool. Mixed practical light reads as unplanned; a single perfect key light reads as a set.

3

An event nobody would fake

Jumping into a fountain fully dressed. When the content itself is implausible to stage, viewers stop asking whether it is generated.

4

One clumsy, unflattering beat

The scramble out, dress clinging heavy, one knee on the stone. Staged videos cut this moment; real ones keep it. It is the strongest single realism signal.

5

People reacting in frame

Friends doubling over, a stranger half-turning to look. Social proof inside the shot does the arguing for you.

6

A payoff pose

Hair wring, hand on hip, big grin. It gives the clip an ending, and endings are what people rewatch.

The frame-by-frame method

The fountain prompt was not written from imagination. It came from studying a real viral clip frame by frame, then writing a brand-new scene with the same realism grammar:

  1. 1

    Save a real viral clip

    Pick a real video (not AI) in the style you want. The kind that made you ask "wait, is this real?" in reverse.

  2. 2

    Step through it frame by frame

    Screenshot 10-12 moments. Slow scrubbing shows you what full-speed watching hides: the camera wobbles, the light sources, the imperfect beats.

  3. 3

    Write down why it reads as real

    For every screenshot, name the realism signal: camera behavior, light, imperfection, background reactions. This list becomes your checklist.

  4. 4

    Reverse-engineer the original prompt

    Write the prompt that would have generated the original clip. This is practice; you will not use it. It forces you to translate observations into promptable language.

  5. 5

    Write a NEW scene against the checklist

    Different event, same grammar. Cloning the original gets you a worse copy; keeping the realism grammar with fresh content gets you your own winner.

The vocabulary bank

Put these in
  • raw vertical phone video
  • one single continuous handheld take
  • no cuts, no color grading
  • slight motion blur, visible night sensor noise
  • filmed by a friend who keeps adjusting framing
  • casual sway / natural handheld tremor
  • natural night lighting / one window light
  • lived-in mess: charger cable, cluttered counter
  • imperfect framing, subject off-center
Keep these out
  • cinematic
  • ultra photorealistic / 8K / hyperreal
  • studio lighting / editorial / premium
  • teal-and-orange grading
  • slow motion
  • “TikTok” / “Reels” (the model draws the app UI)
  • begging: “MUST look real”

Say “smartphone video”, never the platform name. Describe realism instead of pleading for it.

Direct the audio too

Sound sells the last 20%. Three parts, every time: a voice character written as a person (“warm female voice, mid-20s, casual, talking to a friend”), a room tone matched to the space (bathroom reverb, carpeted-bedroom deadness, echoey store), and dialogue that sounds spoken: contractions, fillers, fragments. “Okay so I’ve had this for like two weeks... and honestly? It just works.” A narrator sentence kills the whole illusion.

Four working prompts

All four generated real videos at the 480p tier, first or second take. Where a product appears, attach a clean photo as the reference image and describe yours from the actual photo.

The fountain dare · 15s · the video above
Raw vertical phone video at night, one single continuous handheld take, no cuts, no color grading, slight motion blur and visible night sensor noise, filmed by a laughing friend who keeps adjusting framing. A big European city square after midnight: a wide round stone fountain with tall center jets, lit from inside the water by cool white-blue underwater uplights; warm amber streetlamps and closed cafe signs around the square. A group of friends in going-out clothes stands at the fountain rim. A woman in her mid-20s in an emerald satin slip dress and bare feet, heels left on the stone, hands her little bag to a friend, swings a leg over the fountain rim while the group shrieks. She wades two steps into the glowing water and walks straight under the center jet, arms spread, one slow spin, completely soaked in seconds, screaming with laughter; the blue uplight and the amber lamps both play on the wet satin and her skin. Then she scrambles back over the rim clumsily, dress clinging heavy, one knee on the stone, a friend grabbing her arm and pulling. Standing on the pavement she bends forward, wrings a stream of water out of her long dark hair, flips it back and plants a hand on her hip with a huge grin while the group doubles over and a passing couple turns to look. The camera follows it all in one nervous handheld movement, whip-panning from the group to the rim, tilting with her under the jet, backing up for the pose. Audio: night square ambience, the steady rush of the fountain, overlapping shrieks and laughter from the group, a yelled 'no way, no way!', a distant moped, wet feet slapping stone.
Avoid: cinematic look, teal-and-orange grading, slow motion.
Bedroom yapper · 10s · product review
Raw vertical phone video, casual unedited feel, filmed on a front camera held at arm's length, strong natural handheld tremor throughout. A woman in her mid-20s with dark hair in a messy bun, oversized grey hoodie, sits on her bedroom floor leaning back against the bed. Soft daylight from a window on the left, lived-in room behind her: a charger cable on the carpet, a hoodie draped over the bed corner. She holds [YOUR PRODUCT - describe it exactly from a reference photo: shape, size, color, texture] in her right hand, resting it on her knee, tilting it slightly as she talks. Use the product from the reference image exactly as shown; reproduce shape, size, color and texture faithfully; the hand grip adapts to its true size. She talks straight into the camera. Warm female voice, mid-20s, casual, talking to a friend: "Okay so I've had this for like two weeks... and honestly? It just works. I don't know what to tell you." A small shrug, a tiny laugh at the end, lips gently closed between lines. Bedroom room tone: soft close acoustics, carpeted, minimal echo.
Supermarket aisle · 10s · walk and talk
Raw vertical phone video, front camera held at arm's length, walking slowly through a bright supermarket aisle, strong natural handheld tremor, fluorescent store lighting. A woman in her early 20s, blonde hair half-up, zip hoodie over a gym top, holds [YOUR PRODUCT] up next to her face. Use the product from the reference image exactly as shown; reproduce shape, size, color and texture faithfully. She talks to the camera like she's mid-story with a friend, half-amused. Bright female voice, early 20s, casual, a little conspiratorial: "Okay why does literally everyone at my gym have one of these... fine. I get it now." A small eye-roll smile on 'fine', lips gently closed after. Store ambience: open echoey retail tone, distant murmur, cart wheels.
Night walk · 8s · no-dialogue b-roll
Raw vertical phone video at night, handheld by a friend walking with the group, casual sway, natural night lighting from streetlamps and shop windows only, slight motion blur, no color grading. Three friends in their 20s walk down a city street after the gym, hoodies and gym bags, mid-laugh in an unheard conversation. The person in the middle carries [YOUR PRODUCT] swinging from one hand. Use the product from the reference image exactly as shown. Halfway through they lift it and take a sip without breaking stride. Nobody looks at the camera. No speech directed at camera: natural street ambience, distant traffic, footsteps, muffled laughter.

Run it in the canvas

The manual route needs a generation account, reference hosting and polling scripts. In Krust the same workflow is two nodes: a Text node with your prompt, wired into a Video node set to Seedance 2.5. For product versions, wire your product photo into the reference port.

The fountain prompt wired into a Seedance 2.5 video node on Krust's canvas

The exact fountain prompt from this page, wired and ready.

FAQ

Is the fountain video really 100% AI?

Yes. It was generated on Seedance 2.5 at the 480p tier, in one take, from the exact prompt published on this page. No reference video was used, and nothing was edited afterward.

Which AI video model is best for realism?

Seedance 2.5 currently produces the most convincing casual phone-video realism from pure text prompts, including believable audio. Kling 3.0 is strong for animating a still image, and Veo for single cinematic hero shots. All of them are available inside Krust's canvas.

Do these prompts work with a real product in the video?

Yes. Attach a clean photo of your product as a reference image and keep the "use the product from the reference image exactly as shown" line in the prompt. Describe the product from the actual photo, never from memory: a wrong verbal description overrides the reference.

Why does my AI video still look fake?

Almost always over-polish. Remove cinematic vocabulary, add handheld and sensor-noise anchors, put one clumsy beat in the action, and direct the audio explicitly. Realism comes from subtraction, and pleading ("must look real") actively hurts.