2026/08/12

Text-to-Video vs Image-to-Video: How to Start an AI Video

Text-to-video starts with a description; image-to-video starts with a reference image. A practical guide to picking the right one — both live in Orelon's one workspace.

Text-to-video vs image-to-video: which should you start with?

There's a moment in every AI video workflow where you choose how to begin: type a description, or supply an image. It decides whether your video is anchored to a look you already have, or starts from nothing but words.

Text-to-video and image-to-video are the two main ways to start, and they solve different problems. Here's how to tell them apart, and when each one wins.

The difference in one sentence each

  • Text-to-video turns a written description of a scene into motion. The only input is your words.
  • Image-to-video animates a still image, adding motion and camera moves while keeping that image as the visual anchor.

Both run in the same video workspace in Orelon, sharing one prompt field, model picker, and refinement loop. Switching between them doesn't mean rebuilding your setup.

Start with text when…

  • You're exploring. You have an idea — "golden hour, a figure walking into fog" — and you want to see what the models make of it.
  • You don't have a visual reference yet. For atmosphere, mood, and first drafts, a description is enough.
  • You want to iterate fast. Text is cheap to change. Tweak one phrase and regenerate.

Start with an image when…

  • Consistency matters. A product, a character, a brand look — the model works from that exact image, so the result stays on-brand.
  • You already have the shot. A photo you took, a frame you like, a design you made. You're adding motion, not inventing a scene.
  • You're making something commercial. Product videos almost always want an image anchor, so the output matches the real item.

The middle option: first-last-frame

Want more control than text but don't want to hand over the whole scene? First-last-frame sits in between. You supply a start frame and an end frame, and the model generates the motion between them. It's a precise way to direct a shot — a hero product opening its box, a character turning to camera.

A quick way to choose

Ask one question: do I have a specific look I need to match?

  • Yes → start with an image, or with first-last-frame.
  • No → start with text and let the generation give you something to react to.

Both paths live in the same workspace, so this isn't a commitment. Try text first; if the result is close but the look is off, add a reference image and regenerate.