Skip to content

The edit · AI video production · Now

Directing, with a new kind of crew.

Before software took over my days, I spent years designing, directing, and editing. That part of me never left. It just got a new crew: generated footage, synthetic voices, presenters who were never filmed, and motion graphics written in code. The job is the same as it always was. One person should understand one thing, and want to keep watching.

Before the new crew.

I taught myself design and video years before I wrote software for a living. It always started with a tool I needed for something of my own. Then the tool raised a question I couldn't leave alone: what problem was this invented to solve? Photoshop sent me to photo manipulation, and photo manipulation sent me to the science of images and design. Editing software sent me to directing, camera movement, color, and motion. Within a couple of months of learning each one, people were paying me for it.

I made ads for clothing and shoe stores and for restaurants, and designed restaurant menus, printed packaging, banners, and social media for local businesses. I handled all the media for a couple of student activities in college, and worked with charities and NGOs. When my town ran a campaign to repaint a wall with graffiti, three of my five designs were chosen and painted.

Then I learned to design teaching itself: what keeps the eye comfortable, what holds a student through a twenty-minute lecture, and where things belong on the screen. I produced the complete visual content of five courses for five different instructors, end to end. I also designed study workbooks for about ten private tutors.

  1. The tool

    Photoshop

    The problem it solves

    Photo manipulation

    The science under it

    The science of images

  2. The tool

    Illustrator

    The problem it solves

    Vector drawing

    The science under it

    Design itself

  3. The tool

    Editing software

    The problem it solves

    Cutting a story

    The science under it

    Directing, camera, color

Each tool, followed back to the problem it was made for.

The work, before software.

  • Store ads

  • Restaurant menus

  • Packaging

  • Social media designs

  • Banners

  • Business cards

  • Tutor workbooks

  • Student activities

  • Charities and NGOs

  • Graffiti wall, three of five designs painted

Ten kinds of work from the years before software: stores, restaurants, tutors, student activities, charities, and one wall in my town.
  1. Course 1

  2. Course 2

  3. Course 3

  4. Course 4

  5. Course 5

Five instructors, five courses, all the visuals end to end, about ten years ago.

The tools changed. The instinct under them didn't: find the problem the tool was built for, learn the science under it, then make something a person understands the first time they see it.

The edit log. Picture, voice, captions, and graphics, with a marker for each of the eight moments below.

  1. The two programs

    Right now I'm leading production on two training programs made with AI video. One is a free remote-work learning track, being rebuilt in Egyptian Arabic. The contract said fifty-five lessons. When I reconciled it against the real material, there were fifty-two: fifty teaching videos, one orientation, and one lesson that is only a document. The other is a new set of lesson films for students in grades six to eleven, starting with a two-minute welcome episode built on one idea: technology is the tool, not the hero.

    • Teaching videos50
    • Orientation1
    • Document only1
    • Counted, but not real3
    55 on paper, 52 in reality.One slot for every lesson the contract counted.
  2. How it started

    I didn't pitch with slides. I pitched with a finished eighty-nine-second sample, already in Egyptian Arabic, already cut. It mattered to me for the cause and for the people it reaches, and a finished sample says more than any promise.

  3. The voice that sounded Gulf

    I auditioned three Egyptian voices and picked one for its emotion and movement. My first sample with it came out sounding like the Gulf. The voice was fine. It had been paired with a newer speech model that pulled the accent, and the older model kept it Egyptian. Since then, no lesson ships until a native Egyptian listener approves the accent, whatever a provider's label says. A provider's label is not an ear.

    Label: Egyptian Arabic

    Heard: Gulf

    Approved by a native Egyptian listener

    The same take, labeled one way and heard another. The one that ships is the one a native ear approved.
  4. The letter ق

    Cairo Arabic softens the letter ق, so the easy shortcut is to soften it everywhere. That breaks the words that keep the classical sound, like the phrase for “the job market”. So there's no shortcut: an audited list of words, and a human ear on every script.

    سوق العمل

    The job market. Even in Cairo speech, this ق keeps its classical sound.

    • Softened: wrong here
    • Classical ق: right here
  5. Forty-two thousandths of a second

    One render came back with the narration 42.7 milliseconds late. Most people wouldn't be able to name it. They would just feel that the presenter was slightly off. I put the picture back on the original narration and checked the sync at five points across the film: zero offset at every one. Another early render came back almost silent, because seconds had been passed where frames were expected. Since then, “audio checked” means the audio is actually there.

    PictureVoice

    42.7 ms apart

    Picture and voice

    0 ms at 5 points

    The gap is drawn wider than life. At 42.7 milliseconds you don't see it. You feel it.
  6. Passing is not the same as good

    A draft that put white cards over a talking presenter passed every technical check, and I still sent it back. A lesson needs designed motion and visuals that actually teach, not decoration on top of a face. We rebuilt it in the language of the pilot instead. In the pilot itself, fast, noisy transitions became long scenes, soft one-second fades, and presenter moves timed to the chapter changes.

    Rejected

    Passed

    Every technical check, and still sent back.

    Designed

    The presenter makes room, and the idea takes shape.

    The same lesson, before and after. In the rebuild, every reveal on screen follows the narration.
  7. One picture, two voices

    For the new program's welcome episode we produced two versions with two different voices, on exactly the same picture. The frames are identical down to the last one. Only the voice changes, so the choice can be made on the voice alone.

    Picture

    Identical picture

    Voice A

    Voice B

    Two versions of the welcome episode. Same frames, a different voice, so the choice rests on the voice alone.
  8. The review room

    I also built the place where the client reviews everything: every script, voiceover, and cut. They can comment on a single line of a script, a single moment in the audio, or a single frame of the film. They can leave a voice note, reply in a thread, and approve a whole module, which then locks. And after every release I test the live site myself. Once that caught a check that had quietly read “available” from an empty answer. It was live for minutes.

    • 1A note on the presenter's voice
    • 2A note on the caption
    • A reply, in the same thread
    Feedback, pinned to the frame it's about, with the reply threaded under it.

Every tool, chosen for its limits.

Most of this work is knowing what each model will and won't do before asking it for anything.

  • Eight seconds at a time.

    The video model I use makes clips of about eight seconds. A presenter speaking for two or three minutes would take twelve to twenty-three separate clips, and the voice, the face, the gestures, and the lips would all have to match at every seam. So the background footage comes from that model, and the presenter comes from a tool made for one continuous take.

    One presenter, in eight-second pieces

    12 to 23 clips

    A seam: voice, face, gestures, and lips must match again

    One continuous take

    No seams

    An illustration, not to scale. Background footage comes from the video model. The presenter comes from a tool made for one continuous take.
  • Words go in code.

    I never ask a video model for readable text, a logo, or an Arabic caption. It will try, and it will get letters wrong. Every word on screen is set in code instead, exactly. And writing 8K in a prompt is a wish, not a setting. I check the file itself.

    Generated

    Letters guessed

    Set in code

    ق

    سوق العمل

    Every letter exact

    An illustration. The words on screen are never asked of the video model. They are set in code.
  • The voice sets the timing.

    The approved voice take comes first, and every picture is timed to it. Any sound the video model makes on its own is thrown away.

    Approved voice

    Picture, cut to the voice

    Generated sound: discarded

    An illustration. The cuts fall where the approved voice pauses.

Four phases, one room.

Every lesson passes through the same four phases. Each one asks the client about one thing, in the place where it lives: the words, the voice, the picture, then the whole film.

  1. 01

    Script

    The words, line by line.

    Comment on a single passage, in Arabic or English.

  2. 02

    Voiceover

    The voice, on its waveform.

    Comment on a moment in the waveform.

  3. 03

    Preview cut

    The picture, frame by frame.

    Comment on an exact frame.

  4. 04

    Final cut

    The whole film.

    Approve the module. It locks.

Approved modules lock.

Captions come last, once picture and voice are locked.

The tools change every month. The job is still to make one person understand one thing.

The camera side of this lives in another room: Everything, A to Z