Directing, with a new kind of crew.
Before software took over my days, I spent years designing, directing, and editing. That part of me never left. It just got a new crew: generated footage, synthetic voices, presenters who were never filmed, and motion graphics written in code. The job is the same as it always was. One person should understand one thing, and want to keep watching.
Before the new crew.
I taught myself design and video years before I wrote software for a living. It always started with a tool I needed for something of my own. Then the tool raised a question I couldn't leave alone: what problem was this invented to solve? Photoshop sent me to photo manipulation, and photo manipulation sent me to the science of images and design. Editing software sent me to directing, camera movement, color, and motion. Within a couple of months of learning each one, people were paying me for it.
I made ads for clothing and shoe stores and for restaurants, and designed restaurant menus, printed packaging, banners, and social media for local businesses. I handled all the media for a couple of student activities in college, and worked with charities and NGOs. When my town ran a campaign to repaint a wall with graffiti, three of my five designs were chosen and painted.
Then I learned to design teaching itself: what keeps the eye comfortable, what holds a student through a twenty-minute lecture, and where things belong on the screen. I produced the complete visual content of five courses for five different instructors, end to end. I also designed study workbooks for about ten private tutors.
The tool
The problem it solves
The science under it
Photoshop
Photo manipulation
The science of images
Illustrator
Vector drawing
Design itself
Editing software
Cutting a story
Directing, camera, color
The work, before software.
Store ads
Restaurant menus
Packaging
Social media designs
Banners
Business cards
Tutor workbooks
Student activities
Charities and NGOs
Graffiti wall, three of five designs painted
Course 1
Course 2
Course 3
Course 4
Course 5
The tools changed. The instinct under them didn't: find the problem the tool was built for, learn the science under it, then make something a person understands the first time they see it.
The two programs
Right now I'm leading production on two training programs made with AI video. One is a free remote-work learning track, being rebuilt in Egyptian Arabic. The contract said fifty-five lessons. When I reconciled it against the real material, there were fifty-two: fifty teaching videos, one orientation, and one lesson that is only a document. The other is a new set of lesson films for students in grades six to eleven, starting with a two-minute welcome episode built on one idea: technology is the tool, not the hero.
- Teaching videos
- Orientation
- Document only
- Counted, but not real
55 on paper, 52 in reality. How it started
I didn't pitch with slides. I pitched with a finished eighty-nine-second sample, already in Egyptian Arabic, already cut. It mattered to me for the cause and for the people it reaches, and a finished sample says more than any promise.
The voice that sounded Gulf
I auditioned three Egyptian voices and picked one for its emotion and movement. My first sample with it came out sounding like the Gulf. The voice was fine. It had been paired with a newer speech model that pulled the accent, and the older model kept it Egyptian. Since then, no lesson ships until a native Egyptian listener approves the accent, whatever a provider's label says. A provider's label is not an ear.
Heard: Gulf
Approved by a native Egyptian listener
The same take, labeled one way and heard another. The one that ships is the one a native ear approved. The letter ق
Cairo Arabic softens the letter ق, so the easy shortcut is to soften it everywhere. That breaks the words that keep the classical sound, like the phrase for “the job market”. So there's no shortcut: an audited list of words, and a human ear on every script.
سوق العمل
- Softened: wrong here
- Classical ق: right here
Forty-two thousandths of a second
One render came back with the narration 42.7 milliseconds late. Most people wouldn't be able to name it. They would just feel that the presenter was slightly off. I put the picture back on the original narration and checked the sync at five points across the film: zero offset at every one. Another early render came back almost silent, because seconds had been passed where frames were expected. Since then, “audio checked” means the audio is actually there.
The gap is drawn wider than life. At 42.7 milliseconds you don't see it. You feel it. Passing is not the same as good
A draft that put white cards over a talking presenter passed every technical check, and I still sent it back. A lesson needs designed motion and visuals that actually teach, not decoration on top of a face. We rebuilt it in the language of the pilot instead. In the pilot itself, fast, noisy transitions became long scenes, soft one-second fades, and presenter moves timed to the chapter changes.
Passed
Every technical check, and still sent back.
Designed
The presenter makes room, and the idea takes shape.
The same lesson, before and after. In the rebuild, every reveal on screen follows the narration. One picture, two voices
For the new program's welcome episode we produced two versions with two different voices, on exactly the same picture. The frames are identical down to the last one. Only the voice changes, so the choice can be made on the voice alone.
Picture
Voice A
Voice B
Two versions of the welcome episode. Same frames, a different voice, so the choice rests on the voice alone. The review room
I also built the place where the client reviews everything: every script, voiceover, and cut. They can comment on a single line of a script, a single moment in the audio, or a single frame of the film. They can leave a voice note, reply in a thread, and approve a whole module, which then locks. And after every release I test the live site myself. Once that caught a check that had quietly read “available” from an empty answer. It was live for minutes.
- A note on the presenter's voice
- A note on the caption
- A reply, in the same thread
Feedback, pinned to the frame it's about, with the reply threaded under it.
Every tool, chosen for its limits.
Most of this work is knowing what each model will and won't do before asking it for anything.
Eight seconds at a time.
The video model I use makes clips of about eight seconds. A presenter speaking for two or three minutes would take twelve to twenty-three separate clips, and the voice, the face, the gestures, and the lips would all have to match at every seam. So the background footage comes from that model, and the presenter comes from a tool made for one continuous take.
One presenter, in eight-second pieces
One continuous take
An illustration, not to scale. Background footage comes from the video model. The presenter comes from a tool made for one continuous take. Words go in code.
I never ask a video model for readable text, a logo, or an Arabic caption. It will try, and it will get letters wrong. Every word on screen is set in code instead, exactly. And writing 8K in a prompt is a wish, not a setting. I check the file itself.
Generated
Set in code
ق
سوق العمل
An illustration. The words on screen are never asked of the video model. They are set in code. The voice sets the timing.
The approved voice take comes first, and every picture is timed to it. Any sound the video model makes on its own is thrown away.
Approved voice
Picture, cut to the voice
An illustration. The cuts fall where the approved voice pauses.
Four phases, one room.
Every lesson passes through the same four phases. Each one asks the client about one thing, in the place where it lives: the words, the voice, the picture, then the whole film.
- 01
Script
Comment on a single passage, in Arabic or English.
- 02
Voiceover
Comment on a moment in the waveform.
- 03
Preview cut
Comment on an exact frame.
- 04
Final cut
Approve the module. It locks.
Approved modules lock.
The tools change every month. The job is still to make one person understand one thing.
The camera side of this lives in another room: Everything, A to Z