Brief
I need a program in Python that runs on my PC (Windows) and creates faceless videos in the format of "voiceover + changing visuals from photos and clips" — like historical documentaries on YouTube (I will attach an example separately).
I input a topic → the program writes a script, voices it, selects video/photos from free archives for each piece of text, stitches them together → outputs a ready MP4.
For me only. No website, no users, no sales. One user — me.
How it works (step by step)
1. Input. A simple window opens. I enter the topic of the video and choose a voice from the list (**the list of voices is automatically pulled from ElevenLabs via API** — available on my account). I click "Create." (Second mode: insert a ready script instead of generating.)
2. Script. The program writes a script on the topic of the specified length via LLM API (OpenAI/Anthropic, key in settings).
3. Scene breakdown. LLM divides the script into scenes and returns for each:
- scene text;
- visual type: video or photo;
- search query (a detailed phrase of what should be in the frame);
- highlight mark + importance rating 1-10 (for intro, see below).
4. Voiceover. The text is sent to ElevenLabs API with the chosen voice → audio. The video sequence is cut to match the length of the audio for each scene.
5. Video/photo selection — from free sources via API (see the list below). A specific query is formed for each source (for accuracy). If one doesn't yield results, it tries the next.
6. Verification (maximum 3 steps per scene).
- Step 1: the program takes the first found option (video/photo) based on the query.
- Step 2: sends the frame to LLM — "does it fit the scene?". If it fits → stop.
- Step 3 (if it doesn't fit): the program switches to searching for photos (finding photos based on an exact query is easier than videos — this guarantees relevance) and takes it with enhanced motion (zoom + pan).
Do not perform more than 3 steps for one scene — this saves the LLM budget and ensures that the frame is relevant.
7. Assembly via ffmpeg/moviepy: clips and photos timed to the voiceover, photos are animated with zoom (Ken Burns effect), voice on top, simple transitions. Output: MP4 1920×1080.
Video sequence rules (important — the program is responsible for this)
- First 60 seconds — intro teaser: a montage of the most impactful clips from the entire video
(scenes with the highest importance rating that contain video) under a separate introductory text from LLM ("in this video you will learn..."), frames without explanations, creating intrigue. Then a transition to the main part.
- Alternation: video insert at least every ~6 seconds, no many photos in a row.
- Video share: at least ~40% of the time — live clips, the rest — photos with zoom.
- Frame length: 4-6 seconds (both photos and videos). Do not flicker, do not hold static for too long.
- For purely historical topics where there is no video — photos with enhanced motion (zoom + pan).
Sources (all free, with API)
Modern video + photo: Pexels, Pixabay. Historical / archival (public domain): Wikimedia Commons, Archive.org, Library of Congress, Europeana, NASA, Smithsonian Open Access, Flickr Commons, openverse.
Each source is a separate module, easy to add new ones. Use only public domain / free licenses with the right for commercial use. No parsing of other YouTube/sites, film clips, images "from Google".
Uniqueness of selection
To ensure videos do not match others: take a random clip from the top results (not the first), maintain a database of already used clips (do not repeat), optionally — light processing of the clip (crop/mirror/speed).
Options (enable/disable in settings)
- Subtitles (embed or separate .srt).
- Background music.
- Clip processing for uniqueness.
- Resolution/format, video length, video share, search depth.
Technical requirements
- Python. Modular structure (sources and LLM — through interchangeable modules, to easily
replace or add).
- All API keys — in the settings file, not in the code.
- Simple window (GUI at the discretion of the performer — Tkinter/PyQt), launched with a double click.
- README with instructions, clear logs, comments in the code.
What I provide
- API keys (ElevenLabs, LLM, where registration is needed — I will arrange). I will pay for any fees myself.
- Examples of video references (I will attach) and examples of topics for tests.
Acceptance (ready if)
- I launch → window → I enter the topic, choose a voice → "Create" → I receive a ready MP4.
- The video sequence matches the meaning of the text, alternation of video/photos, intro teaser 60 sec, voiceover on top.
- Works with at least 6 free sources, with fallback between them.
- Frame verification through LLM: max. 3 steps per scene (found → LLM checked → if not,
photo with motion as a sure bet).
- Uniqueness: randomization + database of used clips.
- Only legal sources. There is a README, runs from scratch.
Transfer of results
- All source code — in open form (all files), without obfuscation + a compiled working version.
- I can run it myself from the source according to the instructions (README: installation, keys, launch).
- The code must be clean, commented, and understandable, so that **any other programmer
can continue working** on it if needed (no tie to the author).
- All rights to the code after payment — mine.
Disk space management (important)
The program should not fill up the disk. Implement:
- After assembling the video, all intermediate files (downloaded clips, temporary pieces,
audio cuts) are automatically deleted — only the ready MP4 remains on the disk.
- Cache limit (parameter in settings, e.g., 5 GB): when exceeded, old downloaded
files are automatically deleted (starting with the oldest).
- I set the folder for ready videos and for temporary files in the settings.
- Show how much space is occupied, and a button to "clear cache" manually.
Please specify in your response
- Examples of similar works (ffmpeg/moviepy, working with stock/archive APIs, ElevenLabs/LLM).
- Proposal for GUI.
- Does the solution use a database (which one and why) — or are local files sufficient.
- Deadline and cost.