Back to projects

Repurpose video with less manual editing

AIClipper.video

Turn long YouTube videos into captioned clips with AI-assisted moment selection, visual templates, and face-aware framing. Visit https://aiclipper.video/.

Outcome

A complete preparation and rendering workflow produces downloadable MP4 clips and SRT captions, with repeatable channel imports.

My role

Full Stack Engineer

How I helped

  • Built the workflow from YouTube source to downloadable clips and captions.
  • Connected Gemini transcription and clip scoring to background rendering.
  • Added visual face tracking, podcast layouts, and channel imports.

Product

The workflow in context

Live homepage screenshot captured directly from https://aiclipper.video/.

Screenshot of AIClipper's live homepage showing its video clipping hero and navigation.
Public website screenshot

The solution

Creators submit a video or connect a channel, choose a template and editorial goal, then let background workers transcribe, score, caption, and render candidate moments.

What changed

The current Rails MVP connects video download, Gemini-assisted transcription, bounded clip selection, FFmpeg rendering, and downloadable MP4/SRT output. Seven visual formats include Face Track and Speaker Split. Editorial scores describe content potential, not predicted views.

Behind the solution

Rails 8.1 handles accounts and the creator workspace. Solid Queue runs the processing pipeline; Active Storage holds source media and finished clips. Gemini calls are isolated behind a dedicated client. FFmpeg and yt-dlp handle media processing.

Technical decisions

  • Validate timestamps and model candidates before rendering.
  • Bound source duration and size, scope media to its owner, and limit concurrent jobs.
  • Run face detection locally on reduced-size frames; use FFmpeg to crop the full-resolution video.
  • Deduplicate channel imports by YouTube ID.
  • Keep a real FFmpeg demo flow available without an AI provider.

Key features

  • Goal-weighted clip scoring and seven rendering formats.
  • Burned captions, branding, MP4 downloads, and SRT exports.
  • Face tracking and two-face podcast layouts with one-person fallback.
  • Hourly channel polling with duplicate protection.

Lessons learned

Validate model output before it reaches media tools, and keep each processing stage observable and recoverable. Visual face tracking is not audio speaker identification; full-frame layouts remain useful for difficult footage.

AI system implementation

Gemini handles transcription and candidate clip selection through a dedicated client. Solid Queue coordinates slow processing outside HTTP requests. Model timestamps and candidate boundaries are validated before FFmpeg renders MP4 clips and SRT captions. Local face detection guides crops; it does not identify speakers from audio. Editorial scores express content potential, not predicted views.

Could a similar approach help your business?

Tell me where work gets stuck. We can identify a practical first improvement and what success should look like.

Discuss your challenge