2026 Edition · Our own software · local or cloud APIs

The agentic studio for AI video.

One idea in. A finished, on-brand vertical reel out — storyboard, footage, voiceover, music, presenter and publishing, on a pipeline you can run yourself or drive with an agent.

9:16Native vertical · 16:9 · 1:1
1Idea → finished reel
6+Local engines, one pipeline
~70Agent tools over MCP
01 · Making short-form video in 2026

Great videos are easy to imagine.
They're brutal to actually make.

01

Every reel is six tools.

A script writer, a footage source, a voiceover app, a music library, a video editor, a scheduler. Six tabs for one 20-second clip.

02

Stock footage looks like everyone else.

The same drone shots, the same gradients. Nothing on screen is yours, and nothing is on-brand.

03

Voice and music are a licensing maze.

A consistent narrator costs a subscription. Royalty-free music that does not sound royalty-free costs another.

04

On-brand means redoing it every time.

The logo, the voice, the company line — pasted back in by hand on every single export.

05

Cloud video bills by the second.

Per-clip pricing makes volume expensive, and your raw footage lives on someone else’s server.

06

Five platforms, five exports.

YouTube, TikTok, Instagram, Facebook, LinkedIn — five logins, five captions, five aspect ratios.

02 · One idea, nine stages, one reel

Type the idea. The studio does the rest.

Each stage is its own model and its own editable step — run the whole thing in one call, or stop and tune any stage before the next.

01

Scenario Wizard

Ollama · local LLM

Type a plain-language idea. The local model returns an ordered storyboard of 3–12 cinematic shots — framing, motion, mood, all editable.

02

Narration

Ollama

AI-writes a spoken line per scene, sized to each shot’s duration and matched to a tone you set — warm, punchy, authoritative.

03

Shots

ComfyUI

Each scene renders as text-to-video, or image-to-video when the shot has a reference frame. Re-roll any single shot without touching the rest.

04

Voice

XTTS

Per-scene voiceover or one continuous track. Use a built-in speaker, or your brand’s cloned voice from a short sample.

05

Music

ACE-Step · Suno

A scored soundtrack from style tags, BPM and length — generated locally with ACE-Step, or via Suno with your own API key.

06

Presenter

Wan 2.2 S2V

Optional: turn any scene into a lip-synced talking presenter from a single face image. Fast or high-quality pass.

07

Cover

SDXL

A branded intro still generated to open the reel — your logo and palette baked in.

08

Assemble

Remotion · ffmpeg

Shots, voiceover and soundtrack stitched into one finished 9:16 reel, captions and overlays in place.

09

Publish

Scheduler

Draft a caption, schedule a slot, and post to YouTube, TikTok, Instagram, Facebook and LinkedIn from one screen.

make_reel — the whole run in a single call.Scenario → narration → render every shot + music + speech → assemble → finished reel. One tool, hands-free.
03 · By the numbers

Built to make video at volume.

3–12Cinematic shots per storyboard, from one idea
6+Local engines — video, image, voice, music, LLM, lip-sync
5Publish targets from a single draft
~70MCP tools — drive the whole studio from any AI client
3Aspect ratios — 9:16, 16:9, 1:1, set per reel
0Per-clip cost when you run it on your own hardware
04 · Two ways to render

Render on your GPU, or on cloud APIs you bring.

It's one self-hosted studio — you decide where each stage runs. Own a GPU? Render locally for free. Prefer the cloud? Drop in your own provider keys. Mix them per capability.

ON YOUR GPU

Local render

Your GPU. Your models. Zero per-clip cost.

The whole pipeline runs on your own hardware. Open models do the work, nothing leaves your network, and the marginal cost of another reel is electricity.

  • ComfyUI for text-to-video & image-to-video
  • Ollama for the storyboard, narration & captions
  • XTTS voice synthesis + brand-voice cloning
  • ACE-Step music, Wan 2.2 S2V presenter, SDXL covers
  • Remotion + ffmpeg assembly, a worker queue you control
  • Full privacy — footage and scripts never leave your box
See the engines
YOUR API KEYS

Cloud APIs

No GPU? Bring your own provider keys.

No hardware to spare? Drop in your own cloud API keys and the studio routes that stage to a hosted provider instead — you pay the provider directly, nothing runs through us.

  • Your own keys — e.g. Kling for video, Suno for music
  • Latest hosted models, no drivers or downloads
  • Set a key per capability; leave the rest local
  • You pay the provider directly — no markup, no middleman
  • Same storyboard → reel → publish workflow
  • Mix and match: local where it is cheap, cloud where it is not
See the engine matrix
05 · The engine room

Open models locally. Your API keys in the cloud.

Every capability is a swappable slot. Run it on open weights you control, or point it at a cloud provider with your own key — per capability, your call.

CapabilityLocal · your GPUCloud API · your key
VideoComfyUI — text-to-video & image-to-videoKling (your API key)
Image · CoverSDXL via ComfyUIYour image API key
Script · LLMOllama (any pulled model)Your LLM API key
VoiceXTTS — synthesis + voice cloningYour TTS API key
MusicACE-StepSuno (your API key)
Presenter · Lip-syncWan 2.2 S2VYour lip-sync API key
AssemblyRemotion + ffmpegLocal (both routes)
06 · Brand kit

Every reel comes out on-brand.

Set it once per project. Logo, voice, presenter and company details merge into every render — no copy-paste on export.

Logo

Dropped onto covers and overlays so every reel is unmistakably yours.

Cloned voice

A short sample becomes a reusable narrator — the same voice across every video.

Presenter face

One face image powers lip-synced talking-head scenes on demand.

Company details

Name, phone, website and email merge into captions and end cards automatically.

Brand voice card

A written tone description steers the storyboard and narration writers.

Tag catalogs

Editable libraries of cinematic scene concepts and music styles, reused across projects.

07 · You stay in the director's chair

Generated, but not on autopilot.

The wizard gives you a first cut in one pass. From there every shot, line and beat is yours to nudge — and only what you change re-renders.

  • Edit any shot · rewrite the line, swap the reference image, change motion — then re-render just that scene
  • Reorder & trim · drag shots, adjust durations; narration resizes to fit
  • Rewrite voiceover · regenerate a single shot’s spoken line without touching the video
  • Overlays · text, lower-thirds and brand graphics composited at assembly
  • Aspect per reel · 9:16 for shorts, 16:9 for YouTube, 1:1 for feeds
  • Jobs you can watch · every render is a queued job — poll, wait, or cancel from the dashboard or an agent
Storyboard · “Atelier launch”9:16 · 6 SHOTS
01Macro: thread pulled through linenRENDERED
02Hands shaping clay on the wheelRENDERED
03Presenter: “Made by people, not machines.”PRESENTER
04Product turn on a warm-lit pedestalRENDERING
05Logo cover · champagne on charcoalQUEUED
Music · ACE-Step · cinematic, warm · 92 BPM
08 · Agentic by design

Drive the whole studio from your AI client.

LuwiStudios ships a Model Context Protocol server. Claude Desktop, Cursor, n8n or your own agent can run the entire pipeline — in plain language.

MCP ClientCLAUDE · CURSOR · N8N
LuwiStudios MCP~70 TOOLS
Render workersCOMFYUI · OLLAMA · XTTS
create_projectcreate_scenariowrite_narrationgenerate_shotgenerate_speechgenerate_musicgenerate_avatarassemble_reelmake_reelcreate_post+ ~60 more
09 · Why now

The economics of video just flipped.

Format

Short-form won.

Vertical 9:16 is the default unit of attention across every major platform. The brands that post daily are the brands that get seen.

Models

Open weights caught up.

Wan, SDXL, XTTS and ACE-Step now produce broadcast-credible output on a single GPU box — quality that used to mean a cloud bill.

Cost

Per-second pricing breaks at scale.

Cloud video is fine for a handful of clips. At volume, self-hosting turns marginal cost into electricity — and keeps your footage in-house.

Frequently asked questions

Answers to the most common questions about LuwiStudios.

Our own agentic studio for AI video — end-to-end software we build in-house. You give it an idea; it writes a storyboard, renders each shot, generates voiceover and music, optionally adds a lip-synced presenter, and assembles a finished vertical reel ready to publish. Every stage renders on your own GPU, or through cloud APIs you bring your own keys for — your choice, per capability.

You already have the ideas.

LuwiStudios turns them into finished video — on your hardware or ours, by hand or by agent. Tell us what you want to make.

Our own software · ComfyUI · Ollama · XTTS · ACE-Step · Wan 2.2 · SDXL — or bring your own cloud API keys (Kling · Suno · …)