llms

Hotel Lobby AI Filter: 2-Photo Duo Video Tutorial

September 29, 2026 | Ryan Carter

The Hotel Lobby AI Filter is a viral two-photo image-to-video template that drops two people into a recognizable booth — orange set, overhead mic in the middle, alternating verses like the Quavo & Takeoff COLORS SHOW performance that started the trend. On viddo you can build it yourself with two photos, a short prompt, and the Reference To Video workflow — no studio, no camera crew, no dance rehearsal.

This tutorial walks through the full pipeline: how the trend works, the viddo models that handle it best, a copy-paste prompt template, and the legal guardrails you need before you publish.

1. What Is the Hotel Lobby AI Filter?

The Hotel Lobby AI Filter is an image-to-video format built around three recognizable ingredients:

· Two distinct subjects — one anchored left, one anchored right, never swapping positions

· A single microphone in the middle — overhead or boom, signaling the call-and-response pattern

· A warm orange / terracotta studio background — the COLORS SHOW booth aesthetic, not a real hotel lobby

The name comes from the Quavo & Takeoff “Hotel Lobby” COLORS SHOW performance, a 2022 live rendition that became a meme template. By 2026 it has spawned an entire AI remix culture — Quincy Baal’s AI Century Studios viral how-to, Ronald Isley’s smooth-tribute remix, and dozens of cross-platform recreations where friends, siblings, and even pets are dropped into the booth.

The format is image-to-video by design: you start with two stills and animate each into a verse pose, then stitch them with a microphone-driven cue. That is exactly what viddo’s Image to Video canvas was built for.

2. Why the Trend Took Off

Three forces converged to make the Hotel Lobby AI Filter one of the strongest 2026 remix templates:

1. A short, loopable visual hook. The orange booth, single mic, and two-position split-screen are visually distinctive at any thumbnail size — TikTok, Reels, and Shorts all reward recognizable templates.

2. An open-ended cast. Friends, siblings, parents, exes, fictional characters, even pets — anyone with two photos can participate. The format has no age, gender, or language gate.

3. AI image-to-video reached the threshold. Models like Seedance 2.5, Wan 3.0, and Veo 3.1 on viddo now hold identity across 5-10 second clips and ship native audio. The duo-format that looked gimmicky in 2024 now looks cinematic.

The trend lives in a sweet spot between meme culture and short-form storytelling — which is why it keeps pulling new participants into the format’s format.

3. How to Make a Hotel Lobby AI Video on viddo (3 Steps)

The full workflow lives inside the Image to Video canvas on viddo. Three steps, no editing skills required.

Step 1 — Pick Two Photos and Anchor Left/Right

Take one photo of each subject. The two images should:

· Show the face clearly (no sunglasses, no heavy shadows)

· Use similar lighting and resolution so the model treats them as a matched pair

· Leave the body partially visible — at least chest-up, ideally feet on the shoulder line

Before you generate, commit to who is on the left and who is on the right. Mixing positions between attempts is the single most common cause of identity swap.

Save the photos in the same folder and rename them left.jpg and right.jpg so the upload order stays consistent across iterations.

Step 2 — Choose Your viddo Model and Input Mode

Open the Image to Video canvas and pick the model that matches the look you want:

Model

Best for the Hotel Lobby Filter

Why

Seedance 2.5 (default)

Full 30-second verses

30s per clip + native audio, the closest match to a real COLORS performance

Wan 3.0

Cinematic polish

Smoother motion and color grading when you want a more produced feel

Veo 3.1

Audio-led versions

Native dialogue and ambient booth sound baked into one generation

Kling 3.0

Multiple takes with the same two people

Strongest character consistency across iterations

For a first attempt, Seedance 2.5 with native audio on is the fastest path to a publishable result. Lock the seed once you get a take you like — it makes re-runs consistent.

Step 3 — Generate, Verify, and Iterate

Hit Generate, then watch the full clip — not just the first frame. Check for:

· Identity drift on either subject

· Position swap (the same person appearing on both sides)

· Hand distortion near the microphone

· Mismatch in clothing, hairstyle, or background color

If something is wrong, change one prompt element at a time and re-generate. Identity swaps usually come from prompt ambiguity (“two friends talking” is too vague); face drift usually comes from a photo where the lighting is too dark. Fix the source photo first, then refine the prompt.

Once both verses land, drop them into viddo’s /trimmer for the final cut — trim to the beat, add captions, and export. The full pipeline stays in one workspace.

4. The viddo Models That Work Best for the Hotel Lobby Filter

Not every image-to-video model handles a two-subject booth equally well. Below is the breakdown based on what consistently produces a publishable Hotel Lobby take on viddo.

Seedance 2.5 — Default for 30-Second Verses

Seedance 2.5 generates up to 30 seconds per clip on every viddo plan, with native audio. For the Hotel Lobby format, that means you can capture a full verse per subject without stitching multiple clips. The native audio path renders ambient booth sound — the mic proximity, room tone, and cadence — without a separate sound-design pass. Best for the canonical two-photo duo video.

Wan 3.0 — When You Want a More Cinematic Booth

Wan 3.0 is the second viddo model that ships 30-second clips with native audio. The Wan family leans more cinematic than Seedance, with smoother color grading and a slightly slower motion curve. Use it when the Hotel Lobby take is meant to feel like a music video rather than a meme — slow dolly-in, warm tungsten palette, lens-flare highlights.

Veo 3.1 — When Dialogue and Booth Ambience Matter

Veo 3.1 is Google’s audio-native flagship on viddo. It renders dialogue, mic proximity, and the orange-room ambience together, which makes it the right pick when the Hotel Lobby take is meant to feel like a real studio recording rather than an AI mimic. Strong lip-sync, stronger audio cue alignment than the other models.

Kling 3.0 — For Multiple Takes with the Same Two People

Kling 3.0 on viddo has the strongest character-consistency signal across multiple generations. If you plan to publish a series — same two subjects, multiple verses across a week — Kling 3.0 keeps the faces stable across iterations. Less cinematic than Veo, but more reliable for serialized content.

Reference To Video — When You Want to Match the Original COLORS Booth

The Reference To Video mode on viddo lets you upload a real COLORS SHOW clip (or any reference video) and use it as the visual style guide. The model extracts the orange set, the booth lighting, and the mic placement, then applies them to your two photos. The strongest tool for fidelity to the original trend aesthetic.

5. Hotel Lobby AI Filter Ideas (6 Variations)

The format is open-ended. Below are six variations that work well with the viddo image-to-video pipeline.

Friends

Two friends who already have inside jokes or a recognizable dynamic. The mic-driven call-and-response format maps naturally onto friend banter — one reacts, the other delivers the punchline. Best with Seedance 2.5 for the natural rhythm.

Siblings

Siblings communicate in a way friends do not — shorter cues, more physical shorthand, shared history. The format reads as authentic almost by default. Wan 3.0 cinematic mode gives it a music-video polish if you want to elevate beyond the meme.

Couples

Couples can use the format for anniversary tributes, throwback verses, or playful callouts. Veo 3.1 with native dialogue is the right pick when the verses are supposed to feel like real conversation rather than AI mimicry.

Cross-Generational

Grandparent and grandchild, parent and teen, mentor and mentee. The format is most emotionally effective when the two subjects share a clear age gap. Kling 3.0’s character consistency matters here because the facial differences are wider.

Pets

Two pet photos in the same booth. The motion is subtler — head tilt, ear perk, blink — but the format still works. Use Reference To Video with a real COLORS performance as the style anchor and the model will infer the booth energy even from still animal faces.

Fictional Characters

Two made-up personas in a stylized aesthetic. Anime characters, video game protagonists, book characters — the format works because the orange booth is visually neutral enough to host any art style. Veo 3.1 with cinematic prompt language gives the strongest stylized result.

6. The Hotel Lobby AI Filter Prompt (Copy-Paste Template)

The prompts that produce clean Hotel Lobby takes all follow a five-element structure: subject + position + setting + camera + motion. Below is a template you can copy and adapt.

Create a vertical 9:16 image-to-video of a COLORS booth
performance. Subject A is on the left, Subject B is on the
right. Both perform at a single overhead microphone placed in
the middle of the frame. Both subjects must keep their
hairstyles, clothing, skin tone, and identity from the input
photos.

Setting: warm orange studio, soft natural light, terracotta
back wall, single overhead mic.
Camera: locked tripod, eye-level medium shot, no pan, no zoom.
Motion: Subject A delivers the first verse (8-10 seconds,
subtle hand gestures, head movement, mic proximity). Then
Subject B reacts and delivers the second verse (8-10 seconds,
mirror motion). Alternate naturally between the two subjects.
Audio: ambient booth room tone, mic proximity breath, no
music.

Avoid face blending, identity swap, extra people, background
movement, or camera shake.

Element breakdown:

· Subject + position: “Subject A on the left, Subject B on the right” — explicit left/right anchoring is the strongest defense against identity swap.

· Setting: “warm orange studio, terracotta back wall” — the COLORS booth visual signature; without it, the model defaults to a generic talking-head video.

· Camera: “locked tripod, eye-level medium shot” — any other camera move makes the format read as something other than a Hotel Lobby take.

· Motion: “Subject A delivers first verse, Subject B reacts and delivers second verse, alternate naturally” — the alternating pattern is the format’s defining feature; the model needs to be told.

· Audio: “ambient booth room tone, mic proximity breath, no music” — the absence of music forces the model to render the booth ambiance, which is what makes the format feel like a real performance rather than a music remix.

39 words is the sweet spot. Anything longer and the model starts losing coherence across the five elements.

7. Hotel Lobby AI Filter FAQ

What is the Hotel Lobby AI Filter?

The Hotel Lobby AI Filter is an image-to-video template that drops two subjects into a recognizable COLORS SHOW booth — orange set, overhead microphone in the middle, alternating verses. It originated from the Quavo & Takeoff “Hotel Lobby” COLORS SHOW performance and became a 2026 AI remix trend across TikTok, Reels, and Shorts.

Which viddo model works best for the Hotel Lobby Filter?

Seedance 2.5 is the default — 30-second clips, native audio, identity-stable across the full verse. Wan 3.0 for a more cinematic take, Veo 3.1 when dialogue and booth ambience matter, Kling 3.0 for serialized content with the same two subjects, and Reference To Video when matching the original COLORS booth aesthetic.

How long is a Hotel Lobby AI video?

Each verse runs 5-10 seconds on standard models, or up to 30 seconds on Seedance 2.5, Wan 3.0, or Grok on viddo. A full two-verse take is typically 10-20 seconds total, which fits the TikTok / Reels / Shorts sweet spot.

Can I use the original COLORS SHOW music in my Hotel Lobby video?

You can not legally re-upload the original Quavo & Takeoff “Hotel Lobby” audio to TikTok, Reels, or YouTube without a sync deal. Most Hotel AI takes work without music because the mic-driven cadence is the format’s signature — the absence of background music is part of the trend aesthetic. If you want to add music, use a royalty-free track or an AI-generated instrumental from viddo’s Music to Video tool.

Is it legal to use someone else’s photo for the Hotel Lobby Filter?

You need documented consent from anyone whose likeness you use, unless the subject is clearly fictional (anime, video game character, illustration). For public figures, you also need to comply with deepfake and synthetic media regulations in your jurisdiction (EU AI Act, US NO FAKES Act, China’s Generative AI Services Measures) and label the video as AI-generated content.

What if my Hotel Lobby video has identity swapping?

Identity swap (the same face appearing on both sides, or faces drifting between verses) is the most common Hotel Lobby failure. The two main causes are (1) the prompt does not explicitly anchor Subject A to the left and Subject B to the right, and (2) the two source photos have inconsistent lighting or resolution. Fix the prompt first (use the template in Section 6), then re-shoot the photos with matched lighting.

8. Try the Hotel Lobby AI Filter on viddo

Ready to make your first take? Open the viddo Image to Video canvas, pick Seedance 2.5 as your model, and drop in your two photos. The first generation takes 3-5 minutes; the iteration loop is where the format lives.

 

viddo.ai

Viddo AI 是一款先进的一体式人工智能视频和图像生成平台,可让您快速轻松地从各种输入创建令人惊叹的视频和图像。 这些人工智能模型由 Google、OpenAI、Grok AI、ByteDance、Alibaba、Kling、Runway、Vidu、Minimax、Elevenlabs、Midjourney 等提供支持。

© 2026 viddo.ai. All rights reserved.