Viral AI Trends
By Viral AI Trends · Published and updated

Hotel Lobby AI: how to make the viral orange-booth video

Two photos go in. A short vertical clip comes out — two performers, one hanging microphone, a saturated orange booth. This page explains the format, walks through the whole workflow, and points you at the routes that will not cost you anything to try.

What the Hotel Lobby AI trend actually is

In June 2022, Quavo and Takeoff — recording as Unc & Phew — performed the song "Hotel Lobby" on A COLORS SHOW. The staging is hard to mistake: a seamless orange room, a single vintage microphone hanging from the ceiling between two performers, one locked-off camera, no crowd and no props.

In mid-September 2026, creators started feeding AI video tools two ordinary photos and getting that same booth back with anyone inside it — friends, couples, pets, cartoon characters. Within days the format was everywhere on TikTok and X, and it kept compounding from there. It also circulates as the orange booth video or the COLORS booth trend.

Why it spreads so fast. The staging announces the joke before anyone reads a caption, and the payoff comes from who you pair up rather than from editing skill. The format rewards imagination, not experience — which is also why it travelled so far beyond the audience of the original song.

The five rules that make a clip read as this trend

Swap the cast, keep the structure. These five things are what make a viewer recognise the format instantly:

Everything else — wardrobe, motion, the beat you add before posting — is yours.

How to make one, step by step

You do not need a green screen, a set, or any editing experience. The whole workflow is: two photos, one format decision, one render.

  1. Choose a duo people can read at a glance.

    The joke lands on contrast, not craft. Your two cats, a grandparent and a grandchild, two characters from the same show — anything with an obvious relationship. If you want the least risky result, use pets, drawn characters or fictional performers rather than real people.

  2. Prepare two photos — one subject each.
    • Front-facing, evenly lit, neutral expression
    • No sunglasses, hands or hair away from the face
    • One person (or animal) per photo — a group shot makes the two faces blend
    • Roughly 300–6000 pixels per side; avoid screenshots and distant shots
    • Skip heavy beauty filters — they strip the detail the model needs to keep a face recognisable
    • If you upload several photos of the same subject, vary the angle slightly
  3. Decide who stands left before you upload.

    Almost every tool maps the first photo to the left performer and the second to the right. This is the step people get stuck on — if the faces come back swapped, you re-render. Choosing in advance saves a render.

  4. Pick a route: template, or prompt.

    Template tools hand you the booth, the microphone, the framing and the motion — you upload and go. Prompt-based tools let you describe wardrobe, lighting and camera moves instead, which is more control but more work. Ready-made prompts are further down this page.

  5. Set the format.

    Vertical 9:16 for TikTok, Reels and Shorts. 10–15 seconds is the sweet spot — it matches the loop length of the original and keeps both subjects in frame. 720p is plenty for social.

  6. Render short first, then check five things.

    Watch the whole clip before you post. Check framing (both subjects fully in shot), whether anyone drifted out of the booth, the audio, whether the left/right order survived, and the first moment after each cut. If a face drifts, swap in sharper photos and render again.

  7. Post it properly.

    Add the trending sound from the platform's own sound library, switch on the AI-generated label, and keep the caption short — the format needs no explanation.

The five mistakes that ruin a render

MistakeWhat happens
Group photo instead of one subject per imageThe model blends the two faces and the booth ends up with the wrong people
Heavy beauty filtersFaces lose the detail that keeps them recognisable once they move
Left and right never specifiedPerformers come back swapped — the most common re-render
Posting the first render unwatchedDrifting faces and cut glitches go public
Wrong aspect ratioA 16:9 clip in a vertical feed gets cropped and the booth stops reading

Hotel Lobby AI prompts you can paste

These describe scene, wardrobe, lighting and camera only — the pattern that keeps renders reusable and safe. They never name a real person, and they never reference the original footage.

How to use these. Paste the prompt into a text-to-video or image-to-video tool that supports custom prompts. Where a prompt says "two performers", upload your two photos and set which one stands left. If your tool has a built-in Hotel Lobby template, use that first — a template is faster, a prompt is more controllable.

1. The classic booth

Two performers standing side by side in a seamless saturated orange room, one vintage microphone hanging from the ceiling between them, flat even lighting, locked-off full-body wide shot, no camera movement, no props, vertical 9:16, 12 seconds.

Use when: you want the recognisable look with zero styling decisions.

2. Rap-group styling

A fictional rap duo in matching cream tracksuits and layered gold chains, standing shoulder to shoulder in an orange studio booth, single suspended microphone centred between them, warm key light from above, locked-off full-body shot, shallow depth of field, vertical 9:16, 12 seconds. Performers are fictional — no real-person likeness.

Use when: you want the music-video wardrobe without naming any artist.

3. Pet duo

Two pets, one on each side of an orange studio booth, a single microphone hanging between them, straight-on eye level camera, locked-off full-body shot, soft even lighting, playful alternating head turns, vertical 9:16, 10 seconds.

Use when: the funniest version. Works best with two clearly different silhouettes.

4. Cinematic version

Two performers in an orange studio, deep saturated orange backdrop and floor, one suspended microphone between them, low wide-angle tripod shot with a very slow push-in, warm rim light, subtle haze in the air, cinematic music-video colour grade, vertical 9:16, 15 seconds.

Use when: you want it to look shot rather than generated.

5. Cartoon or animated characters

Two original cartoon characters standing in a flat orange studio booth, one hanging microphone between them, bold clean outlines, flat cel shading, static full-body camera, exaggerated alternating gestures, vertical 9:16, 10 seconds. Original characters only.

Use when: you want to avoid likeness questions entirely.

6. Grandparent and grandchild

An older and a younger performer standing in a seamless orange room, single microphone suspended between them, warm soft lighting, locked-off full-body shot, gentle alternating gestures, no props, vertical 9:16, 12 seconds.

Use when: the contrast is generational rather than stylistic.

7. Silhouette / faceless

Two faceless performers in silhouette against a bright orange studio wall, one microphone hanging between them, backlit with strong rim light, locked-off full-body shot, minimal detail, moody, vertical 9:16, 12 seconds.

Use when: you want the format with no identity attached at all.

8. Retro film version

Two performers in an orange studio booth with a single suspended microphone, shot on 16mm film look, visible grain, slight halation, soft warm light, static full-body camera, vertical 9:16, 12 seconds.

Use when: the trend has been running a while and you want a distinct look.

Adapting a prompt that isn't working

Is there a free way to do this?

Yes, with limits. A couple of browser-based generators produce a clip without an account, and most credit-based tools include free sign-up credits so you can try before paying. Free tiers cap something — usually resolution, clip length, a watermark, or a daily limit.

Paid routes in late September 2026 were reported at roughly $5 for a 15-second clip on one-click template tools, up to a few thousand credits for higher-fidelity motion transfer at low resolution. The important distinction is whether a tool reuses the original performance or track — that changes both the result and your rights position. The full breakdown is on the generator comparison page.

Frequently asked questions

Is Hotel Lobby AI free?

Partly. A few tools render a clip in the browser without an account or a paid plan, and most credit-based tools hand out free sign-up credits so you can try before paying. Free tiers cap something — resolution, clip length, a watermark, or a daily limit.

What's a normal price if I do pay?

Reported figures in late September 2026 ranged from roughly $5 for a 15-second clip on one-click template tools, up to thousands of credits for higher-fidelity motion transfer at low resolution. Check what a credit buys at your chosen resolution before you commit.

Do I need a subscription?

Not on the tools that sell credits — you buy a pack and spend it. Some all-in-one studios also offer monthly plans, but for a single clip a credit pack is usually enough.

How many photos do I need?

Two, one subject each. Some tools accept up to four photos per person for a stronger likeness, but one clear front-facing shot per performer is normally enough.

Why is it so picky about the photos?

Because the model has to keep a face recognisable while it moves. Front-facing, evenly lit, no sunglasses, no group shots, and no heavy beauty filters — those remove exactly the detail the model relies on.

What if the performers come back swapped?

Set the order explicitly if the tool offers it. Otherwise re-upload: the first photo generally becomes the left performer and the second becomes the right. This is the single most common reason for a re-render.

Can I use a photo of a friend or a public figure?

A friend: only with their permission, and ideally with them knowing where it will be posted. A public figure or a performer: not without consent — several tools block prompts that name real people, and platforms treat unconsented likeness as a policy violation.

What about the original song and footage?

Some tools reuse the original performance or track; others build a fresh booth and generate their own audio. If you plan to monetise the clip, prefer the route that reuses nothing, and add the sound from the platform's own library when you post.

Do I have to label it as AI?

TikTok and Instagram both provide an AI-generated label and expect it on synthetic media. Switch it on — it costs nothing and keeps the account clean.

Can I put kids in the video?

Don't. Most tools prohibit content involving minors, and platforms remove it. Use adults, pets or original characters.

Is this trend already over?

The staging will date, the capability will not. The booth trend spread from mid-September 2026 and kept compounding for weeks. Two-photo character swapping in a fixed scene is a format you will be able to use long after this particular room looks old.

Disclosure. This site is an independent guide. It is not affiliated with, endorsed by or sponsored by Quavo, Takeoff, COLORS, Migos, or any generator listed. Some links may be affiliate links. Prices and features are as reported in late September 2026 — verify on the official page before purchase.