Skip to main content
World models

HappyOyster User Guide

HappyOyster world model user guide: capabilities, showcases, and prompt-writing for the Adventure, Directing, and Acting modes.

Overview

HappyOyster is an open-world model series — give it text or an image and it generates an interactive digital world in real time, covering three modes: Adventure, Directing, and Acting. Visit the HappyOyster console to try it online:
ModeWhat it doesInputOutput
Adventure (World Exploration)Enter a generated world in first or third person; move, fight, and ride freelyText + first-frame imageReal-time interactive video stream, 1+ minute continuous interaction
Directing (Real-time Directing)Stream a 3-minute video; inject commands at any time to change the storyText / first-frame image / multi-image referenceStreaming video with pause, rewind, and branching
Acting (Character Roleplay)Talk face-to-face with a generated character who responds with expressions, actions, and voiceText + first-frame imageReal-time interactive video stream

Adventure (World Exploration)

From text + first-frame image input, it generates an open-world scene of any style with multiple forms of interaction. Supports first-person and third-person views with 1+ minute of real-time movement and camera control. An image plus a line of description is all you need to enter a world you can move through, fight in, and interact with — ride horses, fly, drive, and explore worlds of different styles. Define what you want to play at creation time, and the world responds in kind.

Showcases

Perspective comparison

Third personFirst person
See your character's full body — suited for action-adventure and role-play.Immersive POV exploration — suited for driving and stealth.

Scene-aware interaction

The model understands what's in the frame and opens up matching gameplay based on the objects in the scene.
Subject motionEnvironmental interactionPhysical response
Sprint, jump, crouch, attack (varies by character type)Enter vehicles, trigger scene mechanismsCollision, bounce, gravity and other physics feedback

World styles

Supports world generation in any style. Here are some examples:
Felt · Island GameSpace · MoonwalkFantasy · Flight
Cartoon-material rendering, non-realistic explorationLow-gravity physics, space-scene interactionFlight, flying broomstick and other fantasy interactions
Animal ridingCombat & attacksVehicles
Mount, dismount, riding controlPunches, bows, firearms, magicCars, bicycles, skateboards

Interactions

CapabilityDescription
Subject motionSprint, jump, crouch, attack (varies by character type)
CombatPunches, bows, firearms, magic; solo/duo fighting
Mounts & vehiclesMount/dismount horses; cars, bicycles, skateboards
Environmental interactionEnter vehicles, physics collision response
World stylesRealistic, anime, felt, space, fantasy — any style
PerspectiveFirst person (immersive POV) / third person (action-adventure)

Prompt writing

Two core principles: first, whatever is in the scene is what you get to play with; second, whatever reaction you want, write the reaction out too. Objects you don't write can't be played with, and reactions you don't write won't be performed. Write your creation description along these 5 points and the interaction will match your expectations better:
  1. Perspective: First decide first-person or third-person. First-person suits immersive driving and POV exploration; third-person lets you see your character, suited for action-adventure and role-play.
  2. Who you play: Spell out the main character's gender / identity / appearance (e.g., "a female warrior in leather armor"). This matters especially for third-person, since how the character looks directly determines the motion style.
  3. Weapons / equipment: If you want to fight or attack, state what's in hand: fists / longbow / firearm / magic / sword. Different items mean different available actions.
  4. Opponent and reaction: Beyond stating who you hit, write out "what happens after a hit": the opponent stumbles back / drops on impact / raises a shield to block, plus the impact effects you want (sparks, spurting flames, a shockwave). The model performs the reaction you write — the more specific the reaction and effects, the stronger the sense of impact.
  5. Interactable objects: If you want to ride or drive, make the horse, car, bicycle, skateboard, etc. appear in the frame. If an object isn't in the frame, you can't get on it.

Prompt examples

Third person · Action-adventure
Third-person view: a female warrior in leather armor stands at the center of
an arena, holding a longbow, as a giant wolf lunges at her; sparks burst when
the arrow lands, the wolf flinches and stumbles back in pain, dust kicks up
from the sand, and the stands are packed with spectators.
Available interactions: draw and shoot, dodge, fight up close — hits produce sparks and the wolf's hit reactions. First person · Immersive driving
First-person view: I sit in the driver's seat of a vintage classic car, both
hands on the wheel, a dusk coastal highway ahead; the engine roars as I hit the
gas, and the wheels kick up dust from the road.
Available interactions: accelerate, steer, drive along the highway — with engine sound and dust feedback when accelerating.

Reference example

First framePromptGenerated video
Post-apocalyptic survival
Realistic cinematic game visuals. Third-person over-the-shoulder view, a man in a worn dark leather jacket stands on a dirt road in an abandoned village, shotgun raised in a combat stance. In the distant fog, several twisted humanoid creatures are slowly approaching. Surrounding him are collapsed wooden houses, a ruined windmill, and stacked firewood. The sky is overcast and gloomy, the overall atmosphere oppressive and terrifying, tones mainly grey-brown and dark green.

Directing (Real-time Directing)

From text, a first-frame image, or multiple reference images, it streams a video up to 3 minutes long. Inject commands at any time during generation to change the story's direction; supports pause, rewind, and branching. Reference uploaded images to lock character/prop appearance for full consistency.

Showcases

Storytelling

Single-character sceneTwo-character scene
Single-character monologue or action scene.Two-character dialogue, realistic cinematic texture.

POV immersive interaction

POV outdoor explorationRealistic character interaction
First-person roaming through outdoor environments.Face-to-face real-time interaction with realistic characters.

Core capabilities

CapabilityDescription
Streaming generationReal-time generation and playback, no need to wait for full render
Command injectionInsert new commands at any time during generation to change character actions, camera work, or scene environment
Pause & rewindPause then continue, or rewind to any point and re-perform
Multimodal referenceReference uploaded images to lock character/prop appearance for 3-minute consistency

Command injection types

TypeExampleNotes
Action trigger"She suddenly stands up." / "He puts the cup down."Short + clear action, most reliable
Camera work"Cut to a facial close-up." / "Slowly pull the camera back."Cinematic terms are more precise
Scene change"It starts to rain." / "Time jumps to dusk."Inject an environment variable to shift mood
New character"The door is pushed open, a man in a black hat walks in."Spell out appearance + first action

Prompt writing

A 3-minute video needs a structured screenplay. A six-part prompt structure is recommended:
PartNamePurpose
1WorldTime-space + genre + core conflict, 1-2 sentences to set the narrative coordinate
2CharacterGender/age/appearance/clothing/personality/emotion/voice — 7 essentials + reference images to lock appearance
3StyleDirector / color grade / lighting — three anchors
4SceneScene type / ambient mood / layout / ambient sound — four dimensions
5Full scriptSettings lock (viewpoint/subject/voice/pace LOCK) + shot plan (Open/Build/Turn/Sustain/Close — 5 acts)
6Negative listImage/sound elements that must not appear

Section-by-section

① World · Time-space + genre + core conflict
Sentence skeleton: [Year + season], [city / specific location], the genre is [narrative style], the core conflict is [one-line conflict].

② Character · 7 essentials + consistency
1 · Basics — gender / age / build / temperament
2 · Appearance — hairstyle & color / facial features / key identifiers (scar / mole / glasses)
3 · Clothing — top / bottom / shoes / accessories; the more specific, the more stable the consistency
4 · Identity — occupation / class / social role
5 · Personality — 1-2 keywords (reserved · hot-tempered · outwardly tough)
6 · Current emotion — what they're thinking, bothered by, or waiting for in this moment
7 · Voice / timbre — gendered voice + age range + accent + overall feel; a must for multi-character scenes

③ Style · Three anchors: director / color grade / lighting
1 · Director / genre anchor — "Japanese family-drama texture," "Hong Kong neon light and shadow," "hard-edged cold-tone sci-fi"
2 · Color grade — "warm gold tone, low saturation," "cool blue base + neon highlights"
3 · Lighting type — "natural light dominant," "hard top light," "side back-light from a window"

④ Scene · 4 dimensions + key-item reference
1 · Scene type (indie indoor café / abandoned suburban factory / ancient inn hall / school rooftop)
2 · Ambient mood (warm and relaxed / crowded and noisy / cold and oppressive / quiet and suspenseful)
3 · Layout (table & chair placement / key prop positions / spatial depth / main furnishings)
4 · Ambient sound — 2-3 specific sounds (indoor BGM + distant nature sound + occasional voices)

⑤ Full script · Settings lock + shot plan
a) Settings lock · 6 recommended LOCKs
  VIEWPOINT LOCK  Use a single viewpoint throughout; no switching
  POV SELF-LOCK   In POV, "my" face / full body / mirror reflection never appears
  CAST LIMIT      ≤ 3 subjects in frame at once
  SUBJECT LOCK    Character clothing / hairstyle / props unchanged for 3 minutes
  VOICE LOCK      In multi-character scenes, each voice stays fixed; no swapping
  PACE LOCK       No single silence / pause exceeds 2 seconds

b) Shot plan · 5 shots × 4 frames
  Open   [0:00-0:10] HOOK — strong opening, no flat exposition
  Build  [0:10-0:50] Establish character relationships; introduce the first variable
  Turn   [0:50-1:40] The film's one and only surprise (reversal / conflict / choice)
  Sustain[1:40-2:30] Emotional breathing room; density can thin but never fully silent
  Close  [2:30-3:00] End on one image / one line / one action; throw an interaction hook

  Each shot specifies:
  [Context] What's happening + where the lead is and what they're doing + visual anchor
  [Key lines] Format: CharacterName(voice cue) "line"; in multi-character scenes, mark the speaker
  [Pacing] Verifiable hard metrics, e.g., "at least 2 lines in 10 seconds"

⑥ Negative list · What not to show
Only state special taboos that can't be derived from the settings lock; don't repeat positive settings.
Image taboos: e.g., "no night scenes / rain / cars"
Sound taboos: no post-production voiceover / narration / BGM drowning out voices

Acting (Character Roleplay)

From text + first-frame image input, it generates a face-to-face virtual character. When you speak, the character responds in real time with expressions, actions, and voice, while appearance, voice, and personality stay consistent for 3 minutes. Upload an image to bring the character in your mind to life and start a real-time conversation.

Showcases

Acting's core interaction is "you and the character, face to face" — you speak, and the character responds in real time with expressions, actions, and lines. Throughout the session, the character is always the same person.
Realistic · Face-to-face2D stylized
Realistic character; expressions and tone shift naturally with the conversation; appearance locked throughout.2D stylized anime/illustration-style character — also supports face-to-face real-time dialogue with natural expression and tone changes.

Core interaction

CapabilityEffect
Real-time dialogueYou say a line, the character instantly responds with expressions, actions, and lines — like a real conversation
Face-to-face POVThe frame always faces the character's face; only your hand/forearm enters frame, keeping attention on the character
Appearance lockUpload an image to lock the character's look and key props; no face or voice changes throughout
Emotional continuityThe character's expressions and tone shift naturally as the conversation progresses — not a fixed template

Core capabilities

CapabilityDescription
Real-time dialogueYou say a line, the character responds instantly with expressions, actions, and lines
Face-to-face POVDefault first-person view; the camera always faces the character's face; the user never appears on screen
Long-range consistencyThe character's appearance, clothing, voice, and personality stay unchanged for 3 minutes; upload an image to precisely lock appearance

Prompt writing

Acting prompts focus on the character itself and your relationship with the character. Structure along these 5 dimensions:
DimensionDescriptionExample
PersonaGender/age/appearance/clothing/identity/personality"A young woman in her early twenties, long straight black hair, wearing a cream knit sweater, gentle and smiley"
VoiceGendered voice + age range + accent + temperament"Young female voice, bright, with a slightly coy nasal tone"
Relationship & settingYour relationship with the character + current scene"Lovers, a café by the window at dusk, warm yellow light"
Viewpoint lockPOV + "my" hand features"Throughout, only she and my right hand enter frame (adult male, well-defined knuckles)"
OpeningThe character's first action and first line"She rests her chin on her hand, looks at me, and says with a smile: 'You're here so early today'"

Prompt example

First-person face-to-face view. Opposite me is a young woman in her early
twenties, long straight black hair, wearing a cream knit sweater, gentle and
smiley; young female voice, bright, with a slightly coy nasal tone.
We are lovers, right now in a café by the window at dusk, warm yellow light.
She rests her chin on her hand, looks at me, and says with a smile:
"You're here so early today."
Throughout, only she and my right hand enter frame (adult male, well-defined
knuckles, no accessories); my face and reflection never appear.
Interaction: talk to her in real time, joke around, hand her a cup — she responds with expressions, actions, and tone.

Next steps

Once you know the three modes, see HappyOyster Overview to integrate via API / SDK.