Skip to main content
Guide contents

Guide contents

Time to study: 18 min
#higgsfield#cinema-studio#soul-2#video-generation#ai-video#ugc#api
NEWIntermediate18 min

Complete Guide to Higgsfield Cinema Studio 4.0: Scene Management, Soul 2.0, and API Automation

In-depth practical breakdown of the major Higgsfield update: Cinema Studio 4.0, consistent characters with Soul 2.0, API-based generation, and a ready-to-use UGC pipeline.

Published:
FOR AI AGENTChatGPTClaude

The Higgsfield ecosystem has received a major technological update: Cinema Studio 4.0 offers an expanded set of tools for direct scene control, the Soul 2.0 module ensures stable character consistency across frames, and the API catalog provides access to over 50 neural network models for end-to-end video production automation.

This practical guide details all key platform innovations and real-world implementation scenarios—from fine-tuning optics and character emotions manually to programmatic video pipelines and autonomous UGC creative generation.

Tool / ModulePrimary PurposeOptimal Use Case
Cinema Studio 4.0Precise frame-by-frame directorial controlCreating authorial shots with control over camera, lighting, tempo, and optics
Soul 2.0 (Soul ID)Character identity fixationSeries video production featuring a consistent hero in different locations
Higgsfield APIProgrammatic batch generationIntegration into B2B products, automated funnels, mass rendering
UGC Pipeline (Claude + Seedance)Ad generation from product photosRapid launch of marketing creatives without a film crew

1. Key Updates to the Higgsfield Platform

The central event of the release was the launch of Cinema Studio 4.0. The main conceptual shift is moving away from attempts to fit all directorial instructions into one long text prompt toward physically accurate pre-visualization tools before rendering begins.

Key technical improvements include:

  • Increased Duration: The maximum length of a single generated clip has increased from 15 to 30 seconds;
  • Batch Reference Upload: Simultaneous submission of up to 50 images instead of the previous 9;
  • Camera Movement Library: Over 30 curated cinematic presets replacing the 9 basic ones;
  • Camera Types: 4 types of shooting systems instead of 3;
  • Color Styles: More than 50 professional color grading presets instead of 8;
  • Dedicated Controls: Added independent regulators for Tempo, Era Selector, and Emotion Wheel;
  • Bidirectional Scene Extension: Implemented Forward Extend (extending forward) and Backward Extend (reconstructing the beginning) tools;
  • Optical Lens Physics: Lens properties are now embedded directly into the diffusion algorithm during generation, rather than applied as a post-filter overlay.

Simultaneously, the platform is developing two strategic directions: the Soul 2.0 architecture for consistent digital avatars and a REST API with access to 50+ specialized models for building custom video pipelines.


2. Cinema Studio 4.0: Key Changes and Capabilities

In the fourth version of Cinema Studio, basic shooting parameters have been moved out of the text description into independent system selectors. This eliminates neural network hallucinations and provides predictable results.

Generation Length Up to 30 Seconds

A single generation can last up to 30 seconds. This timing is sufficient not just for quick shot changes, but for a full dramatic scene with action development, exposition, and climax, without the need to manually stitch together two-second clips.

Multi-References: Up to 50 References Simultaneously

Up to 50 source images can now be passed into a single generation context. This allows you to simultaneously fixate:

  • Character Face: Portrait angles from different perspectives;
  • Commercial Product: Packaging, shape, logo, physical textures;
  • Location Reference: Interior, architecture, or natural landscape;
  • Visual Aesthetic: Overall shot style and color palette;
  • Prop Details: Key objects in the environment.

The model processes the entire pool of images within a unified embedding, preventing blurring of product details or distortion of the hero's face.

30+ Camera Movements

The Cinema Studio 4.0 toolkit includes more than 30 dynamic virtual camera movement presets, including:

  • POV (Point of View): First-person subjective filming with natural sway;
  • Robot Arm: Ultra-smooth, precise fly-throughs along complex industrial crane-manipulator trajectories;
  • Pan Left / Pan Right: Horizontal panning of the scene while maintaining the horizon;
  • Helicopter Shot: Large-scale wide shots from a bird's-eye view;
  • Tracking: Following a moving object while maintaining distance;
  • Pedestal Down / Up: Vertical tripod movement by height without tilting the optical axis.

Direct Scene Control Parameters

Interface controllers allow direct management of shot physics:

  • Tempo: Sets the internal rhythm of movement and editing dynamics;
  • Era: Stylizes the image according to a historical period;
  • Emotion Wheel: Fixes facial expressions and the emotional intensity of the character;
  • Colour Palette: Defines a professional grading range;
  • Lighting: Forms the layout scheme for studio light sources.

Additionally, the tool allows uploading a ready-made third-party video clip and seamlessly continuing its development forward or reconstructing preceding events via Forward / Backward Extend.


3. Basic Cinematic Configuration: Genre, Era, and Tempo

Before launching generation, the video sequence is calibrated using three fundamental parameters: staging genre, historical era, and editing rhythm.

Genre Stylization (Genre)

The Genre selector comprehensively affects the nature of lighting, framing density, camera tilt angle, and event speed:

  • General: Balanced image with standard cinema contrast without a pronounced genre accent;
  • Action: Active dynamic tracking of the object, tight framing, presence effect;
  • Epic: Panoramic wide shots, deep perspective, maximum environmental detail;
  • Drama: Medium and close-ups, focus on faces, long shot duration, soft chiaroscuro pattern;
  • Comedy: open frontal composition, high exposure, ample space for mise-en-scène and gesturing;
  • Horror: pronounced Dutch angles, claustrophobic framing, contrasting zones of deep darkness;
  • Noir: hard directional light, deep shadows with sharp boundaries, silhouette composition.

Historical Era (Era)

The Era selector automatically adjusts film grain type, color grading, and the character of optical distortions to match the selected decade:

  • Auto: neutral adaptation to the prompt style;
  • 1920s: monochrome contrast, pronounced flicker of early cinema, soft focus;
  • 1950s: saturated early Technicolor, characteristic warm studio lighting;
  • 1970s: muted warm palette, characteristic yellow-orange tones, vintage flare;
  • 1980s: neon accents, slight VHS blur, anamorphic horizontal flares;
  • 1990s: dense 35mm film grain, natural analog color reproduction;
  • 2000s: early digital clarity with a characteristic cold grade;
  • Modern: crystal-clear modern digital image with wide dynamic range (HDR);
  • Future: futuristic hyperrealism, chromatic aberrations, neon reflections.

Tempo and Editing Rhythm (Tempo)

The Tempo parameter controls the intensity of intra-frame movement:

  • Calm: meditative, smooth development of events for dramatic dialogues or landscapes;
  • Single Shot: continuous stable shot for demonstrating unboxing or product presentations;
  • Dynamic: energetic movement, optimal for advertising clips and promos;
  • Chaotic: fast, abrupt changes in angle and high action speed for chase scenes and action movies.

4. Emotion Wheel: Precise Character Emotion Tuning

Emotion Wheel eliminates the need to select dozens of synonyms in text descriptions to make a hero smile or get scared.

The system allows assigning a precise emotional vector to a specific character via text syntax:

text
@character_name Fear

or

text
@character_name Joy

Among the available basic emotional states:

  • Joy: open happiness, natural smile, relaxed facial expression;
  • Sadness: sorrow, downcast gaze, tears, loss of emotional tone;
  • Anger: aggression, furrowed brows, clenched jaws, tense neck;
  • Fear: fright, wide-open eyes, body recoiling backward;
  • Surprise: sudden astonishment, raised eyebrows, slightly open mouth;
  • Disgust: revulsion, wrinkled nose, raised upper lip;
  • Calm: neutral emotional tranquility, relaxed gaze;
  • Anticipation: tense expectation, focused attention.

The controller also allows adjusting the intensity of the feeling and setting complex transitional emotions.


5. Colour Palette and Lighting: Separate Color and Light Settings

In Cinema Studio 4.0, the color scheme and studio lighting no longer conflict with each other during the diffusion process.

Colour Palette

Users have access to over 50 professional cinematic presets, including:

  • Teal & Orange: classic Hollywood blockbuster contrast between warm skin tones and cool shadows;
  • Nostalgic Blue: muted cinematic blue with soft highlights;
  • Vintage Warm: cozy tube spectrum emphasizing amber tones;
  • Cyberpunk Neon: contrasting neon combinations of purple, turquoise, and violet;
  • Bleach Bypass: faded colors with harsh micro-contrast and silver midtones;
  • Black & White: pure high-contrast monochrome with rich gray gradation.

Studio Lighting Schemes (Lighting)

Lighting is configured independently of color correction:

  • Soft Natural: diffused soft daylight from a window without harsh shadow transitions;
  • Rim Light: backlight contour light that clearly separates the hero's silhouette from a dark background;
  • Volumetric (God Rays): volumetric light beams passing through smoke, fog, or blinds;
  • Studio Softbox: calibrated commercial lighting without deep drop-offs in shadows;
  • Harsh Sunlight: noon sun with deep, sharp shadows and maximum contrast;
  • Moody Shadows: dramatic chiaroscuro lighting pattern dominated by semi-darkness.

6. Optics and Frame Character: Camera, Lens, and Aperture

The physical model of optics determines composition, geometry, and the sense of spatial depth.

Shooting Systems (Camera)

  • Cinema Camera: classic operator tripod or rails with perfect stabilization;
  • Handheld: live handheld shooting with slight operator tremor, adding documentary authenticity;
  • Drone View: smooth aerial footage with a wide field of view;
  • Action Cam: ultra-wide dynamic angle with an immersive effect at the center of events.

Lens Sets (Lens)

  • Vintage Anamorphic: horizontal compression, oval flares, soft frame edges;
  • 35mm Film: standard reportage angle with natural perspective rendering;
  • 50mm Standard: field of view closest to the human eye;
  • 85mm Portrait: ideal face geometry without perspective distortion of the nose and cheekbones;
  • Ultra-Wide 14mm: deep dramatic perspective with slight edge stretching;
  • Macro Lens: extreme close-up with maximum detail of product textures.

Aperture and Depth of Field (Aperture)

  • f/1.2 — Shallow: extremely shallow depth of field, pronounced cinematic bokeh, background blurred into a creamy texture;
  • f/2.8 — Moderate: balanced separation of planes, sharp subject and soft surroundings;
  • f/8 — Medium: sharp foreground and midground, optimal for group and studio scenes;
  • f/16 — Deep Focus: maximum sharpness across the entire frame from the foreground boundary to the horizon.

7. Multi-References and Extend: Video Sequence Expansion

Single-shot generation rarely solves complex production tasks. For sophisticated shots, advanced input and timeline tools are employed.

Practical Workflow with 50 References

When batch-submitting up to 50 images, it is recommended to structure the reference stack by layers:

  1. Identity (3–5 photos): The model's face from different angles under neutral lighting;
  2. Product (5–10 photos): Close-ups of the product, including labels, held in hand, and viewed from various angles;
  3. Environment (5–10 photos): Interior references, walls, and environmental textures;
  4. Style / Color (2–3 photos): Target frame samples for lighting and color grading.

This separation ensures that the neural network extracts the appearance from the model references, while textures and logos are taken directly from the product photos, preventing them from blending together.

Forward Extend and Backward Extend Tools

If a shot turns out well but requires development, there is no need to generate an alternative from scratch:

  • Forward Extend: Analyzes the last frame of the clip, preserves the camera movement vector, and generates a seamless continuation forward in time;
  • Backward Extend: Analyzes the starting frame and builds the preceding 5–10 seconds, showing the scene's backstory.

The new fragment is formed within a unified lighting and geometric phase of the original video file.


8. Step-by-Step Scene Assembly in Cinema Studio 4.0

To achieve stable cinematic quality, it is recommended to follow a strict parameter configuration algorithm.

Step 1. Select Genre

Define the dramatic framework of the scene. The genre immediately calibrates composition and lighting. For a dynamic chase, select Action; for an intimate confession or dialogue — Drama.

Step 2. Configure Camera Optics

Sequentially set the camera type, lens, aperture, and palette.

Example of a ready setup:

text
Camera: 35mm Film Lens: Vintage Anamorphic Aperture: f/4 Moderate Colour Palette: Nostalgic Blue

At this same step, upload the prepared hero and product references.

Step 3. Select Editing Tempo

  • Calm — for emotional pauses and portraits;
  • Single Shot — for demonstrating interaction with the product;
  • Dynamic — for promo videos and music clips;
  • Chaotic — for explosive climaxes.

Step 4. Define Hero's Emotional State

Using the Emotion Wheel, bind an emotion to the character marker:

text
@character_name Joy

Step 5. Formulate Text Prompt

Since all cinematographic engineering (camera, lens, palette, light, tempo) is already configured via independent selectors, the text prompt is freed from technical clutter. It describes only the physical action:

text
Молодая женщина в льняной рубашке делает глоток кофе у панорамного окна, смотрит на дождливую улицу и улыбается входящему другу.

Step 6. Final Color Grading

After generating the base clip, the built-in editor allows for post-processing:

  • Temperature: Fine-tuning the warmth/coolness of the frame;
  • Contrast & Saturation: Balancing contrast and saturation;
  • Sharpness & Film Grain: Adding film texture or digital sharpness;
  • Highlights & Exposure: Aligning blown-out highlights and deep shadows.

9. Soul 2.0: Creating and Fixing a Persistent Character

The main problem with AI video generation is the change in the hero's appearance from shot to shot. Soul ID technology solves this issue by creating a digital identity fingerprint.

The neural network trains on a dataset and fixes:

  • Facial architecture and anthropometry;
  • Proportions of cheekbones, nose shape, and eye slits;
  • Skin tone and microtexture;
  • Hair type, growth line, and color.

After saving the Soul ID, the character can be placed in any location, with changes to lighting, age, costume, and shooting angles without losing recognizability.

Step-by-Step Creation of Soul ID

Navigate through the interface path: Character → Soul ID Character → Create

  1. Dataset Preparation: Upload between 20 (minimum threshold) and 70 high-quality photos of one person;
  2. Reference Requirements:
    • Diverse angles (frontal, profile, three-quarter, looking up/down);
    • Stable, neutral diffuse lighting;
    • High resolution without compression artifacts;
    • Absence of strong filters, masks, and other faces in the frame;
  3. Model Training: The system forms a unique ID token (e.g., @alex_soul), which is saved in the profile library.

Composing Prompts for Trained Characters

After training the Soul ID, there is no need to describe eye color or chin shape in detail. The entire appearance is invoked via the system token:

text
@alex_soul сидит за столиком в уличном парижском кафе, в темно-синем блейзере, читает утреннюю газету, мягкое солнце, малая глубина резкости, 9:16

Soul handles anatomical identity, while the prompt handles clothing, mise-en-scène, environment, and movement.


10. Higgsfield API: Autonomous Generation Without Interface

Higgsfield REST API allows integrating photo and video generation directly into your own B2B systems, web services, and automated scripts without manual work in the web interface. The API catalog includes more than 50 advanced models.

Video Generation Model Catalog

  • Seedance 2.5: Synchronous generation of video, realistic voice, lip-sync, and sound environment;
  • Kling 3.0: Top-tier cinematic motion physics and detail;
  • Wan 3.0: Generation of complex physical interactions;
  • MiniMax H3: High-quality human body dynamics;
  • LTX 2.5 Pro: Ultra-fast generation with high pixel density;
  • PixVerse 6: Stable rendering of object motion;
  • Grok Imagine Video: Creative stylization;
  • Higgsfield DoP: Specialized virtual operator.

Image Generator Catalog

  • Soul 2 & Soul Cinema: Generation of photorealistic portraits and cinematic shots with face fixation;
  • Marketing Studio Image: Product and packaging photography;
  • Recraft 4.1, Ideogram 4.0, Qwen Image 3: Generation of graphics, typography, and photorealism.

Connecting to the API

  1. Register a developer account at console.higgsfield.ai;
  2. Top up your USD balance with a card;
  3. Generate a secret API token;
  4. Send HTTP requests via Python SDK, TypeScript SDK, or cURL.

The API format is standardized: switching between Kling, Seedance, or Wan is done by changing the "model_id" parameter.


11. Practical API Pipeline: Character → Frame → Video

The most cost-effective and high-quality way to create videos is a multi-stage pipeline where each step is handled by an optimal specialized model.

Step 1. Generating the Base Character or Product

Generate a static character frame using Soul 2 ($0.0032 per generation) or a product shot using Marketing Studio Image ($0.0059):

bash
curl -X POST https://api.higgsfield.ai/v1/images/generations -H "Authorization: Bearer $HIGGSFIELD_API_KEY" -H "Content-Type: application/json" -d '{ "model": "soul-2", "prompt": "Professional businesswoman in modern office, neutral expression, high detail", "aspect_ratio": "9:16" }'

Step 2. Refining the Frame to Cinematic Quality

The resulting image is passed to Soul Cinema ($0.0032). The model establishes cinematic lighting, depth of field, and color grading. The higher the quality and accuracy of the initial static frame, the fewer artifacts will occur during subsequent animation.

Step 3. Animating the Frame in a Video Model

The prepared static frame is sent to Kling 3.0 or Seedance for dynamics generation:

bash
curl -X POST https://api.higgsfield.ai/v1/videos/generations -H "Authorization: Bearer $HIGGSFIELD_API_KEY" -H "Content-Type: application/json" -d '{ "model": "kling-3.0", "image_url": "https://storage.higgsfield.ai/renders/step2_cinematic.jpg", "prompt": "Camera slowly tracks forward as the woman turns head and speaks, natural movement", "duration_seconds": 10 }'

Step 4. Retrieving the Final Result

Generation is performed asynchronously. The API returns a request_id, whose status can be polled or received via webhook. The finished video is stored on Higgsfield servers for at least 7 days, after which it must be retrieved into your own S3 storage.


12. Pricing and Cost of Working with the Higgsfield API

The API operates on a transparent pay-as-you-go model with prepaid balance top-ups. Payment for video is calculated per second, while images are charged per unit.

Model / EngineBilling TypeUnit Cost
Soul 2Image$0.0032 / generation
Soul CinemaImage$0.0032 / generation
Marketing Studio ImageImage$0.0059 / generation
Qwen Image 3Image$0.03 / generation
Recraft 4.1Image$0.035 / generation
Grok Imagine 2.0Image$0.06 / generation
Ideogram 4.0Image$0.06 / generation
Higgsfield DoPVideo Generation$0.125 / generation
Seedance 2.5Video with Audio$0.0738 / second
Kling 3.0High-Quality Video$0.112 / second
PixVerse 6Video$0.115 / second
MiniMax H3Video$0.13 / second
LTX 2.5 ProVideo$0.17 / second
Wan 3.0Video$0.20 / second
Grok Imagine Video 1.5Video$0.25 / second
Tip

Cost calculation for a full production cycle: Portrait generation in Soul 2 ($0.0032) + cinematic refinement in Soul Cinema ($0.0032) + 10 seconds of video in Kling 3.0 ($1.12) totals $1.126 — less than $1.30 for a finished premium 10-second shot!

Key billing rules:

  • No hidden subscription fees or monthly subscriptions;
  • The API balance is independent of the web subscription at higgsfield.ai (credits do not transfer);
  • Unsuccessful or rejected generations do not deduct funds;
  • Deposits are valid for 1 calendar year from the date of deposit.

13. Tool Selection Matrix for Specific Tasks

Depending on the project format, choose the optimal implementation path:

  • Cinema Studio 4.0: Ideal when you need a bespoke authorial clip with maximum manual control over direction, camera, optics, and character emotions;
  • Soul 2.0 (Soul ID): Essential for narrative clips and series where the same character must appear in different locations without facial deformation;
  • Higgsfield API: Mandatory for integrating video generation into mobile and web applications, creating SaaS services, and automatically rendering hundreds of creatives;
  • Multi-model pipeline (Soul 2 → Soul Cinema → Kling/Seedance): Provides the best quality-to-cost ratio for batch production of commercial video content.

14. Ready-made UGC Workflow: Producing Ads from a Product Photo via Claude and Higgsfield

There is a ready-made open-source production pipeline that allows you to turn a single static product photo into a full-fledged converting UGC video using the combination of Claude Code / Codex and Higgsfield.

StagePipeline PhaseOutput ArtifactRole and Optimization
1Product ProfileIsolated product referenceFixing dimensions, materials, and ergonomics
2Marketing BriefConcept and storyboard for 3 shotsApproving creative before launching render
3Base CharacterHero portrait without productClean face identity reference
4StoryboardSingle panoramic triptych 9:16Synchronous fixation of actor and product geometry
5Script & Audio15-second text with phoneticsDirective no subtitles for clean frame
6Seedance RenderVideo + Voice + Lip-sync + AmbientResolution 720p (67 credits instead of 134 for 1080p)
7Post-productionFinal ad MP4 videoLocal editing: music, zoom accents, subtitles

End-to-end UGC video ad generation pipeline:

(1) Product Profile ➔ (2) Marketing Brief ➔ (3) Base Character ➔ (4) Storyboard (Triptych) ➔ (5) Script & Audio ➔ (6) Seedance Render ➔ (7) Post-production

The pipeline takes a product photo, a target audience description, and a basic brief as input. The process consists of seven sequential stages:

1. Product Profile

Claude performs a detailed analysis of the product:

  • Records dimensions, geometric shape, and materials;
  • Describes packaging features and label readability;
  • Determines real-world ergonomics of use (how fingers hold the item, how the lid opens);
  • Creates an isolated clean product reference on a transparent background.

2. Marketing Brief

A conceptual framework for the ad video is formed:

  • Defining the main hook (audience problem / trigger);
  • Dramaturgical structure of a 15-second video;
  • Storyboarding into three key shots;
  • Technical rendering parameters.

Approving the brief at an early stage prevents rework during the expensive video generation phase.

3. Base Character

One high-quality photorealistic portrait of the model is generated — strictly without the product in hand. This portrait becomes the base identity reference for locking down the actor's appearance.

4. Storyboard

A single panoramic triptych consisting of three vertical frames in 9:16 format is created:

  1. Selfie-hook: Close-up of the hero voicing a problem or intriguing thesis;
  2. Macro-shot: Extreme close-up of interaction with the product (applying cream, pressing a button, taking a sip of a drink);
  3. Call-to-Action: Medium shot of a satisfied hero with a final recommendation.

The hero's photo and the product reference are fed into the model simultaneously, locking down the geometry of both objects before video generation starts.

5. Script & Audio Design

Claude generates a 15-second voiceover script, split into 3 timing blocks.

  • Phonetic adaptation: Complex brand names are spelled out via transliteration exactly as they should be pronounced by the text-to-speech synthesizer to avoid stress distortion;
  • no subtitles directive: A flag excluding system captions is mandatory in the video generator prompt to prevent visual noise generation within the neural network.

6. Video Generation in Seedance

Three references are passed to Seedance simultaneously: the storyboard, the character portrait, the product photo, and the ready audio script. The model creates in a single pass:

  • Realistic video with organic micro-movements;
  • Natural narrator voice;
  • Precise lip-sync (synchronization of lips with sound);
  • Background spatial ambience (ambient noise).

Rendering is recommended at 720p resolution: generation costs only 67 credits, whereas 1080p costs twice as much without visible benefits for mobile social media.

7. Final Local Post-processing

The final stage is performed on the local machine via an automatic ffmpeg script without consuming video credits:

  • Overlaying licensed music background;
  • Generating dynamic word-by-word subtitles with active word highlighting;
  • Adding smooth zoom effects and transitions at cut points.
Important

Philosophy of Production: Never expect a neural network to produce a finished ad video from a single abstract prompt. Break the process into cheap preparatory steps (profile, brief, storyboard) and launch the expensive video generation only after all elements are verified.

This guide is completely free. If it saved you an evening, you can support the project's growth.
Support the author