Skip to content

Introduction

Hello creators and developers! Welcome to the V2Fun API.

V2Fun is an AI-powered platform for creating 3D content from text, images, and other inputs. It enables developers to generate, edit, and animate assets through a unified set of APIs.

With V2Fun, you can build end-to-end workflows, from image generation and 3D modeling to animation and rendering, without managing complex 3D pipelines.

Capabilities

Here is a quick overview of what you can build with the V2Fun API:

🎨 Image Generation and Editing

Create and refine images using natural language or reference inputs.

Supports text-to-image generation, image editing, and prompt enhancement as inputs for downstream 3D and animation workflows.

🧱 3D Model Creation

Generate structured 3D assets from text or images, and prepare them for real-world use.

Includes mesh generation, texture creation, remeshing, and multi-format conversion (for example: GLB, FBX, USDZ).

🕺 Motion and Animation

Bring static models to life with animation workflows.

Supports motion retrieval using natural language, and animation retargeting across different characters.

🎥 Video Motion Capture

Detect human pose from a specific video frame, or extract motion data from a video segment.

Pose detection returns 27-point 2D/3D skeleton keypoints for targeting individuals in multi-person scenes. Motion detection outputs a standard BVH file that can be imported directly into Blender, UE5, or Unity to drive 3D characters.

🎬 Rendering and Output

Generate visual outputs for preview, display, and distribution.

Includes 3D model rendering, animation previews, and format-ready outputs for various platforms.

⚙️ Task-Based Execution

All capabilities are exposed through a task-based API designed for scalability and flexibility.

Supports asynchronous workflows with task polling, as well as optional synchronous execution for quick testing.