jaaari / kokoro-82m
Kokoro v1.0 - text-to-speech (82M params, based on StyleTTS2)
97.9M runs
andreasjansson / clip-features
Return CLIP features for the clip-vit-large-patch14 model
162.6M runs
prunaai / p-image-edit
A sub 1 second 0.01$ multi-image editing model built for production use cases. For image generation, check out p-image here: https://replicate.com/prunaai/p-image
35.6M runs
851-labs / background-remover
Remove backgrounds from images.
27.1M runs
krea
/
krea-2-medium
Foundation image model from Krea, tuned for expressive illustration, anime, and painterly styles. Fast and consistent across artistic directions.
13.8K runs
alibaba
/
happyhorse-1.0
Alibaba's Happy Horse 1.0 generates videos from text prompts or animates a single image into video. Supports 720p and 1080p, 3-15 second durations, and five aspect ratios.
27K runs
openai
/
gpt-image-2
OpenAI's state-of-the-art image generation model. Create and edit images from text with strong instruction following, sharp text rendering, and detailed editing.
12.9M runs
anthropic
/
claude-opus-4.7
Anthropic's most capable model with a step-change improvement in agentic coding, better vision, and stronger multi-step reasoning
145.6K runs
google
/
gemini-3.1-flash-tts
Google's fast, expressive text-to-speech model with 30 voices and 70+ language support
259.6K runs
minimax
/
music-2.6
Generate full-length songs or instrumentals from a text prompt, with optional auto-generated lyrics
17.8K runs
bytedance
/
seedance-2.0
ByteDance's multimodal video generation model with native audio, multimodal reference inputs, and intelligent duration control.
1.1M runs
prunaai
/
p-video-avatar
p-video-avatar is the fastest and cheapest avatar/lipsync video model on the market.
89.8K runs
bytedance
/
seedream-5-lite
Seedream 5.0 lite: image generation with built-in reasoning, example-based editing, and deep domain knowledge
3M runs
xai
/
grok-imagine-video
Generate videos using xAI's Grok Imagine Video model
1.3M runs
black-forest-labs
/
flux-2-max
The highest fidelity image model from Black Forest Labs
3.6M runs
google
/
nano-banana-2
Google's fast image generation model with conversational editing, multi-image fusion, and character consistency
14M runs
Official models are always on, maintained, and have predictable pricing.
Google's fastest image generation model — the lightweight, low-cost version of Nano Banana 2, for rapid creation and editing
Qwen3.7-Plus is Alibaba's cost-effective multimodal model with vision-language understanding, a 1 million token context window, and strong agentic coding and tool use.
A lower-cost variant of Seedance 2.0 for high-volume video generation with multimodal inputs and native audio.
Alibaba's Happy Horse 1.1 generates videos from text, animates a single image, or builds a video from multiple reference images. Supports 720p and 1080p, 3-15 second durations, and five aspect ratios.
Top-quality agentic image model with multi-step reasoning, candidate scoring, and adjustable thinking effort
Speed-optimized variant of Riverflow 2.5 for production and latency-sensitive workflows
Luma's reasoning video model. Generates cinematic 5s or 10s video from text or images, with native HDR and EXR export for professional production pipelines.
Claude Fable 5 from Anthropic: the next generation of intelligence for the hardest knowledge work and coding problems.
The highest quality Ideogram v4 model. v4 creates images with stunning realism, creative designs, and consistent styles
Balance speed, quality and cost. Ideogram v4 creates images with stunning realism, creative designs, and consistent styles
Edit one frame to update an entire video. Aleph 2.0 is Runway's in-context video editor: longer clips (up to 30s), multi-shot edits, and image-level precision via keyframe references.
Image-to-video with synchronized audio using xAI's Grok Imagine Video 1.5 preview model
Krea's flagship foundation image model. Larger and more flexible than Krea 2 Medium, with particular strength in photorealism and expressive artistic styles.
Foundation image model from Krea, tuned for expressive illustration, anime, and painterly styles. Fast and consistent across artistic directions.
Claude Sonnet 4.6 from Anthropic: a full upgrade to coding, computer use, long-context reasoning, agent planning, knowledge work, and design, with a 1 million token context window in beta.
Upscale and enhance video up to 4K at 60fps, with scene-aware presets for AI-generated content, short dramas, UGC, and film restoration.
Google's fast multimodal model with frontier reasoning across agents, coding, and long-context tasks
Create realistic talking avatar videos from text with HeyGen's Avatar V engine — the newest, highest-quality avatar engine with cross-reference-driven animation.
Granite Vision 4.1 4B is a vision-language model (VLM) that delivers frontier-level performance on structured document extraction tasks — chart extraction, table extraction, and semantic key-value pair extraction — in a compact 4B parameter footprint
p-video-animate animates a reference image with the motion and audio of a source video. Optimized for speed and cost — 5.24s per 1s of video.
Use AI to generate images & photos with an API
Use AI to understand, describe, and caption videos with an API
Use AI for text-to-speech or to clone your voice via API
Use AI to generate images from a face with an API
Use AI to generate videos with an API
Use AI to upscale and enhance images with an API
Use AI to generate music with an API
Use AI to edit any image via API
Use AI to transcribe speech to text with an API
Use AI For Optical Character Recognition (OCR) to extract text from images via API
Use AI to remove backgrounds from images and videos with an API
FLUX AI models by Black Forest Labs: image generation & editing via API
Use AI to restore images via API
Use AI to upscale, restore, extend, and enhance videos with an API
Detect NSFW content in images and text
Classify text by sentiment, topic, intent, or safety
Identify speakers from audio and video inputs
Replace faces across images with natural-looking results.
Transform rough sketches into polished visuals
Generate custom emojis from text or images
Create anime-style characters, scenes, and animations
Use AI to generate videos from images with an API
Chat with images — visual Q&A, analysis, and reasoning via API
Use AI to generate captions and descriptions from images with an API
Use AI to edit, restyle, extend, and remix videos with an API
WAN family of models: open-source video, image, and audio generation
Generate 3D objects, meshes, and textures from text or images with an API
Official models are always on, predictably priced, and have a stable API.
Explore Large Language Models (LLMs) for chat, generation & NLP tasks via API
Try AI Models for free: video generation, image generation, upscaling, and photo restoration
Use AI to generate lipsync videos with an API
Use AI to control image generation with an API
Embedding models for AI search and analysis
Use AI object detection and segmentation models to distinguish objects in images & videos
Flux fine-tunes: build and run custom AI image models via API
Kontext fine-tunes: Build custom AI image models with an API
Create songs with voice cloning models via API
AI media utilities: auto-caption, watermark, frame extraction & more via API
Browse the diverse range of qwen-image fine-tunes the community has custom-trained on Replicate.
mewforest / anillustrious-multi-controlnet-lora
Anime image generation with AnIllustrious SDXL, LoRA support, and multi-ControlNet (pose, depth, canny)
11 runs
maxzhao / corridorkey-matte
5 runs
lex2029 / vlogme-avatar-bridge
Create a vertical talking-avatar video from a centered photo and speech audio.
26 runs
0xdino / cyberrealistic-pony-semireal-v6
Semirealistic image generation.
87 runs
myaiteam2 / websitescraper2
Scrapes most websites for emails, phone numbers and social links, Plus html and text content
91.5K runs
google / nano-banana-2-lite
Google's fastest image generation model — the lightweight, low-cost version of Nano Banana 2, for rapid creation and editing
18.7K runs
jaweii / hisam
20 runs
qwen / qwen3-7-plus
Qwen3.7-Plus is Alibaba's cost-effective multimodal model with vision-language understanding, a 1 million token context window, and strong agentic coding and tool use.
2.3K runs
ultralytics / yolo26-sem
Ultralytics YOLO26 semantic segmentation (Cityscapes), selectable size n/s/m/l/x.
21 runs
ultralytics / yolo26-pose
Ultralytics YOLO26 pose estimation (COCO-Pose), selectable size n/s/m/l/x.
8 runs
ultralytics / yolo26-obb
Ultralytics YOLO26 oriented bounding box detection (DOTAv1), selectable size n/s/m/l/x.
5 runs
ultralytics / yolo26-seg
Ultralytics YOLO26 instance segmentation (COCO-Seg), selectable size n/s/m/l/x.
14 runs
const
replicate =
new
Replicate
({
auth
: process.
env
.
REPLICATE_API_TOKEN
})
const
model =
"
const
input = {
prompt
:
"a
};
const
[output] =
await
replicate.
run
(model, { input });
console
.
log
(output);
A poolside patio at sunset with vintage lounge chairs.
black-forest-labs/flux-2-pro
A soft armchair shaped like a peeled banana.
google/nano-banana-pro
A woman relaxing in a french bookstore.
bytedance/seedream-4
A futuristic robot looking into the distance.
black-forest-labs/flux-pro
An abstract painting of a sunrise.
black-forest-labs/flux-proWith Replicate you can
bytedance / seedream-4.5
Seedream 4.5: Upgraded Bytedance image model with stronger spatial understanding and world knowledge
34.3M runs
black-forest-labs / flux-2-flex
Max-quality image generation and editing with support for ten reference images
430.1K runs
openai / gpt-image-1.5
OpenAI's latest image generation model with better instruction following and adherence to prompts
14M runs
bytedance / seedream-5-lite
Seedream 5.0 lite: image generation with built-in reasoning, example-based editing, and deep domain knowledge
3M runs
google / imagen-4-ultra
Use this ultra version of Imagen 4 when quality matters more than speed and cost
1.7M runs
All the latest models are on Replicate. They’re not just demos — they all actually work and have production-ready APIs.
AI shouldn’t be locked up inside academic papers and demos. Make it real by pushing it to Replicate.
openai / gpt-image-2
OpenAI's state-of-the-art image generation model. Create and edit images from text with strong instruction following, sharp text rendering, and detailed editing.
12.9M runs
anthropic / claude-opus-4.7
Anthropic's most capable model with a step-change improvement in agentic coding, better vision, and stronger multi-step reasoning
145.6K runs
google / gemini-3.1-flash-tts
Google's fast, expressive text-to-speech model with 30 voices and 70+ language support
259.5K runs
prunaai / p-video-avatar
p-video-avatar is the fastest and cheapest avatar/lipsync video model on the market.
89.7K runs
bytedance / seedream-5-lite
Seedream 5.0 lite: image generation with built-in reasoning, example-based editing, and deep domain knowledge
3M runs
xai / grok-imagine-video
Generate videos using xAI's Grok Imagine Video model
1.3M runs
black-forest-labs / flux-2-max
The highest fidelity image model from Black Forest Labs
3.6M runs
google / nano-banana-2
Google's fast image generation model with conversational editing, multi-image fusion, and character consistency
14M runs
You can get started with any model with just one line of code. But as you do more complex things, you can fine-tune models or deploy your own custom code.
Our community has already published thousands of models that are ready to use in production. You can run these with one line of code.
import
replicate
output = replicate.run(
"black-forest-labs/flux-dev"
,
input
={
"aspect_ratio"
:
"1:1"
,
"num_outputs"
:
1
,
"output_format"
:
"jpg"
,
"output_quality"
:
80
,
"prompt"
:
"An astronaut riding a rainbow unicorn, cinematic, dramatic"
,
}
)
print
(output)
You can improve models with your own data to create new models that are better suited to specific tasks.
Image models like SDXL can generate images of a particular person, object, or style.
Train a model:
training = replicate.trainings.create(
destination=
"mattrothenberg/drone-art"
version=
"ostris/flux-dev-lora-trainer:e440909d3512c31646ee2e0c7d6f6f4923224863a6a10c494606e79fb5844497"
,
input
={
"steps"
:
1000
,
"input_images"
:
,
"trigger_word"
:
"TOK"
,
},
)
This will result in a new model:
Fantastical images of drones on land and in the sky
0 runs
mattrothenberg / drone-art
Fantastical images of drones on land and in the sky
0 runs
Then, you can run it with one line of code:
output = replicate.run(
"mattrothenberg/drone-art:abcde1234..."
,
input
={
"prompt"
:
"a photo of TOK forming a rainbow in the sky"
}),
)
You aren’t limited to the models on Replicate: you can deploy your own custom models using Cog , our open-source tool for packaging machine learning models.
Cog takes care of generating an API server and deploying it on a big cluster in the cloud. We scale up and down to handle demand, and you only pay for the compute that you use.
First, define the environment your model runs in with cog.yaml:
build:
gpu:
true
system_packages:
-
"libgl1-mesa-glx"
-
"libglib2.0-0"
python_version:
"3.10"
python_packages:
-
"torch==1.13.1"
predict:
"predict.py:Predictor"
Next, define how predictions are run on your model with predict.py:
from
cog
import
BasePredictor, Input, Path
import
torch
class
Predictor
(
BasePredictor
):
def
setup
(
self
):
"""Load the model into memory to make running multiple predictions efficient"""
self
.model = torch.load(
"./weights.pth"
)
# The arguments and types the model takes as input
def
predict
(
self,
image: Path = Input(description=
"Grayscale input image"
)
) -> Path:
"""Run a single prediction on the model"""
processed_image = preprocess(image)
output =
self
.model(processed_image)
return
postprocess(output)
Thousands of businesses are building their AI products on Replicate. Your team can deploy an AI feature in a day and scale to millions of users, without having to be machine learning experts.
Learn more about our enterprise plansIf you get a ton of traffic, Replicate scales up automatically to handle the demand. If you don't get any traffic, we scale down to zero and don't charge you a thing.
Replicate only bills you for how long your code is running. You don't pay for expensive GPUs when you're not using them.
Deploying machine learning models at scale is hard. If you've tried, you know. API servers, weird dependencies, enormous model weights, CUDA, GPUs, batching.
Prediction throughput (requests per second)
Metrics let you keep an eye on how your models are performing, and logs let you zoom in on particular predictions to debug how your model is behaving.
Building with Replicate and Cloudflare, you can wake up with an idea and watch it hit the front page of Hacker News by the time you go to bed.