Precise visual control through text, images, and freehand scribbles
Recently, unified generation and editing models have achieved remarkable success, but their reliance on text prompts makes it difficult to specify exact edit locations and fine-grained visual details. We therefore introduce scribble-based editing and generation, enabling flexible creation through a graphical interface that combines text, images, and freehand scribbles.
DreamOmni3 addresses both data creation and framework design. Its synthesis pipeline covers four editing tasks—scribble and instruction-based editing, scribble and multimodal instruction-based editing, image fusion, and doodle editing—and three generation tasks—scribble and instruction-based generation, scribble and multimodal instruction-based generation, and doodle generation.
Instead of binary masks, DreamOmni3 jointly feeds the original image and its scribble-modified counterpart into the model. Shared index and position encodings precisely align scribbled regions, while joint training with a vision-language model improves the understanding of abstract user marks. Comprehensive benchmarks and extensive experiments demonstrate strong performance across all seven tasks.
Why Scribbles?
Beyond Traditional Binary-mask Inpainting
Binary masks only indicate where to edit. Color-coded scribbles can additionally express what each region means and how multiple source and reference regions relate to one another.
Comparison of the binary-mask input scheme and the scribble input scheme.
01
Color-coded relationships
Different brush colors create explicit correspondences between source and reference regions, making complex spatial and logical requirements easy to describe in natural language. A binary mask can only separate black from white.
02
Multiple controls in one image
Users can draw several regions on a single image and distinguish them by color. Binary-mask workflows require a separate mask layer for every region, increasing both input complexity and computation.
03
Native fit for RGB models
Scribbles remain embedded in natural RGB images, allowing unified generation and editing models to reuse their pretrained visual understanding. Black-and-white masks cannot exploit those capabilities as effectively.
Scribble-based Editing
Editing Results
Four editing tasks · 42 examples · 1–2 images + instruction → result
01
Scribble and Instruction-based Editing
Example S01Scribble and instruction-based editing
Image 1
Result
Delete the person inside the red circle.
Example S02Scribble and instruction-based editing
Image 1
Result
Delete the object inside the red circle.
Example S03Scribble and instruction-based editing
Image 1
Result
Delete the object inside the red circle.
Example S04Scribble and instruction-based editing
Image 1
Result
Delete the cat inside the green circle.
Example S05Scribble and instruction-based editing
Image 1
Result
Add a laptop at the position marked with the red circle.
Example S06Scribble and instruction-based editing
Image 1
Result
Add a woman who is swimming in the red circle.
Example S07Scribble and instruction-based editing
Image 1
Result
Add a Husky dog in the red circle.
Example S08Scribble and instruction-based editing
Image 1
Result
Add a boy inside the black circle.
Example S09Scribble and instruction-based editing
Image 1
Result
Replace the object in the red circle with a cup.
Example S10Scribble and instruction-based editing
Image 1
Result
Replace the object in the red circle with a toy lion.
Example S11Scribble and instruction-based editing
Image 1
Result
Replace the person in the red circle with a panda.
Example S12Scribble and instruction-based editing
Image 1
Result
Replace the plastic bottle in the red circle with a doll.
Example S13Scribble and instruction-based editing
Image 1
Result
Add black eyeliner to the person in the red circle.
Example S14Scribble and instruction-based editing
Image 1
Result
Change the hairstyle of the man in the red circle to a bald head.
Example S15Scribble and instruction-based editing
Image 1
Result
Make the person in the black circle jump.
Example S16Scribble and instruction-based editing
Image 1
Result
Change the material of the bag in the red circle to brown leather.
Example 01Scribble and instruction-based editing
Image 1
Result
Change the color of the car in the black circle to red.
Example 02Scribble and instruction-based editing
Image 1
Result
Replace the car in the red circle with a motorcycle.
Example 03Scribble and instruction-based editing
Image 1
Result
Add a sun inside the black circle in a clear sky.
Example 04Scribble and instruction-based editing
Image 1
Result
Delete the object inside the black circle.
Example 12Scribble and instruction-based editing
Image 1
Result
Add a black car inside the blue circle.
Example 13Scribble and instruction-based editing
Image 1
Result
Replace the woman in the red circle with a man.
02
Scribble and Multimodal Instruction-based Editing
Example S17Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Insert the toy of the second image into the red circle in the first image.
Example S18Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Insert the woman in the red circle from the second image into the red circle in the first image.
Example S19Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Insert the toy of the second image into the red circle in the first image.
Example S20Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Insert the man in the red circle from the second image into the red circle in the first image.
Example S21Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Replace the clothes of the woman in the red circle of the first image with the T-shirt in the second image.
Example S22Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Replace the girl in the red circle of the first image with the girl in the red circle of the second image.
Example S23Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Make the woman's clothing in the red circle of the first image have the same color scheme as the sweater in the red circle of the second image.
Example S24Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Make the woman in the red circle of the first image have the same hairstyle as the woman in the second image.
Example 05Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Insert the woman in the red circle from the second image into the red circle in the first image.
Example 06Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Insert the toy in the red circle from the second image into the red circle in the first image.
Example 07Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Replace the toy in the red circle of the first image with the toy in the red circle of the second image.
Example 08Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Make the object in the red circle of the first image have the same material as the box in the red circle of the second image.
Example 11Scribble and multimodal instruction-based editing
Image 1
Image 2
Result
Make the bag in the red circle of the first image have the same color scheme as the printer in the red circle of the second image.
03
Doodle Editing
Example S27Doodle editing
Image 1
Result
Render the doodled people, dog, and door realistically in the scene.
Example S28Doodle editing
Image 1
Result
Render the mug doodle in a realistic style.
Example 09Doodle editing
Image 1
Result
Render the doodle in a realistic style.
04
Image Fusion
Example S25Image fusion
Image 1
Result
Fuse the objects into the room scene.
Example S26Image fusion
Image 1
Result
Mix the girl. The girl is sitting on the sofa.
Example 10Image fusion
Image 1
Result
Fuse the toy into the scene.
Example 14Image fusion
Image 1
Result
The man walks on the rail. Fuse the image.
Scribble-based Generation
Generation Results
Three generation tasks · 16 examples · 1–4 images + instruction → result
01
Scribble and Instruction-based Generation
Example 05Scribble and instruction-based generation
Image 1
Result
A man is in the red circle, and a woman with long hair is in the green circle. They are shaking hands, with a volcano in the background.
Example 06Scribble and instruction-based generation
Image 1
Result
A girl is standing in the red circle, and a boy is sitting in the green circle. They are watching the sunset.
02
Scribble and Multimodal Instruction-based Generation
Example S01Scribble and multimodal instruction-based generation
Image 1
Image 2
Image 3
Result
In the spaceship, the man from image 2 is standing in the red circle of image 1, and the woman from image 3 is standing in the black circle of image 1.
Example S02Scribble and multimodal instruction-based generation
Image 1
Image 2
Image 3
Result
The woman from image 2 is standing in the black circle of image 1, and the bird from image 3 is flying in the red square of image 1. The background is a park.
Example S03Scribble and multimodal instruction-based generation
Image 1
Image 2
Image 3
Image 4
Result
In front of a house, the man from image 2 is standing in the green circle of image 1, the car from image 4 is parked in the black circle of image 1, and the backpack from image 3 is placed in the red square of image 1 on the car's hood.
Example S04Scribble and multimodal instruction-based generation
Image 1
Image 2
Image 3
Image 4
Result
The woman from image 2 is sitting in the red circle of image 1 and has the same makeup as the woman from image 4. The man from image 3 is standing in the black circle. The background is a beach.
Example 01Scribble and multimodal instruction-based generation
Image 1
Image 2
Result
In the black circle of image 1, the man from image 2 is standing, with a vast blue ocean as the background.
Example 02Scribble and multimodal instruction-based generation
Image 1
Image 2
Image 3
Result
In the living room, the woman from image 2 is standing in the red circle of image 1, and the cat from image 3 is sleeping in the black square of image 1, resting on the table.
Example 03Scribble and multimodal instruction-based generation
Image 1
Image 2
Image 3
Image 4
Result
On the beach, the man from image 2 is standing in the green circle of image 1, the woman from image 3 is standing in the red circle of image 1, and the dog from image 4 is standing in the black box of image 1.
Example 04Scribble and multimodal instruction-based generation
Image 1
Image 2
Image 3
Image 4
Result
In the living room, the bottle from image 3 is placed in the black circle of image 1 and has the same pattern as image 4. The cat from image 2 stands in the red circle.
03
Doodle Generation
Example S05Doodle generation
Image 1
Result
Render the plane in the image to look realistic.
Example S06Doodle generation
Image 1
Result
Render the building in the image to look realistic.
Example S07Doodle generation
Image 1
Result
Render the bed and lamp in the image to look realistic.
Example S08Doodle generation
Image 1
Result
Snowman and house.
Example 07Doodle generation
Image 1
Result
Fish, turtle, and seaweed in the ocean world. Render them realistically.
Example 08Doodle generation
Image 1
Result
A person with rabbit ears is watering a tree.
Qualitative Benchmark
Compare with the Alternatives
Every image is shown in full, without cropping.
Scribble-based Editing
Editing 01
Change the color of the car in the black circle to red.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 02
Replace the car in the red circle with a motorcycle.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 03
Add a sun inside the black circle in a clear sky.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 04
Delete the object inside the black circle.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 05
Insert the woman in the red circle from the second image into the red circle in the first image.
Input 1
Input 2
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 06
Insert the toy in the red circle from the second image into the red circle in the first image.
Input 1
Input 2
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 07
Replace the toy in the red circle of the first image with the toy in the red circle of the second image.
Input 1
Input 2
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 08
Make the object in the red circle of the first image have the same material as the box in the red circle of the second image.
Input 1
Input 2
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 09
Render the doodle in a realistic style.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 10
Fuse the toy into the scene.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 11
Make the bag in the red circle of the first image have the same color scheme as the printer in the red circle of the second image.
Input 1
Input 2
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 12
Add a black car inside the blue circle.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 13
Replace the woman in the red circle with a man.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Editing 14
The man walks on the rail. Fuse the image.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Scribble-based Generation
Generation 01
In the black circle of image 1, the man from image 2 is standing, with a vast blue ocean as the background.
Input 1
Input 2
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Generation 02
In the living room, the woman from image 2 is standing in the red circle of image 1, and the cat from image 3 is sleeping in the black square of image 1, resting on the table.
Input 1
Input 2
Input 3
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Generation 03
On the beach, the man from image 2 is standing in the green circle of image 1, the woman from image 3 is standing in the red circle of image 1, and the dog from image 4 is standing in the black box of image 1.
Input 1
Input 2
Input 3
Input 4
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Generation 04
In the living room, the bottle from image 3 is placed in the black circle of image 1 and has the same pattern as image 4. The cat from image 2 stands in the red circle.
Input 1
Input 2
Input 3
Input 4
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Generation 05
A man is in the red circle, and a woman with long hair is in the green circle. They are shaking hands, with a volcano in the background.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Generation 06
A girl is standing in the red circle, and a boy is sitting in the green circle. They are watching the sunset.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Generation 07
Fish, turtle, and seaweed in the ocean world. Render them realistically.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Generation 08
A person with rabbit ears is watering a tree.
Input 1
DreamOmni3 (Ours)
GPT-4o
Nano Banana
DreamOmni2
Qwen-Edit-2509
Kontext
OmniGen2
Citation
@article{xia2025dreamomni3,
title = {DreamOmni3: Scribble-based Editing and Generation},
author = {Xia, Bin and Peng, Bohao and Liu, Jiyang and Wu, Sitong
and Li, Jingyao and Huang, Junjia and Zhao, Xu and Wang,
Yitong and Chu, Ruihang and Yu, Bei and others},
journal = {arXiv preprint arXiv:2512.22525},
year = {2025}
}