Any-to-Any
Transformers
Diffusers
Safetensors
English
llada2_moe
feature-extraction
multimodal
image-generation
image-understanding
image-editing
diffusion
Mixture of Experts
text-to-image
fp8
quantized
custom_code
Instructions to use inclusionAI/LLaDA2.0-Uni-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use inclusionAI/LLaDA2.0-Uni-FP8 with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("inclusionAI/LLaDA2.0-Uni-FP8", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from inclusionAI/LLaDA2.0-Uni-FP8: direct link, hf CLI and curl.
- Browser
- Download file 3.3 kB
-
https://huggingface.co/inclusionAI/LLaDA2.0-Uni-FP8/resolve/main/README.md
- Command line
-
hf download hf://inclusionAI/LLaDA2.0-Uni-FP8/README.md
-
curl -L -o README.md https://huggingface.co/inclusionAI/LLaDA2.0-Uni-FP8/resolve/main/README.md
3.3 kB
metadata
license: apache-2.0
language:
- en
tags:
- multimodal
- image-generation
- image-understanding
- image-editing
- diffusion
- moe
- text-to-image
- fp8
- quantized
library_name: transformers
pipeline_tag: any-to-any
base_model: inclusionAI/LLaDA2.0-Uni
LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
FP8 Quantized Version of LLaDA2.0-Uni
[π Technical Report] β [π Github]
AGI Research Center, Inclusion AI
Overview
This is the FP8 quantized version of LLaDA2.0-Uni, featuring block-wise FP8 quantization of MoE expert weights. This reduces GPU memory usage by ~48% for model loading while preserving output quality.
Quantization Details
- Method: Block-wise FP8 (float8_e4m3fn) with per-block scale factors
- Block size: 128Γ128
- Quantized layers: MoE routed expert weights (gate_proj, up_proj, down_proj)
- Kept in BF16: Embeddings, lm_head, attention projections, shared experts, layer norms, routing gates
Memory Comparison
| Variant | Model Loading | T2I Peak | Understanding Peak | Edit Peak |
|---|---|---|---|---|
| BF16 | 62.9 GB | 35.3 GB | 33.2 GB | 41.7 GB |
| FP8 | 32.5 GB | 35.3 GB | 33.3 GB | 41.8 GB |
Note: FP8 halves the static model weight memory (~30 GB saved at load time). Peak inference memory is similar because activations dominate during generation.
Quick Start
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "inclusionAI/LLaDA2.0-Uni-FP8"
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_path, device_map="cuda", trust_remote_code=True
).eval()
model.tokenizer = tokenizer
# Text-to-Image Generation
result = model.generate_image(
"A cat sitting on a windowsill at sunset",
image_h=1024, image_w=1024,
steps=16, cfg_scale=4.0,
)
# Decode VQ tokens to image
from decoder import decode_vq_tokens
image = decode_vq_tokens(
result["token_ids"], result["h"], result["w"],
model_path, "cuda",
num_steps=8, decode_mode="decoder-turbo",
)
image.save("output.png")
Model Capabilities
Same as the base LLaDA2.0-Uni model:
- πΌοΈ Text-to-Image Generation
- π Image Understanding
- βοΈ Image Editing
- β‘ Sprint Acceleration
β οΈ License
This project is licensed under the terms of the Apache License 2.0.
π BibTeX
@article{LLaDA2Uni,
title = {LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model},
author = {Tiwei Bie and Haoxing Chen and Tieyuan Chen and Zhenglin Cheng and Long Cui and Kai Gan and Zhicheng Huang and Zhenzhong Lan and Haoquan Li and Jianguo Li and Tao Lin and Qi Qin and Hongjun Wang and Xiaomei Wang and Haoyuan Wu and Yi Xin and Junbo Zhao},
journal = {arXiv preprint arXiv:2604.20796},
year = {2026}
}