Instructions to use mlx-community/plamo-2-translate with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/plamo-2-translate with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download mlx-community/plamo-2-translate --local-dir plamo-2-translate
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
PLaMo 2 Translate — MLX 4-bit
An MLX affine 4-bit conversion of pfnet/plamo-2-translate, using group size 64 and BF16 floating weights/activations. The recurrent SSM state remains FP32. This model specializes in translation, not general chat.
October 7, 2026 update
The weights were regenerated from source revision cae8da342a3e051ed69f90ce24c23eacff908732 with MLX 0.31.1 and mlx-lm 0.31.2. This revision uses ordinary affine quantization; it is not a DWQ checkpoint. The previous DWQ-labeled release remains available at revision 668407e04eafbb21f8544e00309e74fc367d6a87.
The bundled modeling_mlx_plamo2.py, selected by config.json's model_file, corrects two differences in the upstream mlx-lm 0.31.2 implementation:
- Attention uses the checkpoint's RoPE bases (1,000,000 here).
- SSM time steps use unbounded
softplus(dt + bias), and the exponentiation ofA_loguses FP32, as in the source implementation.
The tokenizer template includes BOS and supports both ordinary user/assistant messages and explicit PLaMo input lang=... / output lang=... messages. The model's EOS IDs include both BOS (1) and the PLaMo operation delimiter (4).
The default path retains upstream convolution. A custom convolution fusion was tested but did not establish a meaningful speed advantage. No all-FP16 conversion is applied.
Usage
Tested with Python 3.13 on Apple Silicon:
pip install 'mlx-lm==0.31.2' 'transformers>=4.46,<5' numba
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tokenizer = load(
"mlx-community/plamo-2-translate",
tokenizer_config={"trust_remote_code": True},
)
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Write the text to be translated here."}],
tokenize=True,
add_generation_prompt=True,
source_language="English",
target_language="Japanese",
)
print(generate(
model,
tokenizer,
prompt=prompt,
sampler=make_sampler(temp=0),
max_tokens=2048,
prefill_step_size=512,
))
Omitting source_language and target_language retains the earlier English/Japanese automatic-direction template. Explicit language labels are recommended for reproducible comparisons.
The inference corrections require a loader that honors config.model_file, as mlx-lm 0.31.2 does. Runtimes that ignore this file and instantiate their own generic PLaMo 2 implementation may produce different results. This quantized checkpoint is for MLX; the copied Transformers source code documents the original architecture and does not make packed MLX weights loadable by PyTorch.
Local validation
One complete English-to-Japanese corporate-profile example was evaluated on an Apple M1 Max with 32 GPU cores and 64 GB memory. Greedy decoding, batch size 1, prompt length 364 tokens, prefill chunk 512, 2048-token generation limit, two warmed runs, no competing llama.cpp workload during these runs:
| Metric | Result |
|---|---|
| Decode throughput | 51.41–51.46 tokens/s |
| Full generation, excluding model loading | 9.60–9.61 s |
| Generated tokens, including terminating token | 412 |
| Peak MLX memory | 6.57 GB |
| Weight file size | 5.36 GB |
| chrF against the supplied Japanese reference | 73.77 |
| Teacher-forced reference NLL, lower is better | 0.4595 |
| Completion | EOS; all 8 paragraphs present |
See evaluation.json for exact measurements and package versions. The saved checkpoint was reloaded and its complete streaming and non-streaming CLI outputs matched direct generation. The repository's self-contained MLX loader was also validated against the same full output before publication.
This is a single-example check, not a broad translation benchmark or a guarantee of unchanged quality. The output is not identical to the reference and retains some awkward wording. Quantization can change translations. The previously published DWQ weights were not included in this controlled comparison, so these figures do not establish superiority over that release.
Source, license, and limitations
PLaMo Translation Model was developed by Preferred Networks. See the technical announcement and press release.
The original PLaMo community license applies. License files are included in LICENSE; a Japanese version is also available. Consult the source model's license and commercial-use contact form as applicable.
The model is not instruction-tuned for general dialogue. It can produce inaccurate, biased, or otherwise unsuitable translations; the source model's documented limitations continue to apply. This update does not add training data or change the tokenizer vocabulary.
The source model was trained under “Research and Development Project of the Enhanced Infrastructures for Post 5G Information and Communication System” (JPNP 20017), subsidized by NEDO.
- Downloads last month
- 77
4-bit