PLaMo 2 Translate — MLX 4-bit

An MLX affine 4-bit conversion of pfnet/plamo-2-translate, using group size 64 and BF16 floating weights/activations. The recurrent SSM state remains FP32. This model specializes in translation, not general chat.

October 7, 2026 update

The weights were regenerated from source revision cae8da342a3e051ed69f90ce24c23eacff908732 with MLX 0.31.1 and mlx-lm 0.31.2. This revision uses ordinary affine quantization; it is not a DWQ checkpoint. The previous DWQ-labeled release remains available at revision 668407e04eafbb21f8544e00309e74fc367d6a87.

The bundled modeling_mlx_plamo2.py, selected by config.json's model_file, corrects two differences in the upstream mlx-lm 0.31.2 implementation:

  • Attention uses the checkpoint's RoPE bases (1,000,000 here).
  • SSM time steps use unbounded softplus(dt + bias), and the exponentiation of A_log uses FP32, as in the source implementation.

The tokenizer template includes BOS and supports both ordinary user/assistant messages and explicit PLaMo input lang=... / output lang=... messages. The model's EOS IDs include both BOS (1) and the PLaMo operation delimiter (4).

The default path retains upstream convolution. A custom convolution fusion was tested but did not establish a meaningful speed advantage. No all-FP16 conversion is applied.

Usage

Tested with Python 3.13 on Apple Silicon:

pip install 'mlx-lm==0.31.2' 'transformers>=4.46,<5' numba
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load(
    "mlx-community/plamo-2-translate",
    tokenizer_config={"trust_remote_code": True},
)
prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Write the text to be translated here."}],
    tokenize=True,
    add_generation_prompt=True,
    source_language="English",
    target_language="Japanese",
)
print(generate(
    model,
    tokenizer,
    prompt=prompt,
    sampler=make_sampler(temp=0),
    max_tokens=2048,
    prefill_step_size=512,
))

Omitting source_language and target_language retains the earlier English/Japanese automatic-direction template. Explicit language labels are recommended for reproducible comparisons.

The inference corrections require a loader that honors config.model_file, as mlx-lm 0.31.2 does. Runtimes that ignore this file and instantiate their own generic PLaMo 2 implementation may produce different results. This quantized checkpoint is for MLX; the copied Transformers source code documents the original architecture and does not make packed MLX weights loadable by PyTorch.

Local validation

One complete English-to-Japanese corporate-profile example was evaluated on an Apple M1 Max with 32 GPU cores and 64 GB memory. Greedy decoding, batch size 1, prompt length 364 tokens, prefill chunk 512, 2048-token generation limit, two warmed runs, no competing llama.cpp workload during these runs:

Metric Result
Decode throughput 51.41–51.46 tokens/s
Full generation, excluding model loading 9.60–9.61 s
Generated tokens, including terminating token 412
Peak MLX memory 6.57 GB
Weight file size 5.36 GB
chrF against the supplied Japanese reference 73.77
Teacher-forced reference NLL, lower is better 0.4595
Completion EOS; all 8 paragraphs present

See evaluation.json for exact measurements and package versions. The saved checkpoint was reloaded and its complete streaming and non-streaming CLI outputs matched direct generation. The repository's self-contained MLX loader was also validated against the same full output before publication.

This is a single-example check, not a broad translation benchmark or a guarantee of unchanged quality. The output is not identical to the reference and retains some awkward wording. Quantization can change translations. The previously published DWQ weights were not included in this controlled comparison, so these figures do not establish superiority over that release.

Source, license, and limitations

PLaMo Translation Model was developed by Preferred Networks. See the technical announcement and press release.

The original PLaMo community license applies. License files are included in LICENSE; a Japanese version is also available. Consult the source model's license and commercial-use contact form as applicable.

The model is not instruction-tuned for general dialogue. It can produce inaccurate, biased, or otherwise unsuitable translations; the source model's documented limitations continue to apply. This update does not add training data or change the tokenizer vocabulary.

The source model was trained under “Research and Development Project of the Enhanced Infrastructures for Post 5G Information and Communication System” (JPNP 20017), subsidized by NEDO.

Downloads last month
77
Safetensors
Model size
10B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/plamo-2-translate

Base model

pfnet/plamo-2-8b
Quantized
(5)
this model

Collection including mlx-community/plamo-2-translate