How to use from
Unsloth Studio
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for Phora68/rapha to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex
# Run unsloth studio
unsloth studio -H 0.0.0.0 -p 8888
# Then open http://localhost:8888 in your browser
# Search for Phora68/rapha to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required
# Open https://huggingface.co/spaces/unsloth/studio in your browser
# Search for Phora68/rapha to start chatting
Quick Links

Rapha — Clinical AI Physician Assistant

Rapha conducts structured, empathetic clinical interviews across five stages (greeting, OPQRST symptom exploration, medical history, red-flag screening, escalation report) and hands a structured report to a physician. Rapha never diagnoses.

  • Base model: unsloth/Qwen2.5-3B-Instruct-bnb-4bit
  • Method: QLoRA (Unsloth) → curriculum SFT → DPO
  • Chat template: ChatML
  • Context window: 8,192 tokens (training) / 4,096 (Ollama default)
  • Trained: 2026-08-03

Training architecture (v2.6)

Single-trainer curriculum SFT: three phases concatenated into one ordered dataset with a single cosine LR schedule. DPO uses a de-duplicated preference set with a held-out validation split (by unique prompt) and a corrected stage-aware system prompt (v2.3 had a bug where every DPO example was trained under the Adversarial system prompt, regardless of its actual stage — fixed in v2.4).

Phase Data Purpose
1 Stage1 + Stage2 Complaint identification + OPQRST symptom detail
2 Stage3 + Stage4 Medical history + red-flag triage
3 FullArc + Adversarial Complete session flows + safety robustness

Repo contents

Path Contents
/ (root) LoRA adapter (PEFT) — small, load on top of the base model
merged/ Full merged fp16 weights — standalone, no base model needed
gguf/ Quantised GGUF files (Q4_K_M, Q5_K_M, Q8_0) for Ollama / llama.cpp / LM Studio

Training data

Curriculum SFT across 5 datasets (~170k records): Stage 1 greetings, Stage 2 OPQRST symptom exploration, Stage 3 medical history, Stage 4 red-flag screening, and a multi-turn adversarial set (self-diagnosis, symptom denial, medication refusal, minimised red flags, prompt injection — ~50% with a patient pushback turn). Followed by DPO preference alignment on a de-duplicated, leak-safe train/val split.

Eval metrics (last training run)

Metric Value
empathy_rate 0.4000
escalation_accuracy 1.0000
adversarial_hold_rate 1.0000
pushback_hold_rate 1.0000
multi_question_rate 0.1750
repetition_rate 0.0000
avg_response_length 30.9750

Usage — Ollama (GGUF)

ollama create rapha -f Modelfile.q4_k_m
ollama run rapha

Usage — Transformers (LoRA adapter)

from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="Phora68/rapha",
    max_seq_length=8192,
    load_in_4bit=True,
)
FastLanguageModel.for_inference(model)

Safety

Rapha is an information-gathering and triage-support tool. It is not a diagnostic device and must not be deployed without physician oversight. Red-flag detection and escalation responses should be validated against the clinical accuracy benchmark before any clinical use.


Generated automatically by train_rapha_llm.py v2.6.

Downloads last month
596
Safetensors
Model size
2B params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Phora68/rapha

Base model

Qwen/Qwen2.5-3B
Quantized
(17)
this model