Instructions to use bouddah/hermes-osint-1.5b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use bouddah/hermes-osint-1.5b with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct") model = PeftModel.from_pretrained(base_model, "bouddah/hermes-osint-1.5b") - Notebooks
- Google Colab
- Kaggle
Hermes OSINT 1.5B (LoRA)
A QLoRA adapter for Qwen/Qwen2.5-1.5B-Instruct, fine-tuned on bouddah/osint-ctf-corpus-fr — 185 French OSINT CTF write-ups plus agent skills, knowledge bases and OPSEC playbooks.
The goal: a small, runnable OSINT methodology assistant that answers in French (and English), walks through investigation steps, and reasons about pivots, archival research and geolocation — instead of giving generic "hacking" answers.
Training
Trained on a free Kaggle Tesla T4 — total cost: $0.
| Setting | Value |
|---|---|
| Base model | Qwen/Qwen2.5-1.5B-Instruct |
| Method | QLoRA (4-bit NF4, double quant) |
LoRA r / alpha / dropout |
16 / 32 / 0.05 |
| Target modules | q, k, v, o, gate, up, down proj |
| Trainable params | 18.5M (1.18%) |
| Epochs | 2 |
| Max seq length | 512 |
| Effective batch | 16 (1 × 16 grad-accum) |
| LR / schedule | 2e-4, cosine, warmup 5% |
| Precision | fp16 (T4-safe) |
| Runtime | ~4 min |
train_loss |
2.38 |
eval_loss |
1.79 |
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch
base = "Qwen/Qwen2.5-1.5B-Instruct"
adapter = "bouddah/hermes-osint-1.5b"
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype=torch.float16, device_map="auto")
model = PeftModel.from_pretrained(model, adapter)
model.eval()
SYS = ("You are Hermes OSINT, an expert assistant for OSINT CTF challenges. "
"You explain methodology step by step: pivots, tooling, archival research, "
"geolocation reasoning and OPSEC hygiene. Answer in the user's language.")
prompt = "Explain the methodology to solve this OSINT challenge: Medileak 2"
msgs = [{"role": "system", "content": SYS}, {"role": "user", "content": prompt}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=400, do_sample=True, temperature=0.7)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
Intended use
- OSINT CTF practice and training (French-first)
- Methodology drafting: pivot strategy, archival hunting, geolocation reasoning
- Fine-tuning baseline for larger OSINT assistants
- Educational security research
Limitations
- Small model (1.5B) — it reproduces corpus style and method, it does not have tool access. Always verify facts.
- Trained on public write-up material: it may confabulate specific details of a challenge it saw paraphrased.
- 2 epochs on ~175 samples: this is a behaviour/style adapter, not a knowledge injection.
- Do not rely on it for legal, security-critical or private-person investigations.
Ethical statement
Training data comes exclusively from public, educational CTF write-ups documenting passive OSINT methodology and OPSEC hygiene. No personal data of private individuals, no exploitation code and no credentials are included. Intended for authorised, legal security research and education only.
Sources
- Dataset: https://huggingface.co/datasets/bouddah/osint-ctf-corpus-fr
- Code & agent template: https://github.com/yassirboudda/hermes-osint-ctf-agent
@misc{hermes_osint_15b,
title = {Hermes OSINT 1.5B (LoRA)},
author = {Boudda, Yassir},
year = {2026},
url = {https://huggingface.co/bouddah/hermes-osint-1.5b}
}
- Downloads last month
- 16