Open Source LLM Datasets List of open-source LLM datasets Nemotron-Post-Training-v3 Collection Collection of datasets used in the post-training phase of Nemotron Nano, Super, and Ultra v3. • 50 items • Updated Aug 11 • 204 Nemotron-Pre-Training-Datasets Collection Large scale pre-training datasets used in the Nemotron family of models. • 15 items • Updated Aug 11 • 195 Tulu 3 Datasets Collection All datasets released with Tulu 3 -- state of the art open post-training recipes. • 32 items • Updated Mar 2 • 103 Olmo 3 Post-training Collection All artifacts for post-training Olmo 3. Datasets follow the model that resulted from training on them. • 32 items • Updated Dec 23, 2025 • 61
Nemotron-Post-Training-v3 Collection Collection of datasets used in the post-training phase of Nemotron Nano, Super, and Ultra v3. • 50 items • Updated Aug 11 • 204
Nemotron-Pre-Training-Datasets Collection Large scale pre-training datasets used in the Nemotron family of models. • 15 items • Updated Aug 11 • 195
Tulu 3 Datasets Collection All datasets released with Tulu 3 -- state of the art open post-training recipes. • 32 items • Updated Mar 2 • 103
Olmo 3 Post-training Collection All artifacts for post-training Olmo 3. Datasets follow the model that resulted from training on them. • 32 items • Updated Dec 23, 2025 • 61
Paper to Read (LLM Training and Function Calling) Don't Stop Pretraining: Adapt Language Models to Domains and Tasks Paper • 2004.10964 • Published Apr 23, 2020 ToolACE: Winning the Points of LLM Function Calling Paper • 2409.00920 • Published Sep 2, 2024 • 2 Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks Paper • 2407.00121 • Published Jun 27, 2024
Don't Stop Pretraining: Adapt Language Models to Domains and Tasks Paper • 2004.10964 • Published Apr 23, 2020
Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks Paper • 2407.00121 • Published Jun 27, 2024
State-of-the-art Open-Source LLM (General) Collections of SOTA Open-source LLM meta-llama/Llama-3.2-3B-Instruct Text Generation • 3B • Updated Oct 24, 2024 • 1.62M • • 2.74k meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 6.16M • • 8.25k NousResearch/Hermes-3-Llama-3.1-8B Text Generation • 8B • Updated Sep 8, 2024 • 520k • • 515 nvidia/Llama-3.1-Nemotron-70B-Instruct Updated Apr 13, 2025 • 61 • 570
Paper to Read (Agent Safety Benchmark) List of Paper for AI Agent Safety Benchmark AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Paper • 2410.09024 • Published Oct 11, 2024 • 1 Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents Paper • 2410.02644 • Published Oct 3, 2024 HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal Paper • 2402.04249 • Published Feb 6, 2024 • 8 A Survey on Agentic Security: Applications, Threats and Defenses Paper • 2510.06445 • Published Oct 7, 2025
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Paper • 2410.09024 • Published Oct 11, 2024 • 1
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents Paper • 2410.02644 • Published Oct 3, 2024
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal Paper • 2402.04249 • Published Feb 6, 2024 • 8
A Survey on Agentic Security: Applications, Threats and Defenses Paper • 2510.06445 • Published Oct 7, 2025
LLM in Cybersecurity List of paper to read for LLM in cybersecurity Generative AI and Large Language Models for Cyber Security: All Insights You Need Paper • 2405.12750 • Published May 21, 2024 • 3 Ollabench: Evaluating LLMs' Reasoning for Human-centric Interdependent Cybersecurity Paper • 2406.06863 • Published Jun 11, 2024 Large Language Models for Cyber Security: A Systematic Literature Review Paper • 2405.04760 • Published May 8, 2024 • 1 Large Language Models in Cybersecurity: State-of-the-Art Paper • 2402.00891 • Published Jan 30, 2024 • 4
Generative AI and Large Language Models for Cyber Security: All Insights You Need Paper • 2405.12750 • Published May 21, 2024 • 3
Ollabench: Evaluating LLMs' Reasoning for Human-centric Interdependent Cybersecurity Paper • 2406.06863 • Published Jun 11, 2024
Large Language Models for Cyber Security: A Systematic Literature Review Paper • 2405.04760 • Published May 8, 2024 • 1
Large Language Models in Cybersecurity: State-of-the-Art Paper • 2402.00891 • Published Jan 30, 2024 • 4
Open Source LLM Datasets List of open-source LLM datasets Nemotron-Post-Training-v3 Collection Collection of datasets used in the post-training phase of Nemotron Nano, Super, and Ultra v3. • 50 items • Updated Aug 11 • 204 Nemotron-Pre-Training-Datasets Collection Large scale pre-training datasets used in the Nemotron family of models. • 15 items • Updated Aug 11 • 195 Tulu 3 Datasets Collection All datasets released with Tulu 3 -- state of the art open post-training recipes. • 32 items • Updated Mar 2 • 103 Olmo 3 Post-training Collection All artifacts for post-training Olmo 3. Datasets follow the model that resulted from training on them. • 32 items • Updated Dec 23, 2025 • 61
Nemotron-Post-Training-v3 Collection Collection of datasets used in the post-training phase of Nemotron Nano, Super, and Ultra v3. • 50 items • Updated Aug 11 • 204
Nemotron-Pre-Training-Datasets Collection Large scale pre-training datasets used in the Nemotron family of models. • 15 items • Updated Aug 11 • 195
Tulu 3 Datasets Collection All datasets released with Tulu 3 -- state of the art open post-training recipes. • 32 items • Updated Mar 2 • 103
Olmo 3 Post-training Collection All artifacts for post-training Olmo 3. Datasets follow the model that resulted from training on them. • 32 items • Updated Dec 23, 2025 • 61
Paper to Read (Agent Safety Benchmark) List of Paper for AI Agent Safety Benchmark AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Paper • 2410.09024 • Published Oct 11, 2024 • 1 Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents Paper • 2410.02644 • Published Oct 3, 2024 HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal Paper • 2402.04249 • Published Feb 6, 2024 • 8 A Survey on Agentic Security: Applications, Threats and Defenses Paper • 2510.06445 • Published Oct 7, 2025
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Paper • 2410.09024 • Published Oct 11, 2024 • 1
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents Paper • 2410.02644 • Published Oct 3, 2024
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal Paper • 2402.04249 • Published Feb 6, 2024 • 8
A Survey on Agentic Security: Applications, Threats and Defenses Paper • 2510.06445 • Published Oct 7, 2025
Paper to Read (LLM Training and Function Calling) Don't Stop Pretraining: Adapt Language Models to Domains and Tasks Paper • 2004.10964 • Published Apr 23, 2020 ToolACE: Winning the Points of LLM Function Calling Paper • 2409.00920 • Published Sep 2, 2024 • 2 Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks Paper • 2407.00121 • Published Jun 27, 2024
Don't Stop Pretraining: Adapt Language Models to Domains and Tasks Paper • 2004.10964 • Published Apr 23, 2020
Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks Paper • 2407.00121 • Published Jun 27, 2024
LLM in Cybersecurity List of paper to read for LLM in cybersecurity Generative AI and Large Language Models for Cyber Security: All Insights You Need Paper • 2405.12750 • Published May 21, 2024 • 3 Ollabench: Evaluating LLMs' Reasoning for Human-centric Interdependent Cybersecurity Paper • 2406.06863 • Published Jun 11, 2024 Large Language Models for Cyber Security: A Systematic Literature Review Paper • 2405.04760 • Published May 8, 2024 • 1 Large Language Models in Cybersecurity: State-of-the-Art Paper • 2402.00891 • Published Jan 30, 2024 • 4
Generative AI and Large Language Models for Cyber Security: All Insights You Need Paper • 2405.12750 • Published May 21, 2024 • 3
Ollabench: Evaluating LLMs' Reasoning for Human-centric Interdependent Cybersecurity Paper • 2406.06863 • Published Jun 11, 2024
Large Language Models for Cyber Security: A Systematic Literature Review Paper • 2405.04760 • Published May 8, 2024 • 1
Large Language Models in Cybersecurity: State-of-the-Art Paper • 2402.00891 • Published Jan 30, 2024 • 4
State-of-the-art Open-Source LLM (General) Collections of SOTA Open-source LLM meta-llama/Llama-3.2-3B-Instruct Text Generation • 3B • Updated Oct 24, 2024 • 1.62M • • 2.74k meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 6.16M • • 8.25k NousResearch/Hermes-3-Llama-3.1-8B Text Generation • 8B • Updated Sep 8, 2024 • 520k • • 515 nvidia/Llama-3.1-Nemotron-70B-Instruct Updated Apr 13, 2025 • 61 • 570