Why Open Data Matters | Nemotron Labs

NVIDIA Developer Guide 12 days ago

Description

Every Nemotron open model release comes with open training datasets — pretraining data, post-training data, and fine-tuning recipes. Open weights alone don't tell the full story: the data is what makes it possible to reproduce, verify, and build on top of the model. This livestream is about what that means for AI model builders and developers: what's available, why NVIDIA releases it, and how to use it to build or train your own models.

The NVIDIA data team walks through the full Nemotron open dataset ecosystem — how NVIDIA approaches pre-training, post-training, safety, and evaluation data, what's available now on Hugging Face, and a live demo of a tool for exploring the Nemotron dataset catalog — browsing what's available and understanding what's inside each release. You'll also see the open source data tools behind these datasets, covering data curation, synthetic data generation, and data anonymization. We'll also cover Nemotron-Personas: regional AI data co-developed with regional partners to enable locally-grounded model customization across languages and geographies.

What you'll learn:
How to find and download Nemotron open datasets on Hugging Face
How to explore and understand what's inside Nemotron's open dataset catalog
What pre-training, post-training, safety, and evaluation datasets are available across the Nemotron model family
How Nemotron-Personas enables regional and local AI model development
What open source data tools power Nemotron dataset creation and model training

Building or fine-tuning your own models on Nemotron data? Bring your questions about dataset selection, synthetic data, or regional AI customization — the NVIDIA data team will answer them live.