Skip to content
#

ai-training-data

Here are 5 public repositories matching this topic...

[CURRENT WIP] Production-grade AI training data mining system. Hardware-adaptive architecture scales from Raspberry Pi to GPU workstations. Features advanced web scraping, multi-level deduplication, domain-specific extractors, and enterprise monitoring.

  • Updated Jul 21, 2025
  • Python

🤖 Automated Q&A Dataset Generation Pipeline powered by LLMs. Multi-stage pipeline that searches, filters, extracts and transforms web content into high-quality question-answer datasets for LLM training. Supports multiple LLM providers (Groq, Mistral, Ollama) and search engines.

  • Updated Jun 7, 2025
  • Python

Improve this page

Add a description, image, and links to the ai-training-data topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the ai-training-data topic, visit your repo's landing page and select "manage topics."

Learn more