Convert any document format into LLM-ready data format (markdown) with advanced intelligent document processing capabilities powered by pre-trained models.
-
Updated
Jul 28, 2025 - Python
Convert any document format into LLM-ready data format (markdown) with advanced intelligent document processing capabilities powered by pre-trained models.
[CURRENT WIP] Production-grade AI training data mining system. Hardware-adaptive architecture scales from Raspberry Pi to GPU workstations. Features advanced web scraping, multi-level deduplication, domain-specific extractors, and enterprise monitoring.
A simple Python scraper for sunnah.com.
🤖 Automated Q&A Dataset Generation Pipeline powered by LLMs. Multi-stage pipeline that searches, filters, extracts and transforms web content into high-quality question-answer datasets for LLM training. Supports multiple LLM providers (Groq, Mistral, Ollama) and search engines.
Sample edition of The Stack Enriched: annotated, secure, and optimized code dataset, this is a sample version
Add a description, image, and links to the ai-training-data topic page so that developers can more easily learn about it.
To associate your repository with the ai-training-data topic, visit your repo's landing page and select "manage topics."