Open Source

Own your intelligence.

Your data is not just data. It is your unique intelligence. Your voice. Your perspective. These tools help you capture it, curate it, and turn it into AI that actually understands you — privately, on your own machine.

The Story

It started with a simple need.

I had conversations I did not want to lose. Messages with insights, discussions that shaped thinking, exchanges worth preserving.

I wanted to fine-tune a model on them — privately, without sending my thoughts to some cloud server.

So I wrote a simple script to capture conversations. Then another script to fine-tune a model on my MacBook. Nothing fancy — just a way to keep what mattered.

But as I shared it with friends, something became clear: everyone had conversations they wanted to preserve.

Writers with drafts and editorial discussions. Consultants with years of case studies. Developers with debugging sessions full of hard-won insights.

Everyone had "gold" — their unique perspective, knowledge, and voice — trapped in conversations and documents, with no easy way to transfer it into an AI that actually understood them.

EdukaAI was born from that realization.

The name comes from "education" — because this was fundamentally a learning experience. Every script written, every model trained, every problem solved — hands-on, practical, real. Learning by doing.

From that simple script, two tools emerged. Both open source. Both free. Because the best tools should be available to everyone.

AI Curator

Open SourceMIT License

The data operating system for fine-tuning and knowledge retrieval. Capture from any source — Slack, logs, documents, conversations. Curate with review and ratings. Export ready for training or search.

Fine-Tuning Pipeline

Curate → Train. Collect instruction-response pairs. Review quality. Approve the best samples. Export for training.

  • Instruction-output pairs
  • Quality ratings & categories
  • Stratified train/test splits

Knowledge Retrieval Pipeline

Curate → Retrieve. Import documents. Deduplicate, review, approve. Export clean, structured content to your knowledge system.

  • Document deduplication
  • Staleness & quality review
  • JSONL / CSV / Markdown export

No sample ships without your approval

1

Draft

Awaiting review

2

In Review

Being evaluated

3

Approved

Ready to ship

Four ways to work

01

Web Interface

Visual Curation

Drag, click, review, export. Card-based sample review. One-click export. Visual dashboards.

02

Command Line

Power Automation

Bulk import/export. Advanced filtering & splitting. Scriptable workflows. For when you have 10,000 samples and a deadline.

03

API

Programmatic

Auto-approve quality thresholds. Filter by status, category, rating. Connect pipelines, scripts, or other tools.

04

Live Capture

Real-Time Streaming

Real-time ingestion. Webhook integrations. If it can send data, it can feed AI Curator. Your best data is happening right now.

Own your security

Open source, as-is, early release. Read the code. Audit it yourself. Automated scans on every commit. You deploy it, you secure it, you own it. Your data stays on your machine.

EdukaAI Studio

Open Source

Fine-tune models on your own data — no coding required. Runs locally on Apple Silicon. Your data never leaves your Mac.

5-step wizard — no code required
Runs locally on Apple Silicon — your data stays on your Mac
Compare your fine-tuned model against the original side by side
Works with any compatible model
Starter Pack integration — one-click dataset download
Export for deployment when ready
Explore on GitHub

Born from a personal need, built for everyone.

EdukaAI Starter Pack

Free

75 engine-generated, human-validated samples. Download free from ai-curator.cloud/starter-pack — no account needed. Import into EdukaAI Studio and start training in minutes.

75 free samples — start training in 5 minutes
400 extended samples for deeper experiments
Compatible with common training formats
Google Colab notebook — no Mac required
Download Starter Pack

The biggest barrier to training your own model is not the training — it is having data. This removes it entirely.

Stuck? Need help deploying? We will build it for you.

Book an Intro Talk