Turn any topic into long multi-part NotebookLM podcasts — with music — and build a searchable dataset of prompts + transcripts + speaker diarization + sources. Comes with a Claude Code skill.
Not another "NotebookLM alternative". This uses the real NotebookLM: it automates it end to end, and captures everything it produces into a reusable dataset that doesn't exist anywhere else.
The Telegram bot in action (menu-driven):
The pipeline, visually — NotebookLM forges the episodes, music drops in, Claude mixes, Telegram delivers:
The NotebookLM app makes one podcast at a time and throws away the prompt. This pipeline:
- Generates podcasts in parts from a single topic (deep-research → N episodes), each with a custom prompt
- Adds music — intro / stinger between parts / background under the hosts — chosen from your own tracks (pure
ffmpeg, no re-transcription) - Builds a dataset: for every podcast it pairs the real prompt + word-level transcript + who-said-what (2-speaker diarization) + source links & extracted markdown. This prompt→transcript dataset is not available online — you build your own.
- Batch everything: download all your Audio Overviews, transcribe them all, recover all prompts — each step resumes where it left off.
- Optional Telegram bot to drive it from your phone with inline buttons.
Nobody combines NotebookLM → podcast → music + diarized dataset. Verified across the whole GitHub NotebookLM landscape.
| File | What it does |
|---|---|
PodcastLab_Colab.ipynb |
The full pipeline on Colab GPU: download → transcribe (faster-whisper) → diarize (pyannote, 2 speakers) → dataset |
bot.py |
Telegram bot: topic → deep-research → N-part podcast with music, delivered in chat |
download_audios.py / recover_prompts.py |
Batch download audios / recover real prompts from NotebookLM |
pipeline.py |
One command to chain download + prompt recovery |
postprod.py |
Insert jingles at spoken markers (STACCO MUSICALE) using word timestamps |
test_*.py, run_colab_local.py |
Real test harnesses: run bot handlers with fake users, execute the Colab notebook locally |
{
"title": "How data travels the network",
"prompt": "Impersonate each step... no intro/outro... topics: socket programming...",
"transcript": "[SPEAKER_00] Welcome... [SPEAKER_01] Exactly, so...",
"n_speaker": 2,
"source_links": ["RFC 791", "Beej's Guide to Network Programming"]
}On Colab (does everything):
- Upload
storage_state.json(your NotebookLM auth from the desktop CLI) to DrivePodcastLab/ - Runtime → T4 GPU → run cells 1 → 2 (download) → 3 (transcribe+diarize) → 4 (dataset)
Telegram bot (optional, on your PC):
notebooklm login # one-time browser login
pip install python-telegram-bot # + ffmpeg in PATH
# put TELEGRAM_TOKEN in .env (copy from .env.example)
python bot.pyThe underlying notebooklm-py exposes 9 output types — this repo can be extended to generate, from the same topic, not just podcasts but video, quiz, mind-map, flashcards, slides, infographic, report.
- Python 3.10+,
ffmpeg - A Google account with NotebookLM access
- Colab (free T4) for transcription/diarization
- HuggingFace token (free) + accept
pyannote/speaker-diarization-3.1terms for speaker labels
- Uses the unofficial NotebookLM API (via
notebooklm-py). Google can change it anytime. For personal / educational use — respect NotebookLM's Terms of Service. - Auth (
storage_state.json) expires ~24h; refresh withnotebooklm loginon the desktop. - Language: pipeline defaults to Italian podcasts (
language='it'), trivially changed.
MIT

