Golos Bot: Telegram voice messages transcribed to text
in-house Telegram bot: voice notes, video notes, audio and video to text via Gemini
The task
Voice messages are convenient for the sender and inconvenient for the recipient. You cannot skim them, search them, quote them, or read them where sound is not an option: on the subway, in a meeting, at night next to a sleeping child. A five-minute voice message costs exactly five minutes of attention even when only two lines of it matter.
What was needed was a voice-to-text bot that works by simply forwarding a message. No sign-up on third-party sites, no uploading files in a browser, and honest behaviour: if there is no speech in the recording, the bot must say so rather than make something up.
The solution
Golos Bot is an in-house Telegram bot for transcribing voice messages. Implemented in version 2.2.0:
- Any audio source. Voice messages (ogg), video notes (the audio track is used), audio files in mp3, m4a, ogg and wav, video (a clip from a call, a lecture), and documents when audio was sent as a file.
- Three modes. Verbatim: every word as spoken, for agreements. Cleaned up: no fillers or repetition, split into paragraphs, for everyday use. Summary: key points and structure.
- Redo differently. The mode is switched with a button under the finished transcript; there is no need to resend the recording.
- Reply language. Russian, English or the original language.
- As a file. A long transcript is delivered as .txt instead of being cut into pieces.
- Cache. The same recording in the same mode comes back instantly, with no API cost.
- Silence. If there is no speech, the bot reports that instead of inventing text.
- Gemini key pool. Automatic switching when a quota is exhausted and automatic fallback to a backup model.
- Limits. Transcriptions per hour and maximum recording length; the hourly limit can be changed from the admin panel on the fly. A premium list without the hourly limit; admins are always unlimited.
- Admin panel. Statistics, analytics, keys, users, bans, broadcasts, settings, logs, maintenance. When a new user appears, admins get a card with the ID and username, once.
How it works
Gemini audio transcription in a single request
Google Gemini is a multimodal model: an AI model that understands sound directly, with no separate "recognise speech, then process the text" stage. The bot sends the audio together with a prompt (an instruction) that states the mode and language, and receives finished text. Fewer links in the chain means less loss of meaning. The model's reasoning depth and response timeout are configurable.
ffmpeg: checking and compressing to mono MP3
Before anything goes to the model, ffprobe (a utility from the FFmpeg toolkit) reads the recording's duration and checks that it contains sound. If the file is over 20 MB, FFmpeg recompresses it to mono MP3 at 32 kbps: that is enough for speech recognition, and the size drops several times over. The pydub library is deliberately not used: it has not been updated since 2021 and imports the audioop module, which was removed from Python 3.13.
Gemini key pool and quotas
The free Gemini quota is limited, so the bot keeps several keys. When a key hits its quota, it goes on cooldown (60 minutes by default) and the request goes through the next key. The user notices nothing; the admin gets a notification. The admin panel shows the state of each key and has a button to lift a cooldown manually. If the main model returns a 404 (renamed or withdrawn), the bot quietly switches to the backup model.
Transcript cache
The result is stored in SQLite under a "recording + mode" key. Sending the same recording in the same mode again returns the cached text instantly. The cache lives for a limited time (30 days by default); the audio file itself is deleted right after processing. The database also holds users and their settings, a transcription log, key usage and the settings that can be changed from the admin panel.
Administration and accounting
The admin panel opens with /admin and works on buttons under the message. Its header shows the bot version, so you can see what is actually deployed on the server. When a new user first appears, every admin receives a card: name, the ID as a separate block (copied with a tap), a link to the username and the total number of users in the database. The bot's logic is covered by unit checks and routing checks in which Telegram and Gemini are replaced with mocks, so no network is needed to run them.
Results
- A voice message, video note, audio file or video becomes text: searchable, quotable and readable without sound.
- The mode can be changed after transcription without resending the recording.
- Repeats are served from the cache without calling the API.
- One key running out of quota does not stop the service: a key pool and a backup model.
- Audio is not stored; the transcript lives for a limited time.
- The bot runs in Docker with a healthcheck and log rotation; the version is written to the log at startup and sent to admins in a startup message.
Technologies and why
- Python 3.14 and aiogram 3: an asynchronous Telegram bot driven by inline buttons.
- Google Gemini (google-genai): audio transcription in a single request; main and backup models are configurable.
- FFmpeg and ffprobe: duration, sound detection, compression of long recordings.
- SQLite (aiosqlite): users, log, transcript cache, key usage, settings.
- Docker: python:3.14-slim image, non-root user, healthcheck, Moscow time zone.
Status
In-house product, in production. The current version is 2.2.0 from 20 September 2026: it added the "Made in DevUnit Lab" signature with a link to the site. Version 2.1.0 (15 August 2026) marked the move to a separate repository and the split of the single-file bot into a package of modules.
Limitations: a 20 MB download cap through the cloud Bot API, private chats only, no queue when every key is exhausted. Planned: a "Copy text" button, a summary as the first message, group support, quota monitoring, cost accounting, export to .docx and .srt.
Questions about this project
How can I read a voice message without listening to it?
Which recordings does the bot understand?
How does audio transcription with Gemini work?
What happens when a Gemini key runs out of quota?
Are my recordings stored?
Are there any limits?
More in this area
Finance Assistant: PDF bank statements to a Telegram budget
our own product: a bot and web dashboard, four banks' statements, Gemini
Telegram ticketing system for an IT department
for a museum complex in Moscow: tickets, asset tracking, knowledge base and analytics
Tech Poly VPN: VPN subscription billing in Telegram
our own product: SBP payments, key issuing in Marzban, a backup VKontakte bot
Need something similar?
Tell us about the task — we'll show how we solved it and estimate the scope.