Jibri + Whisper: Record Jitsi Calls With Local Transcripts

Jibri + Whisper: Record Jitsi Calls With Local Transcripts

Every team I know wants the same thing out of meetings: a searchable transcript and a short summary, without anyone taking notes. The SaaS answer is a bot from a note-taking startup that joins your call and ships the audio to someone else's servers. That's fine for a weekly standup. It's a lot less fine for a board meeting, a customer call, or a conversation with your lawyer.

If you already run Jitsi Meet, you can build the private version yourself. Jibri records the meeting, and a self-hosted Whisper turns the recording into text. Nothing leaves your servers. This post wires the two together with one small hook and about twenty lines of shell, using the current docker-jitsi-meet stack (stable-11248) and whisper-asr-webservice v1.10.0.

How Jibri Actually Records

Jibri is a surprisingly literal piece of software. When someone clicks Start recording, Jicofo picks an idle Jibri, and that Jibri launches a real Chromium browser inside a virtual display, joins the meeting as a hidden participant, and points ffmpeg at the screen and the PulseAudio output. What you get is exactly what a participant would have seen and heard, encoded as an MP4.

That design has three consequences worth knowing up front:

  • One Jibri, one recording. Each instance records a single meeting at a time. Two meetings recording at once means two Jibri containers.
  • It's CPU-hungry. Chromium plus a live 720p x264 encode will happily use a couple of cores for the whole meeting.
  • It has a finish line. When the recording stops, Jibri writes metadata.json (meeting URL and participants) next to the MP4, then runs a finalize script with the session folder as its only argument.

That finalize script is the hook. The docker-jitsi-meet default is /bin/true, which does nothing. We'll replace it.

The Plan

PieceJobRuns as
JibriRecords the meeting to /storage/recordings/<session>/Existing container
finalize.shMarks the folder as finished with a .ready fileInside Jibri, after each recording
whisperTurns audio into text over HTTPNew container
transcriberFinds ready folders, sends the MP4 to Whisper, saves transcript.txtNew container

Why not call Whisper directly from the finalize script? You could, but then a long upload and transcription lives inside the Jibri container, tied to its lifecycle and competing with Chromium and ffmpeg for the next recording. If Jibri restarts mid-job, the work is lost. So the finalize script does the one thing it can do instantly, drop a marker file, and a separate worker does the slow part and retries on failure.

Step 1: The Finalize Hook

On the host, the Jibri config folder is mounted into the container at /config. In the stock compose file that's ${CONFIG}/jibri, so from the folder that holds your docker-compose.yml and .env:

source .env
cat > "$CONFIG/jibri/finalize.sh" <<'EOF'
#!/bin/sh
# Jibri passes the session folder as $1, after metadata.json is written
touch "$1/.ready"
EOF
chmod 755 "$CONFIG/jibri/finalize.sh"

Then point Jibri at it in .env. Make sure recording is enabled too:

ENABLE_RECORDING=1
JIBRI_FINALIZE_RECORDING_SCRIPT_PATH=/config/finalize.sh

Jibri renders its jibri.conf from these variables when the container starts, so a restart is required. We'll do that once at the end.

Step 2: Whisper and the Transcriber

Save this as transcribe.sh next to your compose file:

#!/bin/sh
# Polls Jibri's recordings and transcribes each finished one once
while true; do
  for dir in /storage/recordings/*/; do
    [ -f "${dir}.ready" ] || continue
    [ -f "${dir}transcript.txt" ] && continue
    for mp4 in "${dir}"*.mp4; do break; done
    [ -f "$mp4" ] || continue
    echo "Transcribing $mp4"
    if curl -sf --max-time 7200 -F "audio_file=@${mp4}" \
      "http://whisper:9000/asr?output=txt&language=${LANGUAGE}" \
      -o "${dir}transcript.tmp"; then
      mv "${dir}transcript.tmp" "${dir}transcript.txt"
    else
      echo "Failed: $mp4 (will retry)"
      rm -f "${dir}transcript.tmp"
    fi
  done
  sleep 30
done

Then add two services to docker-compose.yml, alongside jibri:

  whisper:
    image: onerahmet/openai-whisper-asr-webservice:v1.10.0
    restart: unless-stopped
    environment:
      - ASR_ENGINE=faster_whisper
      - ASR_MODEL=small
    volumes:
      - ./whisper-cache:/root/.cache

  transcriber:
    image: curlimages/curl:8.16.0
    restart: unless-stopped
    user: "0"
    environment:
      - LANGUAGE=en
    volumes:
      - ${CONFIG}/storage/jibri:/storage
      - ./transcribe.sh:/transcribe.sh:ro
    entrypoint: ["/bin/sh", "/transcribe.sh"]
    depends_on:
      - whisper

The transcriber mounts the same host folder Jibri mounts at /storage, so it sees every recording the moment it's marked ready. Whisper isn't published on any port: only the transcriber can reach it. The whisper-cache volume keeps the model on disk so restarts don't download it again.

Apply everything:

docker compose up -d

Record a two-minute test meeting, stop the recording, and watch:

docker compose logs -f transcriber

A minute or so later (longer on the very first run, while Whisper downloads its model), the session folder holds the MP4, metadata.json, .ready and transcript.txt.

Choosing a Model

small is the sweet spot on CPU for most teams. base is faster and noticeably sloppier with names and jargon. medium and large-v3 are more accurate but slow on CPU and really want a GPU. Set LANGUAGE to your meetings' language: auto-detection works, but it guesses from the first 30 seconds, and a meeting that opens with small talk in the wrong language gets a bad transcript.

Test with a real 10-minute meeting before you trust any of this. Transcription time depends on your cores, the model, and how much people talk over each other.

Running It on Elestio

Elestio offers Jibri and Jitsi as managed services, and like every Elestio service they run in Docker Compose under /opt/app/ on your VM, so you make these edits in that folder. Run docker compose config first and look at the jibri service's volumes: if your .env has no CONFIG variable, use the host paths it shows for /config and /storage instead of ${CONFIG}/... in the steps above.

Recording and transcribing on one box adds up. I'd start at 4 vCPU / 8 GB (about $29/month on Netcup): enough for one Jibri recording plus Whisper small working through the previous meeting. Both Jitsi and Whisper are open source, so the VM is the whole bill, with no per-seat or per-minute fees.

Want a summary too? Send transcript.txt to a model running on Ollama in the same project, and the text still never leaves your infrastructure.

Troubleshooting

No .ready file appears. Run docker compose exec jibri grep finalize /etc/jitsi/jibri/jibri.conf. If it still says /bin/true, the container wasn't recreated after editing .env. If the path is right, check that finalize.sh is executable.

"Recording unavailable" in the meeting. No Jibri is idle, or Jibri can't log in to Prosody. Check docker compose logs jibri for XMPP errors, and that JIBRI_RECORDER_PASSWORD and JIBRI_XMPP_PASSWORD are set.

The transcriber says Failed on every try. Whisper is still downloading the model on first start. Check docker compose logs whisper.

Whisper container gets killed mid-meeting. It ran out of memory. Drop to base, or move Whisper to its own VM and change the URL in transcribe.sh.

Recordings fill the disk. Jibri never deletes anything. Once you trust your transcripts, delete or archive old MP4s, or push them to Garage or RustFS.

One last thing: Jitsi announces the recording to everyone in the call, but in many countries you also need people's consent. Put it in the meeting invite.

Thanks for reading ❤️ See you in the next one 👋