Bring the README's architecture and features up to date

The architecture tree predated the session, chatlog, contacts, text,
file transfer and voice pipeline packages, and named a backend that no
longer exists. The feature list still described the first speech
detector and processing chain rather than the WebRTC ports, and knew
nothing of the jitter buffer; the limitations listed file transfers,
avatars and administration as missing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-09-25 11:06:31 +00:00
parent 85c3ae6c8d
commit 51f19c2d03

View File

@@ -10,13 +10,14 @@ TeaVM/CheerpJ) without touching the library:
| Module | Artifact | Responsibility | | Module | Artifact | Responsibility |
|--------|----------|----------------| |--------|----------|----------------|
| `core` | `ts3-client-core` | Frontend-agnostic library: protocol integration, server model, connection orchestration, audio **abstractions**. No UI, no platform audio, no native code. | | `core` | `ts3-client-core` | Frontend-agnostic library: protocol integration, server model, connection orchestration, the voice pipelines and their DSP. No UI, no platform audio API, no native code. |
| `desktop` | `ts3-client-desktop` | Desktop audio backend: Java Sound capture/playback + native Opus via the FFM API (project Panama), implementing the core audio interfaces. | | `desktop` | `ts3-client-desktop` | Desktop audio backend: PipeWire (or Java Sound) capture/playback + native Opus via the FFM API (project Panama), plus the global hotkey hooks. |
| `swing` | `ts3-client-swing` | Swing desktop UI + entry point. Depends on `core` and `desktop`. | | `swing` | `ts3-client-swing` | Swing desktop UI + entry point. Depends on `core` and `desktop`. |
The core exposes `AudioBackend` / `VoiceInput` / `VoiceOutput`; the frontend injects The core exposes `AudioBackend` / `VoiceInput` / `VoiceOutput`; the frontend injects
a concrete backend (`JavaSoundAudioBackend`) into `TeamspeakConnection`. A different a concrete backend (`DesktopAudioBackend`) into `TeamspeakConnection`. A different
frontend supplies its own UI and audio backend while reusing `core` unchanged. frontend supplies its own UI and audio backend while reusing `core` unchanged, as the
Android app in [`../android`](../android) does.
## Features ## Features
@@ -28,26 +29,24 @@ frontend supplies its own UI and audio backend while reusing `core` unchanged.
elsewhere install libopus from your package manager. Encoding at 48 kHz, 20 ms frames. elsewhere install libopus from your package manager. Encoding at 48 kHz, 20 ms frames.
- **Voice Activation Detection (VAD)** with the same three modes as the TS3 client: - **Voice Activation Detection (VAD)** with the same three modes as the TS3 client:
- **Volume Gate** — RMS/dBFS threshold with a live input meter and hangover. - **Volume Gate** — RMS/dBFS threshold with a live input meter and hangover.
- **Automatic** — a dependency-free speech detector (short-term energy + spectral - **Automatic** — a Java port of WebRTC's RNN speech detector (`rnn_vad`), the one the
flatness + dominant frequency against an adaptive noise floor). TS3 client uses, with the same weights.
- **Hybrid** — transmit only when loud enough **and** detected as speech. - **Hybrid** — transmit only when loud enough **and** detected as speech.
- Optional **VAD over Push-To-Talk** (keep detecting voice while PTT is available). - Optional **VAD over Push-To-Talk** (keep detecting voice while PTT is available).
- **Capture pre-processing** — a mini audio-processing chain in the same order as the - **Capture pre-processing** — a Java port of the WebRTC audio-processing stages the TS3
TS3 client's WebRTC APM, applied before activation and encoding (and feeding the VAD): client enables, in its order: high-pass filter → **noise suppression** (WebRTC NS, TS3's
input gain → **high-pass filter** (80 Hz, always-on rumble/DC removal) → **noise four levels) → **typing attenuation** (the transient suppressor, told about key presses)
suppression** → **typing attenuation** → **AGC**: → **automatic gain control** (AGC2, WebRTC's successor to the AGC1 TS3 uses). Checked
- **Remove background noise** — spectral denoiser (decision-directed Wiener with against WebRTC built from source.
minimum-statistics noise tracking), with an adjustable removal level.
- **Typing attenuation** — detects impulsive keystroke transients (short, broadband,
high-frequency bursts) and ducks them while leaving sustained speech intact.
- **Automatic gain control (AGC)** — normalises mic loudness to a target level
(fast attack / slow release, noise-gated so silence is never amplified).
- Echo cancellation (WebRTC AEC3) is intentionally omitted — it needs the loudspeaker - Echo cancellation (WebRTC AEC3) is intentionally omitted — it needs the loudspeaker
reference signal and matters mainly for open speakers, not the typical headset. reference signal and matters mainly for open speakers, not the typical headset.
- **Push-to-Talk** — bind any key; transmits only while held (while the app is focused). - **Push-to-Talk** — bind any key or mouse button as a global hotkey; transmits only while held.
- **Continuous** transmission mode. - **Continuous** transmission mode.
- **Playback mixing** — each speaker gets its own Opus decoder and audio line, so - **Adaptive jitter buffer** — a port of the libspeex jitter buffer the TS3 client uses,
multiple simultaneous talkers are mixed and one slow decode never blocks others. set up and driven as it does: the delay follows the measured network jitter, and what it
has learned carries over between talk bursts.
- **Playback mixing** — each speaker gets their own Opus decoder and audio line (a stream
of their own in the PipeWire mixer), so one slow decode never blocks the others.
- **Whisper receive** — targeted voice is decoded and played like normal voice. - **Whisper receive** — targeted voice is decoded and played like normal voice.
- Per-client **mute**, master **deafen**, adjustable mic gain / playback volume. - Per-client **mute**, master **deafen**, adjustable mic gain / playback volume.
- **On-the-fly Opus tuning** — bitrate, complexity, VBR, FEC and voice/music codec - **On-the-fly Opus tuning** — bitrate, complexity, VBR, FEC and voice/music codec
@@ -212,21 +211,32 @@ voice activation while watching the live meter.
## Architecture ## Architecture
``` ```
core/ com.ts3client core/ com.ts3client
├── config.Settings persisted prefs (~/.ts3jclient/settings.properties) ├── config the profile (~/.ts3jclient): Settings, Bookmarks, IdentityStore,
├── config.Bookmarks persisted server bookmarks (with per-server identity) │ AwayMessages, …; AppDirs finds it, ProfileFiles writes it
├── config.IdentityStore managed identities (~/.ts3jclient/identities/*.ini) │ privately and all at once
├── teamspeak the official client's own data files ├── session
│ ├── TeamSpeakSettingsDb locates settings.db; reads Contacts and the ProtobufItems sync store │ ├── ServerSession one server tab's state and actions, whatever the frontend
│ ├── SqliteReader minimal read-only SQLite 3 file reader (no driver) │ └── Sessions what the servers share: the microphone, away state, nickname
│ ├── SyncItem a synced bookmark / identity / folder (Item_Data) ├── net
│ ├── SyncItemDecoder protobuf wire decoding of Item_Data, via ProtobufReader │ ├── TeamspeakConnection ties socket + audio + model; runs requests on its own pool
│ └── TeamSpeakImporter merges synced items into IdentityStore / Bookmarks │ ├── ConnectionEventHandler server events → model changes, sounds and log lines
├── audio abstractions + reusable DSP (no platform code) │ ├── ServerModel channel/client state; ChannelNode/ClientEntry view models
│ ├── AudioBackend factory for a platform's VoiceInput/VoiceOutput │ ├── ChannelAdmin/BanAdmin/AvatarAdmin channel editing, the ban list, our own avatar
│ ├── VoiceInput capture source (extends ts3j Microphone) + gating controls │ ├── IconRepository server icons (avatar/ holds the client avatar cache)
│ ├── VoiceOutput voice-packet playback sink │ ├── filetransfer channel file browsing, up- and downloads
│ ├── OpusParameters live-tunable encoder settings │ └── ConnectionListener frontend callbacks
│ └── SpeechDetector feature-based speech-probability VAD (Automatic/Hybrid) ├── audio the voice pipelines every platform shares
│ ├── AudioBackend/AudioIo what a platform supplies: devices, lines, Opus
│ ├── CaptureVoiceInput capture → processing → VAD/PTT gate → Opus encode
│ ├── StreamingVoiceOutput per-speaker decode, each through a VoiceStream
│ ├── JitterBuffer port of the libspeex jitter buffer TS3 uses
│ ├── PerSpeakerPlayout a line per speaker (desktop); MixedPlayout mixes onto one (Android)
│ ├── processing port of TS3's WebRTC capture chain: HPF, NS, transients, AGC2
│ ├── vad port of WebRTC's rnn_vad speech detector
│ └── opus codec interfaces
├── chatlog TS3-compatible chat logs, shared with the TS3 client by default
├── contacts friends and blocked clients, in TS3's own record format
├── text BBCode, TeamSpeak links, chat search
├── gfx ├── gfx
│ ├── IconPack an icon pack zip/folder + its settings.ini mapping │ ├── IconPack an icon pack zip/folder + its settings.ini mapping
│ └── IconPacks discovery of installed packs │ └── IconPacks discovery of installed packs
@@ -242,21 +252,21 @@ core/ com.ts3client
│ ├── SoundPacks discovery of installed packs │ ├── SoundPacks discovery of installed packs
│ ├── NotificationSettings per-action enabled / important flags │ ├── NotificationSettings per-action enabled / important flags
│ ├── SoundNotifier decides what is heard (pack + config + mute state) │ ├── SoundNotifier decides what is heard (pack + config + mute state)
│ └── SoundPlayer platform hook for rendering a sound │ ├── SoundPlayer renders a sound
└── net │ └── WavSoundPlayer decodes, resamples and mixes sounds onto one line
├── TeamspeakConnection ties socket + audio backend + model, translates events └── teamspeak the official client's own data files
├── ServerModel thread-safe channel/client state ├── TeamSpeakSettingsDb locates settings.db; reads Contacts and the ProtobufItems sync store
├── ChannelNode/ClientEntry view models ├── SqliteReader minimal read-only SQLite 3 file reader (no driver)
└── ConnectionListener frontend callbacks ├── SyncItem a synced bookmark / identity / folder (Item_Data)
├── SyncItemDecoder protobuf wire decoding of Item_Data, via ProtobufReader
└── TeamSpeakImporter merges synced items into IdentityStore / Bookmarks
desktop/ com.ts3client.audio.desktop + com.ts3client.hotkey.desktop desktop/ com.ts3client.audio.desktop + com.ts3client.hotkey.desktop
├── Opus Panama (FFM) binding to native libopus ├── DesktopAudioBackend PipeWire where it runs, Java Sound elsewhere, and native Opus
├── OpusEncoder/OpusDecoder thin codec wrappers ├── Opus/NativeOpusCodec Panama (FFM) binding to native libopus
├── AudioDevices device enumeration + line opening (48 kHz/16-bit) ├── AudioDevices device enumeration + line opening (48 kHz/16-bit)
├── JavaSoundVoiceInput capture + VAD/PTT gating + Opus encode ├── JavaSoundCapture/JavaSoundPlayback Java Sound lines
├── JavaSoundVoiceOutput per-client Opus decode + playback + mixing ├── pipewire PipeWire streams through FFM
├── WavSoundPlayer sound-pack playback: decode, resample and mix on one line
├── JavaSoundAudioBackend wires the above into the core AudioBackend
└── hotkey.desktop └── hotkey.desktop
├── XInput2InputHook global key/button capture via XInput2 raw events ├── XInput2InputHook global key/button capture via XInput2 raw events
├── XRecordInputHook the same via X11's RECORD extension, as a fallback ├── XRecordInputHook the same via X11's RECORD extension, as a fallback
@@ -274,7 +284,7 @@ swing/ com.ts3client
├── Spacers TS3 spacer-channel name parsing/rendering ├── Spacers TS3 spacer-channel name parsing/rendering
├── InfoPanel channel description / client group + details view ├── InfoPanel channel description / client group + details view
├── ChatPanel chat log + input ├── ChatPanel chat log + input
├── SettingsDialog audio + VAD options with live meter ├── SettingsDialog the Options dialog; one panel per tab (DevicesPanel, …, ChatLogsPanel)
├── HotkeysPanel hotkey list (Options → Hotkeys) ├── HotkeysPanel hotkey list (Options → Hotkeys)
├── HotkeyDialog add/edit one hotkey: action, combination, trigger, scope ├── HotkeyDialog add/edit one hotkey: action, combination, trigger, scope
├── HotkeyService bindings + engine + input hook, for the dialogs ├── HotkeyService bindings + engine + input hook, for the dialogs
@@ -293,11 +303,7 @@ swing/ com.ts3client
``` ```
## Known limitations / next steps ## Known limitations / next steps
- The **Automatic/Hybrid** VAD uses a lightweight energy/spectral detector rather
than the WebRTC GMM model the official client ships; it is intentionally
dependency-free and reusable in the core library.
- Whisper is received/played but not yet **sendable** from the UI. - Whisper is received/played but not yet **sendable** from the UI.
- Playback decodes streams as mono; stereo music-bot audio is down-mixed.
- Hotkey actions TeamSpeak offers but this client does not perform yet (listed but - Hotkey actions TeamSpeak offers but this client does not perform yet (listed but
greyed out in the dialog): capture/playback/hotkey profiles and sound packs, greyed out in the dialog): capture/playback/hotkey profiles and sound packs,
whisper and push-to-whisper, recording, plugins, server groups, talk power, 3D whisper and push-to-whisper, recording, plugins, server groups, talk power, 3D
@@ -309,4 +315,3 @@ swing/ com.ts3client
covered by tests. covered by tests.
- When no backend can start, push-to-talk does not work at all, not even with the window - When no backend can start, push-to-talk does not work at all, not even with the window
focused; a focus-scoped fallback hook would restore the pre-hotkey behaviour. focused; a focus-scoped fallback hook would restore the pre-hotkey behaviour.
- No file transfer, avatars, or server/channel administration UI yet.