From 51f19c2d033a2a6734462e9671b2acaf60148373 Mon Sep 17 00:00:00 2001 From: ericek111 Date: Fri, 25 Sep 2026 11:06:31 +0000 Subject: [PATCH] Bring the README's architecture and features up to date The architecture tree predated the session, chatlog, contacts, text, file transfer and voice pipeline packages, and named a backend that no longer exists. The feature list still described the first speech detector and processing chain rather than the WebRTC ports, and knew nothing of the jitter buffer; the limitations listed file transfers, avatars and administration as missing. Co-Authored-By: Claude Opus 5.5 --- ts3-client/README.md | 109 ++++++++++++++++++++++--------------------- 1 file changed, 57 insertions(+), 52 deletions(-) diff --git a/ts3-client/README.md b/ts3-client/README.md index c1c2aab..1c6b31d 100644 --- a/ts3-client/README.md +++ b/ts3-client/README.md @@ -10,13 +10,14 @@ TeaVM/CheerpJ) without touching the library: | Module | Artifact | Responsibility | |--------|----------|----------------| -| `core` | `ts3-client-core` | Frontend-agnostic library: protocol integration, server model, connection orchestration, audio **abstractions**. No UI, no platform audio, no native code. | -| `desktop` | `ts3-client-desktop` | Desktop audio backend: Java Sound capture/playback + native Opus via the FFM API (project Panama), implementing the core audio interfaces. | +| `core` | `ts3-client-core` | Frontend-agnostic library: protocol integration, server model, connection orchestration, the voice pipelines and their DSP. No UI, no platform audio API, no native code. | +| `desktop` | `ts3-client-desktop` | Desktop audio backend: PipeWire (or Java Sound) capture/playback + native Opus via the FFM API (project Panama), plus the global hotkey hooks. | | `swing` | `ts3-client-swing` | Swing desktop UI + entry point. Depends on `core` and `desktop`. | The core exposes `AudioBackend` / `VoiceInput` / `VoiceOutput`; the frontend injects -a concrete backend (`JavaSoundAudioBackend`) into `TeamspeakConnection`. A different -frontend supplies its own UI and audio backend while reusing `core` unchanged. +a concrete backend (`DesktopAudioBackend`) into `TeamspeakConnection`. A different +frontend supplies its own UI and audio backend while reusing `core` unchanged, as the +Android app in [`../android`](../android) does. ## Features @@ -28,26 +29,24 @@ frontend supplies its own UI and audio backend while reusing `core` unchanged. elsewhere install libopus from your package manager. Encoding at 48 kHz, 20 ms frames. - **Voice Activation Detection (VAD)** with the same three modes as the TS3 client: - **Volume Gate** — RMS/dBFS threshold with a live input meter and hangover. - - **Automatic** — a dependency-free speech detector (short-term energy + spectral - flatness + dominant frequency against an adaptive noise floor). + - **Automatic** — a Java port of WebRTC's RNN speech detector (`rnn_vad`), the one the + TS3 client uses, with the same weights. - **Hybrid** — transmit only when loud enough **and** detected as speech. - Optional **VAD over Push-To-Talk** (keep detecting voice while PTT is available). -- **Capture pre-processing** — a mini audio-processing chain in the same order as the - TS3 client's WebRTC APM, applied before activation and encoding (and feeding the VAD): - input gain → **high-pass filter** (80 Hz, always-on rumble/DC removal) → **noise - suppression** → **typing attenuation** → **AGC**: - - **Remove background noise** — spectral denoiser (decision-directed Wiener with - minimum-statistics noise tracking), with an adjustable removal level. - - **Typing attenuation** — detects impulsive keystroke transients (short, broadband, - high-frequency bursts) and ducks them while leaving sustained speech intact. - - **Automatic gain control (AGC)** — normalises mic loudness to a target level - (fast attack / slow release, noise-gated so silence is never amplified). +- **Capture pre-processing** — a Java port of the WebRTC audio-processing stages the TS3 + client enables, in its order: high-pass filter → **noise suppression** (WebRTC NS, TS3's + four levels) → **typing attenuation** (the transient suppressor, told about key presses) + → **automatic gain control** (AGC2, WebRTC's successor to the AGC1 TS3 uses). Checked + against WebRTC built from source. - Echo cancellation (WebRTC AEC3) is intentionally omitted — it needs the loudspeaker reference signal and matters mainly for open speakers, not the typical headset. -- **Push-to-Talk** — bind any key; transmits only while held (while the app is focused). +- **Push-to-Talk** — bind any key or mouse button as a global hotkey; transmits only while held. - **Continuous** transmission mode. -- **Playback mixing** — each speaker gets its own Opus decoder and audio line, so - multiple simultaneous talkers are mixed and one slow decode never blocks others. +- **Adaptive jitter buffer** — a port of the libspeex jitter buffer the TS3 client uses, + set up and driven as it does: the delay follows the measured network jitter, and what it + has learned carries over between talk bursts. +- **Playback mixing** — each speaker gets their own Opus decoder and audio line (a stream + of their own in the PipeWire mixer), so one slow decode never blocks the others. - **Whisper receive** — targeted voice is decoded and played like normal voice. - Per-client **mute**, master **deafen**, adjustable mic gain / playback volume. - **On-the-fly Opus tuning** — bitrate, complexity, VBR, FEC and voice/music codec @@ -212,21 +211,32 @@ voice activation while watching the live meter. ## Architecture ``` core/ com.ts3client -├── config.Settings persisted prefs (~/.ts3jclient/settings.properties) -├── config.Bookmarks persisted server bookmarks (with per-server identity) -├── config.IdentityStore managed identities (~/.ts3jclient/identities/*.ini) -├── teamspeak the official client's own data files -│ ├── TeamSpeakSettingsDb locates settings.db; reads Contacts and the ProtobufItems sync store -│ ├── SqliteReader minimal read-only SQLite 3 file reader (no driver) -│ ├── SyncItem a synced bookmark / identity / folder (Item_Data) -│ ├── SyncItemDecoder protobuf wire decoding of Item_Data, via ProtobufReader -│ └── TeamSpeakImporter merges synced items into IdentityStore / Bookmarks -├── audio abstractions + reusable DSP (no platform code) -│ ├── AudioBackend factory for a platform's VoiceInput/VoiceOutput -│ ├── VoiceInput capture source (extends ts3j Microphone) + gating controls -│ ├── VoiceOutput voice-packet playback sink -│ ├── OpusParameters live-tunable encoder settings -│ └── SpeechDetector feature-based speech-probability VAD (Automatic/Hybrid) +├── config the profile (~/.ts3jclient): Settings, Bookmarks, IdentityStore, +│ AwayMessages, …; AppDirs finds it, ProfileFiles writes it +│ privately and all at once +├── session +│ ├── ServerSession one server tab's state and actions, whatever the frontend +│ └── Sessions what the servers share: the microphone, away state, nickname +├── net +│ ├── TeamspeakConnection ties socket + audio + model; runs requests on its own pool +│ ├── ConnectionEventHandler server events → model changes, sounds and log lines +│ ├── ServerModel channel/client state; ChannelNode/ClientEntry view models +│ ├── ChannelAdmin/BanAdmin/AvatarAdmin channel editing, the ban list, our own avatar +│ ├── IconRepository server icons (avatar/ holds the client avatar cache) +│ ├── filetransfer channel file browsing, up- and downloads +│ └── ConnectionListener frontend callbacks +├── audio the voice pipelines every platform shares +│ ├── AudioBackend/AudioIo what a platform supplies: devices, lines, Opus +│ ├── CaptureVoiceInput capture → processing → VAD/PTT gate → Opus encode +│ ├── StreamingVoiceOutput per-speaker decode, each through a VoiceStream +│ ├── JitterBuffer port of the libspeex jitter buffer TS3 uses +│ ├── PerSpeakerPlayout a line per speaker (desktop); MixedPlayout mixes onto one (Android) +│ ├── processing port of TS3's WebRTC capture chain: HPF, NS, transients, AGC2 +│ ├── vad port of WebRTC's rnn_vad speech detector +│ └── opus codec interfaces +├── chatlog TS3-compatible chat logs, shared with the TS3 client by default +├── contacts friends and blocked clients, in TS3's own record format +├── text BBCode, TeamSpeak links, chat search ├── gfx │ ├── IconPack an icon pack zip/folder + its settings.ini mapping │ └── IconPacks discovery of installed packs @@ -242,21 +252,21 @@ core/ com.ts3client │ ├── SoundPacks discovery of installed packs │ ├── NotificationSettings per-action enabled / important flags │ ├── SoundNotifier decides what is heard (pack + config + mute state) -│ └── SoundPlayer platform hook for rendering a sound -└── net - ├── TeamspeakConnection ties socket + audio backend + model, translates events - ├── ServerModel thread-safe channel/client state - ├── ChannelNode/ClientEntry view models - └── ConnectionListener frontend callbacks +│ ├── SoundPlayer renders a sound +│ └── WavSoundPlayer decodes, resamples and mixes sounds onto one line +└── teamspeak the official client's own data files + ├── TeamSpeakSettingsDb locates settings.db; reads Contacts and the ProtobufItems sync store + ├── SqliteReader minimal read-only SQLite 3 file reader (no driver) + ├── SyncItem a synced bookmark / identity / folder (Item_Data) + ├── SyncItemDecoder protobuf wire decoding of Item_Data, via ProtobufReader + └── TeamSpeakImporter merges synced items into IdentityStore / Bookmarks desktop/ com.ts3client.audio.desktop + com.ts3client.hotkey.desktop -├── Opus Panama (FFM) binding to native libopus -├── OpusEncoder/OpusDecoder thin codec wrappers +├── DesktopAudioBackend PipeWire where it runs, Java Sound elsewhere, and native Opus +├── Opus/NativeOpusCodec Panama (FFM) binding to native libopus ├── AudioDevices device enumeration + line opening (48 kHz/16-bit) -├── JavaSoundVoiceInput capture + VAD/PTT gating + Opus encode -├── JavaSoundVoiceOutput per-client Opus decode + playback + mixing -├── WavSoundPlayer sound-pack playback: decode, resample and mix on one line -├── JavaSoundAudioBackend wires the above into the core AudioBackend +├── JavaSoundCapture/JavaSoundPlayback Java Sound lines +├── pipewire PipeWire streams through FFM └── hotkey.desktop ├── XInput2InputHook global key/button capture via XInput2 raw events ├── XRecordInputHook the same via X11's RECORD extension, as a fallback @@ -274,7 +284,7 @@ swing/ com.ts3client ├── Spacers TS3 spacer-channel name parsing/rendering ├── InfoPanel channel description / client group + details view ├── ChatPanel chat log + input - ├── SettingsDialog audio + VAD options with live meter + ├── SettingsDialog the Options dialog; one panel per tab (DevicesPanel, …, ChatLogsPanel) ├── HotkeysPanel hotkey list (Options → Hotkeys) ├── HotkeyDialog add/edit one hotkey: action, combination, trigger, scope ├── HotkeyService bindings + engine + input hook, for the dialogs @@ -293,11 +303,7 @@ swing/ com.ts3client ``` ## Known limitations / next steps -- The **Automatic/Hybrid** VAD uses a lightweight energy/spectral detector rather - than the WebRTC GMM model the official client ships; it is intentionally - dependency-free and reusable in the core library. - Whisper is received/played but not yet **sendable** from the UI. -- Playback decodes streams as mono; stereo music-bot audio is down-mixed. - Hotkey actions TeamSpeak offers but this client does not perform yet (listed but greyed out in the dialog): capture/playback/hotkey profiles and sound packs, whisper and push-to-whisper, recording, plugins, server groups, talk power, 3D @@ -309,4 +315,3 @@ swing/ com.ts3client covered by tests. - When no backend can start, push-to-talk does not work at all, not even with the window focused; a focus-scoped fallback hook would restore the pre-hotkey behaviour. -- No file transfer, avatars, or server/channel administration UI yet.