Skip to content

feat(player): auto-switch source when stuck buffering (opt-in) - #3218

Draft
RjBiermann wants to merge 7 commits into
recloudstream:masterfrom
RjBiermann:smart-source-failover
Draft

RjBiermann wants to merge 7 commits into
recloudstream:masterfrom
RjBiermann:smart-source-failover

Conversation

@RjBiermann

@RjBiermann RjBiermann commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Opt-in player setting — "Auto-switch stuck source" (Player settings, off by default). When a source stalls during playback, the player switches to the next ranked source instead of spinning forever. Fixes the silent-stall case behind #3096 and #3097.

Problem

Hard errors already switch mirrors automatically (PlayerView.playerError → nextMirror()), and #2968 added mid-stream timeout recovery by re-preparing. But when a source just silently stalls — player sits in STATE_BUFFERING, position frozen, no exception thrown — nothing reacts. The user gets an endless spinner unless they manually open the source picker.

How it works

StuckBufferingWatcher polls the player once per second (cheap: one status check + one position read). When enabled, it switches to the next source via the existing nextMirror() / loadLink(sameEpisode = true) path — position is preserved, no player code changes.

Triggers (either one fires a switch):

Trigger Condition
Single stall buffering with no position advance for 10 s
Repeated stalls buffering ≥ 15 of the last 60 s (~25% duty) — catches choppy streams that keep creeping forward

Guardrails:

  • 60 s cooldown between switches
  • Budget reset on success — a switch that leads to 30 s+ of playback progress past the switch point restores the full budget, so working sources get effectively unlimited switches. Only a cycle where nothing ever plays is capped at 3, preventing an infinite loop when all mirrors are dead
  • hasNextMirror() checked first — never triggers the "no links found" page-pop
  • Manual switch respected — picking a source yourself resets the stall tracking, so the previous source's stall history can't auto-switch away from your pick
  • Live streams & torrents excluded — mirror-switching fights the live-edge logic on live TV, and torrent/magnet buffering is inherent to the source, so neither triggers a switch
  • Firing clears the stall state (hysteresis); a toast informs the user

Telemetry: per-host failure counters are recorded locally (DataStore, incremented on playerError and stall switches). Nothing reads them yet — they're the data source for a possible future reliability-ordering PR. Local only, never synced.

Non-breaking / scope

  • Off by default — with the setting off, the tick early-returns and behavior is identical to today (covered by unit tests)
  • Touches only ui/player wiring, a settings row, strings, and one preference. No changes to CS3IPlayer.kt, PlayerView.kt, MainAPI, or any extension-facing API
  • Interaction with fix: skip broken CF bypass on Android TV + recover from mid-stream ti… #2968: the timeout-recovery prepare() retry and this watcher can alternate (retry → buffer → retry…). The frozen-position detection acts as a circuit breaker for that loop after ~10 s. fix: skip broken CF bypass on Android TV + recover from mid-stream ti… #2968's path itself is untouched
  • Known limitation: the single-stall trigger can't distinguish user-scrubbing from real stalls; the cooldown and success-reset budget bound the damage (noted in code)

Testing

Unit tests (in this PR): fire timing, cooldown, session cap, budget reset on recovery, host parsing — fake clock, no Android dependencies.

Manual, TV emulator, throttled network (3G / 3.5G), multiple sessions:

  • Switches fire on both triggers; position preserved; toast shown
  • Budget reset verified live: after each successful switch that resumed playback, a later stall correctly fired again as "switching source (1/3)"
  • Feature off → no behavior change observed

Transparency note: one session hit an unrelated OutOfMemoryError (heap saturated at the 512 MB debug-build limit during source resolution). GC logs show the heap hit 511 MB ~16 s before the watcher fired; the watcher's footprint is a deque of ~60 longs. Pre-existing emulator/extension behavior, not introduced here.

AI usage disclosure

This change was written with AI assistance (GLM 4.6 / glm-5.3-flash via OpenCode Go, coding-agent workflow), per AI-POLICY. Unit tests pass locally; the feature was manually smoke-tested on an emulator as described above.

@RjBiermann
RjBiermann marked this pull request as ready for review September 25, 2026 21:50
TEMP added 2 commits September 25, 2026 22:13
Off-by-default setting that reacts to a source stuck buffering: polls the
player once per second and switches to the next ranked source when either
buffering without position progress for 10s, or buffering >=15 of the last
60s. Guardrails: 60s cooldown between switches, max 3 per session.

Reuses the existing nextMirror()/loadLink(sameEpisode=true) machinery, so
position is preserved and no player code path changes. Also records
per-host failure counts locally (used by a later ranking PR).

With the setting off, behavior is unchanged.
The 3-switch session cap treated working and dead sources alike: a
long-stalling episode burned the whole budget on mirrors that
ultimately played fine. Now a switch that leads to 30s+ of playback
progress past the switch point resets the counter, so working sources
are effectively unlimited while an all-mirrors-dead cycle still stops
after 3 instead of looping forever.
@RjBiermann
RjBiermann force-pushed the smart-source-failover branch from 1441644 to 0983245 Compare September 26, 2026 02:56
@RjBiermann

Copy link
Copy Markdown
Contributor Author

Rebased onto master and pushed a follow-up fix: the per-session switch cap (3) burned its whole budget on mirrors that eventually played fine. Now a switch that leads to 30s+ of playback progress past the switch point resets the counter — working sources get effectively unlimited switches, while a cycle where nothing plays is still capped at 3 so it can't loop forever.

Verified on the TV emulator with network throttling (3G / 3.5G): switches fire on both triggers, budget restores after each recovery, unit tests + lint pass.

TEMP added 3 commits September 25, 2026 23:11
CodeFactor flags tick() as a complex method (complexity 16). Extract
the buffering-window bookkeeping, stall clock, trigger check and switch
side effects into small helpers — no behavior change.
The source picker goes through the same loadLink(sameEpisode=true)
path as the auto-switch, so the stall clock and buffering window
carried over: a source the user manually picked could be switched
away from ~10s later because the previous source's stall history was
still armed. Reset stall tracking on a manual pick; the switch budget
is untouched (a manual pick is not the watcher's switch).
@RjBiermann

Copy link
Copy Markdown
Contributor Author
untitled.mp4

Tested with network throttle and skipping ahead.

TEMP added 2 commits September 25, 2026 23:55
- reset stall tracking on every same-episode mirror switch, not just
  manual picks: the error-failover path loaded fresh sources with the
  previous source's armed stall clock, risking a spurious switch ~10s later
- skip auto-switch for live streams (fights live-edge/zapping) and for
  torrent/magnet sources (buffering is inherent, switching cannot help)
- stop the watcher in onDestroyView: after the host view is torn down the
  status getter defaults to IsBuffering and position reads 0, which could
  fire a switch into a dying fragment
- nits: use the tick's timestamp in updateStallClock, init
  lastSwitchPositionMs to 0 to avoid the Long.MIN_VALUE overflow
…itch

isStuck() took a buffering param it never read (the stall clock already
encodes it), and fireSwitch re-inlined the exact resetStallTracking()
body. Net -3 lines, no behavior change.
@RjBiermann
RjBiermann marked this pull request as draft September 26, 2026 04:01

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant