Files
shorebird-updater/library
Eric Seidel 90c349287d refactor: per-patch lifecycle state machine (#352)
* refactor: introduce per-patch lifecycle state machine (types + persistence)

First slice of shorebirdtech/shorebird#3737 — replace the scattered
storage of per-patch state across download_state.rs sidecars, bare
files in downloads/, and PatchesState fields (next_boot_patch,
last_booted_patch, known_bad_patches) with a single per-patch
state.json driven by an explicit state machine.

This commit only adds the types and the storage layer; nothing wires
into update_internal or report_launch_* yet. Subsequent commits will
build out transitions and migrate the call sites.

States:
  - Downloading { url, hash, signature, partial_size }
  - Downloaded  { url, hash, signature, size }
  - Installed   { hash, signature, size }
  - Bad         { reason, hash?, signature?, size? }      // tombstone

Operations:
  - mark_bad(n, reason) — sugar over write Bad{} + cleanup
  - cleanup(n) — state-aware: keeps Bad tombstone, else forgets dir

Per-release pointers (next_boot, last_booted, currently_booting,
boot_started_at) move to a separate pointers.json holding patch
numbers; metadata lives once per patch in state.json instead of
duplicated across pointers.

Persistence rides on the existing atomic disk_io::write (sibling-
write + rename) so partial writes can't leave torn files.

* refactor: add download/install/boot transitions to lifecycle

Extends the lifecycle module with the methods update_internal and
report_launch_* will need:

  - decide_start(n, url, hash) → DownloadAction (Fresh, Resume, Complete, Skip)
  - record_download_started / _complete
  - record_install_complete (transitions Downloaded → Installed,
    removes the now-unneeded compressed download file)
  - promote_to_next_boot (transitions a freshly Installed patch into
    pointers.next_boot, retiring any unbooted predecessor)
  - record_boot_start / _success / _failure
  - detect_boot_crash_on_init (handles the breadcrumb left when a
    prior process crashed during boot)
  - validate_next_boot_patch (size + signature checks; marks
    Bad{ValidationFailed} on failure)
  - recompute_next_boot

Nothing wires into the existing updater.rs / report_launch_* yet —
those changes follow in the cutover commit.

* refactor: replace UpdaterState's PatchManager backend with PatchLifecycle

Drops the patch_manager dependency from updater_state.rs. UpdaterState
now owns a PatchLifecycle directly, with all patch-related methods
delegating to it:

  install_patch        → write Installed state + promote to next_boot
  next_boot_patch      → pointers.next_boot_patch + installed_artifact_path
  is_known_bad_patch   → read_state matches Bad
  uninstall_patch      → cleanup + recompute_next_boot
  record_boot_*        → lifecycle's boot transitions
  validate_next_boot_patch → lifecycle's signature/size check

UpdaterState's own state.json shrinks to {client_id, release_version,
queued_events}. Per-release per-patch state lives entirely in the
lifecycle (pointers.json + patches/{N}/state.json).

Other notable behavior changes:

  - record_boot_failure_for_patch no longer requires
    currently_booting_patch to match. Matches the prior PatchManager
    semantics: clear breadcrumb, mark Bad{BootCrash}, recompute.
  - recompute_next_boot leaves a valid Installed next_boot_patch alone
    instead of promoting last_booted over it. Without this, processing
    server rollbacks would clobber a freshly-installed newer patch.

Tests in updater.rs that asserted file paths (patches_state.json) update
to the new layout (pointers.json). The MockManagePatches-backed unit
tests in updater_state.rs are gone — replaced by direct end-to-end tests
that exercise PatchLifecycle through UpdaterState.

patch_manager.rs is still on disk and compiled (nothing references it
from cache/mod.rs anymore) — deletion happens after the updater.rs
cutover.

* refactor: cut update_internal over to PatchLifecycle

Replaces the download_state.rs sidecar machinery and the bespoke
DownloadStartState / should_install_patch helpers in update_internal
with calls into PatchLifecycle.

Concrete changes:

  - Drops compute_resume_offset / determine_download_start_state /
    DownloadStartState — replaced by PatchLifecycle::decide_start
    returning DownloadAction.
  - Drops should_install_patch / ShouldInstallPatchCheckResult —
    folded into the same DownloadAction match. KnownBad → UpdateIsBadPatch,
    AlreadyInstalled → NoUpdate, the rest proceed.
  - Drops install_downloaded_patch's bespoke "stage in downloads/, rename
    on install" choreography. The lifecycle owns the per-patch directory;
    inflate writes directly into patches/{N}/dlc.vmcode and the
    Downloaded → Installed transition removes the now-unneeded compressed
    download file.
  - Drops cleanup_download_artifacts and clean_download_dir. mark_bad
    handles tombstone-aware cleanup for failures, and the lifecycle owns
    its directory.
  - Marks the patch Bad{InvalidPatchBytes} when inflate fails and
    Bad{InstallHashMismatch} when check_hash fails. Subsequent attempts
    short-circuit at decide_start (Skip(KnownBad)) instead of
    re-downloading and re-failing on every cycle. This is the marks-bad-
    on-install-failure behavior we deferred from #351.
  - Preserves the explicit "server-side Content-Length vs total_bytes"
    mismatch check — surfaces the contract violation directly instead of
    obliquely via inflate failure.

decide_start consults the on-disk download file as the source of truth
for "how many bytes do we have so far." The denormalized partial_size
in PatchState::Downloading is kept for diagnostics/serde stability but
not consulted for the resume offset; this matches the prior
file-size-on-disk behavior and avoids needing to update state.json
mid-stream from the network layer.

For Downloading / Downloaded states where the OS evicted the artifact
file out from under us (e.g. iOS code-cache eviction), decide_start
falls through to Fresh so the next attempt re-downloads from scratch.

The failing-test fixes update file paths and pre-staged state from the
old layout (downloads/{N} + downloads/{N}.download.json) to the new
layout (patches/{N}/download + patches/{N}/state.json). Several tests
that targeted the now-deleted helpers directly are removed; their
behavior is covered by the new lifecycle unit tests.

* refactor: delete patch_manager.rs and download_state.rs

Both modules are now subsumed by PatchLifecycle:

  - patch_manager.rs (~1600 lines) managed the per-release `patches/`
    tree, the `next_boot_patch` / `last_booted_patch` /
    `currently_booting_patch` / `known_bad_patches` fields of
    `PatchesState`, and the fall-back-from-bad-patch logic. All of
    this is now in lifecycle.rs split between `PatchLifecycle`
    (per-patch state.json) and `ReleasePointers` (a single
    pointers.json).
  - download_state.rs (~125 lines) wrote the per-download sidecar
    JSON. Now folded into PatchState::Downloading /
    PatchState::Downloaded variants in state.json, owned by the
    lifecycle module.

The MockManagePatches / ManagePatches trait machinery in
updater_state.rs's tests goes away with patch_manager. The lifecycle
operations are tested directly in cache::lifecycle::tests against a
real on-disk filesystem under TempDir, which catches issues the
mocks couldn't (e.g. file-existence cross-checks in `decide_start`).

The cargo workspace now has 212 passing tests vs. 257 before, the
delta being the patch_manager unit tests whose subjects no longer
exist. End-to-end coverage in `updater.rs::tests` is unchanged and
all green.

* refactor: address self-review feedback on lifecycle PR

Nine fixes from the self-review pass on shorebirdtech/updater#352:

  - Drop `partial_size` from PatchState::Downloading. The field was
    misleading (decide_start reads from disk, not from the recorded
    value) and unused. record_download_started loses its 5th arg.
  - Gate `UpdaterState::install_patch` to `#[cfg(test)]`. Production
    no longer routes through it; only test_utils and the updater_state
    tests do. The gate makes the divergence intentional and prevents
    future production callers.
  - install_patch defensively removes any prior `dlc.vmcode` before
    rename so behavior is OS-agnostic (POSIX rename overwrites silently;
    Windows fails). Also mirrors record_install_complete's cleanup of
    a stale `download` file in the patch dir.
  - Document on `lifecycle()` / `lifecycle_mut()` that the direct
    accessors are intentional — wrapping every transition would be
    churn for no reader benefit.
  - recompute_next_boot now clears `last_booted_patch` when its
    on-disk record is gone (Unknown), so pointers.json doesn't
    accumulate references to nothing. A `Bad` last_booted is left
    alone — that's a useful breadcrumb and recompute simply doesn't
    promote it.
  - Promote `download_artifact_path` and `installed_artifact_path` to
    methods on PatchLifecycle for symmetry with state_path /
    pointers_path. Updates all call sites.
  - Document mark_bad-on-Bad merge semantics: latest reason wins,
    other fields preserved. Hypothetical in practice but no longer
    silent.
  - Add tests for recompute_next_boot's stale-pointer clearing
    (Unknown → cleared, Bad → kept). Brings the suite to 214 tests.

The mark_bad-cleanup-stale-bytes recovery path I flagged was already
covered by `cleanup_on_bad_patch_keeps_tombstone` — no new test needed.

* refactor: undo two over-corrections from the prior self-review fix

Two issues introduced by 07ea84a that the second self-review caught:

  - update_internal computed download_path / installed_path via
    `with_state(...).lifecycle().download_artifact_path(n)`. That
    routed a pure-function path computation through a state read,
    serializing it against other state operations and adding a disk
    read per call. Restore the free `download_artifact_path` /
    `installed_artifact_path` functions as the canonical entry point;
    the methods on PatchLifecycle are kept as thin wrappers for
    callers that already hold a lifecycle handle.
  - install_patch's "defensive remove_file before rename" added a
    TOCTOU window between the exists() check and the rename for no
    real benefit — install_patch is now `#[cfg(test)]`-only, the
    Windows-rename concern was theoretical for tests, and POSIX rename
    atomically overwrites. Drop the explicit remove_file.

* ci: fix warnings-as-errors and cspell

Three CI failures from the prior push:

  - `Context` and `bail` imports in updater_state.rs are only used
    inside the `#[cfg(test)] install_patch` helper. Moved them under
    `#[cfg(test)]` so non-test builds don't see unused imports.
  - `PathBuf` import in updater.rs became unused after switching the
    download/install paths to take `Path::new(&config.storage_dir)`
    directly. Dropped.
  - `FileOperation::DeleteDir` is no longer constructed by production
    code (was used by the deleted patch_manager.rs). Mirrored the
    existing `#[allow(dead_code)]` annotation already on `DeleteFile`
    rather than removing the variant — the format/handling code for
    it is still useful for any future code that needs to surface a
    "delete dir failed" error.

CSpell additions: `unparseable`, `tombstoned`, `roundtrips` — words
introduced by the lifecycle module's docs and test names.

* ci: gate PathBuf import to platforms that actually use it

The previous fix removed `PathBuf` from the top-level use, which
broke the lib build on non-android non-test platforms (macOS,
Linux, Windows) where `libapp_path_from_settings` uses `PathBuf`.
That function is itself gated to `cfg(not(any(target_os = \"android\",
test)))`, so the PathBuf import needs the same gate.

Restoring it as a cfg-gated separate import — keeps the lib build
green on all three runners and keeps the lib-test build clean
under -D warnings.

* test: cover the gaps surfaced by the coverage report

Adds 7 tests targeting branches that were uncovered:

  - mark_bad_from_downloading_records_partial_file_size: the
    Downloading source state in mark_bad reads bytes-on-disk for
    the recorded `size` field; the existing tests only covered
    Installed → Bad.
  - mark_bad_overwrites_reason_when_already_bad: documented merge
    semantics (latest reason wins, other fields preserved) now
    has a test backing the doc.
  - validate_next_boot_patch_marks_bad_when_artifact_missing: the
    "Installed state but dlc.vmcode is gone" path. Existing test
    covered size mismatch; this covers the missing-file branch.
  - validate_next_boot_patch_marks_bad_when_pointer_targets_non_installed:
    the case where the next_boot pointer references a state other
    than Installed (state.json + pointers can diverge through
    corruption).
  - install_patch_install_only_{accepts_valid,rejects_missing,rejects_bad}_signature:
    cover the InstallOnly verification path in
    UpdaterState::install_patch using the existing test keypair
    from signing.rs.

Coverage improvements:
  - cache/lifecycle.rs:    94.65% → 96.06% lines, 98.72% → 100% fns
  - cache/updater_state.rs: 94.34% → 96.94% lines, 95.65% → 98% fns
  - total:                 94.70% → 95.11% lines

214 → 221 tests, all passing under -D warnings.

* ci: add 'keypair' to cspell dictionary

* test: audit deleted patch_manager tests; restore strict checks

Goes through every test deleted with patch_manager.rs (~40) and either
maps it to a new test (with a `Ports patch_manager.rs::mod::name`
comment) or adds a port. Three behavioral regressions caught and fixed:

  - record_boot_start now requires next_boot_patch to match the arg.
    Old PatchManager had this defensive check; my refactor dropped it.
    Carries forward the engine-vs-state agreement guard.
  - Added strict-mode signature tests for validate_next_boot_patch
    (valid, missing, invalid, no public key, bad public key). Ports
    five `validate_next_boot_patch_tests::strict_mode_*` tests.
  - Added InstallOnly + no public_key test (ports
    add_patch_tests::install_only_succeeds_with_any_signature_if_no_public_key).

Plus two scenario tests that didn't have direct coverage:

  - rolled_back_patch_not_resurrected_when_replacement_fails: ports
    fall_back_tests::rollback_then_failed_replacement_does_not_resurrect_rolled_back_patch
    against the new state-based fallback (cleanup forgets, recompute
    won't promote a None state).
  - validate_then_promote_catches_corrupted_last_booted: ports
    validate_next_boot_patch_tests::does_not_fall_back_to_last_booted_patch_if_corrupted
    with a note: new code catches the corruption at the next validate
    pass instead of at fall-back time, but the user-visible outcome
    (boot the base release) is the same.

Test helper split: `install_state` (writes Installed without touching
pointers) vs callers explicitly setting pointers. Avoids the prior
helper's auto-promote masking pointer-management behavior the tests
were trying to exercise.

Tests deliberately not ported (with reasoning):

  - debug_tests::manage_patches_is_debug, patch_manager_is_debug:
    Trait-Debug tests for types that no longer exist. Replaced
    implicitly by `#[derive(Debug)]` on PatchLifecycle.
  - fall_back_tests::succeeds_if_deleting_artifacts_fails: needs
    filesystem fault injection that wasn't worth the test
    infrastructure to bring back.
  - record_boot_success_for_patch_tests::deletes_unrecognized_directories_in_patches_dir:
    behavior change — new cleanup_older_than skips non-numeric
    directory names rather than deleting them. Defensive; the prior
    `rm -rf` was overly aggressive.

Coverage: 222 → 230 tests, all green under -D warnings.

* test: port the two patch_manager tests I'd handwaved

Both turned out trivial to port — neither was actually doing fault
injection, just exploiting graceful-failure paths.

  - record_boot_success_deletes_unrecognized_directories_in_patches_dir:
    pre-creates a junk dir and a stray file in patches/ before the
    install/boot, asserts they're gone after record_boot_success.
    Restores the prior `cleanup_older_than` behavior of removing
    non-numeric entries — we own patches/, so anything not named like
    a patch number is corruption to sweep up.
  - record_boot_failure_succeeds_if_artifact_dirs_are_already_gone:
    pre-deletes both patch dirs and verifies record_boot_failure still
    completes correctly (mark_bad recreates the dir for the tombstone;
    cleanup paths are graceful when the dir is missing).

Reverts the prior behavior change in cleanup_older_than that skipped
non-numeric entries instead of deleting them. The "defensive" framing
was wrong — we own patches/, anything unexpected there came from a
different version of our code or actual corruption, and leaving it
behind is the wrong default.

* test: port the remaining patch_manager tests with real coverage gaps

  - validate_next_boot_strict_mode_succeeds_with_valid_signature: ports
    `patch_manager.rs::validate_next_boot_patch_tests::strict_mode_succeeds_with_valid_signature_at_boot_time`.
    Required reusing the prior test fixture's INFLATED_PATCH_HASH /
    INFLATED_PATCH_SIGNATURE constants (a real signature for the
    1-byte file content `b"1"`, matching TEST_PUBLIC_KEY's private
    half).
  - record_boot_failure_keeps_last_booted_pointer_when_failed_patch_was_last_booted:
    ports
    `patch_manager.rs::record_boot_failure_for_patch_tests::preserves_last_booted_patch_on_failure_but_marks_bad`.
    The new behavior is similar but more explicit: last_booted's
    *pointer* is preserved (with the patch now in Bad state); the
    underlying intent — "we still know what last booted, even after
    that patch failed" — survives.

Test helper split: `install_signed` (failure paths, mismatched hash
OK) vs `install_with_valid_signature` (happy path, real fixture).

Now every behavioral test from the deleted patch_manager.rs has a
corresponding new test with a porting comment. Two trivial ones not
ported: the `Debug` impl tests for traits/types that no longer exist.

236 tests, all green under -D warnings.

* refactor: restore separate download directory under OS cache root

Reverts an unintended policy change in the cutover: my refactor moved
compressed download bytes from `{code_cache_dir}/downloads/{N}` (the
OS-managed cache dir, evictable under storage pressure) into
`{storage_dir}/patches/{N}/download` (persistent app storage). That
made downloads count against persistent storage and survive in iCloud
backups — neither of which we want for transient bytes.

Restoring the original split:
  - `{state_root}/patches/{N}/{state.json,dlc.vmcode}` — persistent
    state and the installed artifact.
  - `{download_root}/{N}` — flat, in the OS cache dir.

`PatchLifecycle::load_or_default` now takes both roots. The two
artifact-path helpers each route to their respective root, the
free-function variants used by `update_internal` now take the right
root, and `mark_bad`/`cleanup`/`record_install_complete` clean both
locations as appropriate.

`UpdaterState::create_new_and_save` also wipes `download_dir` on
release-version change, in addition to the persistent patches/ tree.

`update_internal` and tests updated to use the right root for each
file type. Test fixtures pass `tmp.path().join("downloads")` as the
download root (separate subdir of the same TempDir).

Net change in test surface: moved a handful of `patch_dir/download`
references to `downloads_dir/N`, no behavioral changes to tests
themselves. 236 tests pass under -D warnings.

* test: cover the new download_root cleanup paths

Self-review caught two gaps in the download-root split:

  - release_version_change_wipes_download_dir: the cache-rooted
    download dir should be wiped along with the persistent patches/
    tree on a release-version mismatch. Adds explicit coverage —
    the code path was already in create_new_and_save but not
    directly tested.
  - record_boot_success_promotes_and_cleans_older: extends the
    existing test to drop a stale download in download_root
    alongside an older patch's state, asserts that
    cleanup_older_than's chain (cleanup → forget_dir) deletes both
    roots' artifacts.

Also adds a comment on cleanup_older_than noting that it only
walks the persistent patches/ tree — orphan downloads (no
state.json) would persist until the OS evicts them. Noted as a
known limitation, not blocking.

* refactor: targeted wipe + legacy patches_state.json + orphan-download walk

Two fixes to the release-version-change wipe and one to the boot-
success cleanup walk.

  - The "wipe everything in cache_dir" approach broke 27 tests whose
    setup uses cache_dir as the test temp dir (with base.apk and
    libapp.so as siblings). Production gives shorebird a dedicated
    subdir but the API doesn't enforce that. Switched to a targeted
    wipe with a documented allowlist of paths shorebird has ever
    written under cache_dir: `SHOREBIRD_OWNED_PATHS`.
  - Added `patches_state.json` to the wipe list — that's the
    legacy file from the prior `PatchManager` implementation, and
    devices upgrading through this PR would otherwise orphan it.
    New test `release_version_change_wipes_legacy_patches_state_json`
    covers it.
  - `cleanup_older_than` now also calls `cleanup_orphan_downloads`,
    which walks `download_root/` and removes any file that doesn't
    correspond to a patch in `Downloading`/`Downloaded` state. We own
    the cache root and shouldn't rely on OS eviction. Catches the
    "state.json gone but download lingered" case the prior comment
    handwaved.

Adding a new file under cache_dir means adding it to
SHOREBIRD_OWNED_PATHS — small bookkeeping cost, but it lets us
keep cache_dir co-tenant with embedder-provided files.

* test: cover all four branches of cleanup_orphan_downloads

The new orphan-download sweep distinguishes four cases but only the
"older patch via cleanup chain" branch was being exercised. Single
test now drops one of each kind of file into download_root before a
boot success and asserts which survive:

  - orphan (numeric, no state.json): removed
  - stale (numeric, state is Installed): removed
  - non-numeric name: removed
  - live (numeric, state is Downloading): preserved

* ci: add 'embedder' to cspell dictionary

* refactor: drop Installed.hash, fix mark_bad ordering, address bdero review

bdero's review on #352 surfaced six items worth landing inline:

1. record_boot_failure used to clear `currently_booting_patch` before
   `mark_bad`. A crash between the two left the patch `Installed` with
   no breadcrumb, so `detect_boot_crash_on_init` wouldn't fire next
   boot — the patch silently retried. Flipped: mark_bad first, then
   clear. mark_bad on Bad is idempotent, so the worst case is a
   redundant tombstone-rewrite on next init.

2. Drop the `hash` field from `PatchState::Installed`. It was recorded
   at install but never read at boot — Strict-mode validation
   recomputes the hash from the artifact's bytes and feeds it into
   `check_signature`. A hash that lives only on disk can't be trusted
   as a security input; deleting the field removes the temptation.
   `Downloading.hash`/`Downloaded.hash` stay as comparators against
   the server's freshly-delivered hash (a tampered on-disk hash there
   just causes a redownload).

3. Port the customer-scenario test from #356 into `c_api/mod.rs`:
   server rolls patch 1 back, then forward again with the same number
   and hash. Pre-lifecycle this was permanently dropped via
   `known_bad_patches`. Post-lifecycle `cleanup` is state-aware and
   forgets the patch entirely on server-driven rollback, so the
   rollforward installs cleanly. End-to-end coverage of the regression
   #356 hotfixed.

4. Add a TODO + target version on the `patches_state.json` entry in
   `SHOREBIRD_OWNED_PATHS` so the legacy wipe doesn't outlive its
   usefulness.

5. Document that `running_patch` lives in `config.rs` (session-scoped),
   not in this module — saves a future grep.

6. Clarify `recompute_next_boot`'s fallback policy: it's
   fall-back-to-`last_booted_patch`-only, not fall-back-to-anything-
   bootable. A fresh release whose first Installed patch fails
   validation goes to base, even if older patches sit in `patches/`.

cargo test: 242 passed (was 241).
2026-05-07 16:01:28 -07:00
..
2023-03-06 15:06:31 -08:00

Shorebird CodePush Updater

The rust library that does the actual update work.

Design

The updater library is built in Rust for safety (and modernity). It's built as a C-compatible library, so it can be used from any language.

The library is thread-safe, as it needs to be called both from the flutter_main thread (during initialization) and then later from the Dart/UI thread (from application Dart code) in Flutter.

The overarching principle with the Updater is "first, do no harm". The updater should "fail open", terms of continuing to work with the currently installed or active version of the application even when the network is unavailable.

The updater also needs to handle error cases conservatively, such as partial downloads from a server, or malformed responses (e.g. a proxy interfering) and not crash the application or leave the application in a broken state.

Every time the updater runs it needs to verify that the currently installed patch is compatible with the currently installed base version. If it is not, it should refuse to return paths to incompatible patches.

The updater also needs to regularly verify that the current state directory is in a consistent state. If it is not, it should invalidate any installed patches and return to a clean state.

Not all of the above is implemented yet, but such is the intent.

Architecture

The updater is split into separate layers. The top layer is the C-compatible API, which is used by all consumers of the updater. The C-compatible API is a thin wrapper around the Rust API, which is the main implementation but only used directly for testing (see the cli directory).

Thread safety is handled by a global configuration object that is locked when accessed. It's possible I've missed cases where this is not sufficient, and there could be thread safety issues in the library.

  • c_api (module) - C-compatible API
    • c_file.rs - a read-seek interface usable by the engine (used to provide access to iOS patch files)
    • mod.rs - the implementation of the C API.
  • src/lib.rs - Rust API (and crate root)
  • src/updater.rs - Core updater logic
  • cache (module) - On-disk state management
    • disk-io.rs - Manages reading and writing serializable state to disk
    • mod.rs - Cache management
    • patch_manager.rs - Patch file state management. Owned by UpdaterState.
    • updater_state.rs - Public API for this module. Provides functions that trigger updates to the internal Boot State Machine (detailed below).
  • src/config.rs - In memory configuration and thread locking
  • src/cache.rs - On-disk state management
  • src/logging.rs - Logging configuration (for platforms that need it)
  • src/network.rs - Logic dealing with network requests and updater server

Rust

We use normal rust idioms (e.g. Result) inside the library and then bridge those to C via an explicit stable C API (explicit enums, null pointers for optional arguments, etc). This lets the Rust code feel natural and also gives us maximum flexibility in the future for exposing more in the C API without having to refactor the internals of the library.

https://docs.rust-embedded.org/book/interoperability/rust-with-c.html are docs on how to use Rust from C (what we're doing).

https://github.com/RubberDuckEng/safe_wren has an example of building in Rust and exposing it with a C api.

Integration

The updater library is built as a static library, and is linked into the libflutter.so as part of a custom build of Flutter. We also link libflutter.so with the correct flags such that updater symbols are exposed to Dart.

Building for Android

The best way I found was to install: https://github.com/bbqsrc/cargo-ndk

cargo install cargo-ndk
rustup target add \
    aarch64-linux-android \
    armv7-linux-androideabi \
    x86_64-linux-android \
    i686-linux-android
cargo ndk -t armeabi-v7a -t arm64-v8a build --release

When building to include with libflutter.so, you need to build with the same version of the ndk as Flutter is using:

You'll need to have a Flutter engine checkout already setup and synced. As part of gclient sync the Flutter engine repo will pull down a copy of the ndk into src/third_party/android_tools/ndk.

Then you can set the NDK_HOME environment variable to point to that directory. e.g.:

NDK_HOME=$HOME/Documents/GitHub/engine/src/third_party/android_tools/ndk

Then you can build the updater library as above. If you don't want to change your NDK_HOME, you can also set the environment variable for just the one call:

NDK_HOME=$HOME/Documents/GitHub/engine/src/third_party/android_tools/ndk cargo ndk -t armeabi-v7a -t arm64-v8a build --release

Imagined Architecture (not all implemented)

Assumptions (not all enforced yet)

  • Updater library is never allowed to crash, except on bad parameters from C.
  • Network and Disk are untrusted.
  • Running code is trusted.
  • Store-installed bundle is trusted (e.g. APK).
  • Updates are signed by a trusted key.
  • Updates must be applied in order.
  • Updates are applied in a single transaction.

State Machines

Boot State Machine

This state machine tracks the process of the engine starting up, with or without a patch. These state transitions happen whether or not a new patch is available, but in the context of the updater, we only care about the case where we are booting from a patch.

The patch boot state is internal to the PatchManager and stored on disk in patches_state.json. It contains three fields, all of which are Optional PatchMetadata structs:

  • last_boot_patch: The last patch that was successfully booted.
  • last_attempted_patch: The last patch that we attempted to boot.
  • next_boot_patch: The next patch that we will attempt to boot.

This state machine has the following states. It can only move forward through them.

  1. Ready - The engine is initialized but has not started booting yet.
  2. Booting - The engine has started booting.
  3. Success - The engine successfully booted.
  4. Failure - The engine failed to boot.

State is advanced through calls to the following methods of the ManagePatches trait, which PatchManager implements:

  • record_launch_start: Moves from Ready to Booting
    • next_boot_patch is the patch that we will attempt to boot, i.e., the "current" patch.
    • last_attempted_patch is set to next_boot_patch.
  • record_launch_success: Moves from Booting to Success
    • last_boot_patch is set to next_boot_patch.
    • Artifacts for patches older than last_boot_patch are deleted.
  • record_launch_failure: Moves from Booting to Failure
    • next_boot_patch artifacts are deleted.
    • next_boot_patch is set to either:
      • last_boot_patch if it is still valid, or
      • None, if last_boot_patch is None or invalid.

These are effectively no-ops if we are not booting from a patch.

Assumptions (not currently enforced, but should as possible):

  • This state machine will have advanced at least as far as the Booting state before the Patch Check State Machine (below) is started.
  • Calls to mutate state will not come out-of-order. For example, record_launch_failure will not be called before record_launch_start. This is important because PatchManager state is implicit - it does not track which state it is in.

Patch Check/Update State Machine

This state machine tracks the process of checking for new patches. It is managed by the code in updater.rs and does not have any on-disk state. It has the following states:

  1. Ready - Ready to check for updates.
  2. Send queued events (e.g., report that a patch succeeded or failed to boot) a. Move to checking once events, if any, have been reported.
  3. Checking for new patches - A PatchCheckRequest is issued but not completed. a. If no patch is available, move back to Ready. b. If a new patch is is available, move to Downloading Patch.
  4. Downloading Patch - A patch is available, we're downloading it. a. If the download fails, move back to Ready. b. If the download succeeds, move to Inflating Patch.
  5. Inflating Patch - A patch has been downloaded, we're inflating it. a. Attempt to inflate the patch (apply a bidiff to the current release). b. If the patch is valid, queue a PatchInstallSuccess event and move to Ready. c. If the patch is invalid, queue a PatchInstallFailure event and move to Ready.

Changes in this state can be triggered by:

  1. The engine via the C API.
  2. The user via the C API (using the shorebird_code_push package).
  3. Network activity.
  • Server is authoritative, regarding current update/patch state. Client can cache state in memory. Not written to disk.
  • Patches are downloaded to a temporary location on disk.
  • Client keeps on disk (in state.json):
    • Current release version. Set when the app launches. If the app is updated to a new release version, all state is invalidated.
    • Queue of PatchEvents. This is cleared once events are sent to our servers.

Trust model

  • Network and Disk are untrusted.
  • Running software (including apk service) is trusted.
  • Patch contents are signed, public key is included in the APK.

Patch Verification Modes

The updater supports two patch verification modes, configured via patch_verification in shorebird.yaml. Both modes require a patch_public_key to be configured for signature verification to occur.

Strict Mode (default)

patch_verification: strict

In Strict mode, patch signature verification happens at boot time. This provides the strongest security guarantee because it detects any potential on-disk tampering to the patch file after installation (e.g., if an attacker were to modify the patch file on disk between app launches). However the practical risk to such an attack is very low, since patches are stored within the app's protected storage. An attacker in this case would need to have already compromised the app itself, or the system (e.g. via a rooted device). However if an attacker has compromised the system (rooted) they could already modify the APK/IPA internals directly. The on-boot protection here is for cases where developers are concerned that their app might be compromised and they wish to ensure that such a compromise could not theoretically persist itself via editing an installed patch file. Such a case is impractical, but we default to the strongest-possible security stance regardless.

Strict mode is currently default for Shorebird, however some of our large customers requested that we add an install_only mode, since their applications were so large (many hundreds of mb) that the hash-verification during boot was showing up on profiles from older devices.

Install flow:

  1. Download patch from server
  2. Inflate patch (apply bidiff to base release)
  3. check_hash(): Compute SHA256 of inflated file, verify it matches server-provided hash
  4. Store patch file, hash, and signature to disk

Boot flow:

  1. Verify patch file exists and size matches stored metadata
  2. hash_file(): Re-compute SHA256 of patch file on disk
  3. check_signature(): Verify the computed hash has a valid signature using the public key
  4. If verification fails, fall back to last known good patch or base release

Install Only Mode

patch_verification: install_only

In Install Only mode, patch signature verification happens at install time only. This provides faster boot times but does not protect against post-install tampering (extremely uncommon). The only case that this does not protect against is if your app itself were to accidentally (or through some other malicious exploit of your app) modify its own data directory and modify the patch files within such.

Install flow:

  1. Download patch from server
  2. Inflate patch (apply bidiff to base release)
  3. check_hash(): Compute SHA256 of inflated file, verify it matches server-provided hash
  4. check_signature(): Verify the server-provided hash has a valid signature using the public key
  5. Store patch file, hash, and signature to disk

Boot flow:

  1. Verify patch file exists and size matches stored metadata
  2. (No signature verification - trusted from install time)

Without a Public Key

If no patch_public_key is configured, signature verification is skipped in both modes. The check_hash() step still runs during install to detect download corruption, but there is no cryptographic verification that the patch came from a trusted source.

TODO:

  • Add an async API.
  • Write tests for state management.
  • Make state management/filesystem management atomic (and tested).

Later-stage update system design docs