Adaptive Physical Communication System (APCS)

Design Decisions

A record of the major technical decisions: the context, the options considered, what was chosen and why, and the consequences. Each record uses the Architecture Decision Record (ADR) format. When you change one of these decisions, add a new record rather than silently editing an old one.

Back to the documentation index.


Index

# Decision Status
ADR-01 Physical channels only Accepted
ADR-02 Rateless fountain coding for Light and Sound Accepted
ADR-03 Replace robust soliton with cycled/dense/sparse rows and an exact decoder (APCF v3) Accepted, supersedes the v2 mapping
ADR-04 Raw QR byte mode instead of base64 text Accepted
ADR-05 Sparse QR codes and Auto density Accepted, supersedes the 800 B default
ADR-06 QR error-correction level L and fixed mask 0 Accepted
ADR-07 Camera tuned for screens: zoom, exposure, Y plane, single isolate Accepted
ADR-08 zxing2 decode sequence: GlobalHistogram → pureBarcode → Hybrid Accepted
ADR-09 Content-derived Light session IDs and Resume Accepted
ADR-10 Multi-tone FSK for sound Accepted, supersedes two-tone FSK
ADR-11 Reed-Solomon per frame with soft-decision erasures Accepted
ADR-12 Per-frame sync marker with leading-edge detection Accepted
ADR-13 Receiver auto-detects the Sound profile Accepted
ADR-14 No back-channel: sender streams until stopped Accepted
ADR-15 Extend, don’t overwrite: legacy modems stay Accepted
ADR-16 Single controller plus fine-grained notifiers Accepted
ADR-17 Headless simulators before hardware Accepted
ADR-18 Photos compressed to ≤ 120 KB; bundled demo samples Accepted
ADR-19 Automatic Gallery saving with gal Accepted
ADR-20 A Silent (near-ultrasonic) band for Sound Accepted, extends ADR-10 and ADR-13
ADR-21 Text and link envelopes without name or MIME type Accepted
ADR-22 Live frequency readout from the tone schedule and a throttled FFT Accepted, applies ADR-15 and ADR-16

ADR-01: Physical channels only

Context. The project’s goal is communication with no radio or network infrastructure: air-gapped, radio-silent, emergency and demonstration settings.

Decision. Only three channels exist: optical (screen → camera), acoustic (speaker → microphone) and vibration (motor → accelerometer). platform_capabilities.dart lists the excluded transports (Internet/IP, Wi-Fi, Bluetooth, NFC, cellular/SMS, cloud), and test/physical_only_test.dart checks that exactly three physical channels exist. The Android release manifest requests no INTERNET permission.

Consequences. Throughput is low compared with radio. The app can never “cheat” by falling back to a network, so every result is a genuine physical-channel result.


ADR-02: Rateless fountain coding for Light and Sound

Context. A screen can’t hear a camera and a speaker has no reliable return path. With a “send each chunk once, then repeat” carousel, a receiver that misses chunk 7 of 40 must wait a whole cycle to see it again. The expected time to collect all N chunks at loss rate p grows like a coupon-collector problem.

Options.

  1. Repeating carousel of numbered chunks (the old APCS1 text QR).
  2. Carousel plus per-chunk ACK over another channel (needs a return path).
  3. Fountain (LT) code: an endless stream of distinct encoded symbols, any K of which (plus a few) decode.

Decision. Option 3, systematic LT, shared by Light and Sound (lt_codec.dart).

Consequences.


ADR-03: Replace robust soliton with cycled/dense/sparse rows and an exact decoder

Context. The first fountain used the classic robust-soliton degree distribution with a peeling decoder. Field reports showed transfers “stuck at 1 / 3 symbols”. Measurement showed that at small K the robust soliton’s spike put about 47% of repair symbols on the same all-blocks equation (K = 3). Only 56% of random K+2 symbol sets for K = 1…10 were full rank, and a 3-block transfer completed from five distinct repair symbols only 69% of the time.

Decision (APCF v3).

Consequences. Mean overhead is 0–2.2 symbols across K = 1…600 at 40% loss; K = 3 always completes from 5 consecutive repair symbols. The cost is O(K²/32) word operations per symbol in the dense regime, which is trivial at the K values used (≤ ~1 000). See Fountain Code.


ADR-04: Raw QR byte mode instead of base64 text

Context. The first QR transport encoded binary chunks as base64 text, which costs 33% capacity and forces a denser QR for the same payload.

Decision. Encode APCF frames in true QR byte mode (QrCode.fromUint8List) and read the raw byte segment back from zxing2. A text form (FQR3: + base64url) remains only for the string-only web preview sampler.

Consequences. 25% fewer bytes inside each QR for the same payload (base64 turns 3 bytes into 4 characters), hence smaller QR versions and higher decode rates.


ADR-05: Sparse QR codes and Auto density

Context. The old default was 800 B per frame (QR v20, 97×97 modules). In the field it produced “DEC 2.0/s from CAP 18 fps” (≈11% decode rate).

Investigation. A camera simulator (test/optical_camera_sim.dart) rendered QR frames with perspective, blur, motion blur, washed-out contrast, glare, noise and a bezel. A density sweep with the production decoder gave, in the “typical” hand-held tier: 160 B 100%, 240 B 94%, 330 B 81%, 600 B 31%, 800 B 13%, 1 200 B 0%. The simulator reproduced the field report almost exactly.

Decision.

Consequences. Effective goodput rose even though each frame carries less, because far more frames decode. Worked example: 800 B × 12 fps × 13% ≈ 1.25 KB/s, against 160 B × 12 fps × 100% × a 0.7 refresh yield ≈ 1.34 KB/s. On top of that, the sparse code keeps working in the “hard” tier, where the dense one reads 0%.


ADR-06: QR error-correction level L and fixed mask 0

Context. Higher QR EC levels (M, Q, H) add Reed-Solomon parity inside each code, making it denser. Mask selection evaluates 8 masks with penalty rules.

Decision.

Consequences. The steady 12 fps cadence stays within budget; the next symbol is built while the current one is displayed.


ADR-07: Camera tuned for screens: zoom, exposure, Y plane, single isolate

Decision and rationale.

Choice Why
1280×720 (ResolutionPreset.high), YUV420 Enough pixels for v8–v12 at arm’s length; a cheap stream
1.5× zoom by default (chips 1×/1.5×/2×/3×) Zoom crops the sensor before it is downscaled, adding real pixels per module; the user can stay at 15–25 cm where every camera focuses
−0.7 EV exposure offset Screens are bright. A shorter exposure means less motion blur and no bloom
Centre focus and metering, tap to refocus The code is in the centre; lets the user recover focus
Use the Y plane directly It is already luminance, so there is no colour conversion; the crop is a row memmove
Centred crop of 98% of the short side Matches the on-screen aim brackets
One long-lived decode isolate, 1 in flight Keeps the UI thread free; with fountain symbols, dropping a frame while busy costs nothing

ADR-08: zxing2 decode sequence: GlobalHistogram → pureBarcode → Hybrid

Context. The sweep compared binarizers and hints. Random payload bytes occasionally form fake 1:1:3:1:1 finder patterns that outrank the true corners.

Decision.

  1. GlobalHistogramBinarizer first: the fastest, and it beat Hybrid in almost every sweep cell.
  2. The same bitmap with pureBarcode, which reads the grid from the black bounding box. It rescued 100% of the fake-finder failures (3–8% of frames).
  3. HybridBinarizer last, for uneven lighting.

tryHarder stays off: it’s cheaper to drop a frame and take the next symbol.


ADR-09: Content-derived Light session IDs and Resume

Decision. sessionId = (CRC32(envelope) ^ blockLen·0x9E3779B1) & 0xFFFFFFFF (0 → 1). The sender remembers the next unshown symbol index per session (last 16 sessions).

Consequences. Re-sending the same file at the same density continues the same fountain with new symbols. A receiver that got 80% of the way keeps its progress (it holds up to 3 partial sessions). Changing density creates a new session, as it must, because the block layout differs.


ADR-10: Multi-tone FSK for sound

Context. The legacy two-tone FSK (1 800 / 3 200 Hz, 18 ms per bit) delivered about 55.6 b/s raw, had no error correction, and assumed symbol boundaries at fixed offsets from the buffer start. Microphone chunks arrive at arbitrary offsets, so windows straddled tones.

Options. The three M-ary keying families carry data in amplitude, phase or frequency:

  1. M-ary ASK (amplitude levels). Rejected. The received level changes with distance (about 6 dB from 20 to 40 cm), volume, the user’s hand, the speaker’s frequency response and room echoes, which can notch single frequencies by 10–20 dB. The receiver can’t tell which level was sent.
  2. M-ary PSK (phase angles). Rejected. It needs coherent detection against a carrier reference that two phones don’t share:
    • The two phones’ sample clocks drift apart.
    • Multipath smears the phase.
    • At 19 kHz a wavelength is 1.8 cm, so moving the phone 1 cm shifts the phase by about 200°.
    • Hand motion at 0.1 m/s gives about 5.5 Hz of Doppler, which rotates the phase by about 90° over one 46 ms Silent symbol.
  3. M-ary FSK (which frequency is playing). Chosen.
    • It is detected non-coherently: the receiver compares energies at known bins with Goertzel filters, with no phase tracking, and a common gain change cancels out.
    • It has a constant envelope, so a small speaker can play it near full level without distortion.
    • Unlike ASK and PSK, whose points crowd together as M grows, orthogonal FSK needs less energy per bit as M grows. It pays in bandwidth, which the audio band has to spare, rather than in power, which a phone speaker lacks.

    The cost is lower spectral efficiency than PSK or QAM in a clean channel. OFDM remains on the roadmap for quiet rooms.

Decision. Multi-tone FSK in the style of ggwave: 6 or 8 simultaneous tones, each choosing 1 of 16 frequencies (4 bits). Every tone sits on an exact FFT bin of a 1 024-sample frame at 44.1 kHz (43.066 Hz spacing), so tones are orthogonal and phase-continuous. The band is about 1.2–7.2 kHz.

Consequences. Standard reaches 345 b/s raw and 27 B/s net after RS parity and sync, about 4× the old modem, with error correction. The cost is audible chords and a need for per-frame sync (ADR-12).


ADR-11: Reed-Solomon per frame with soft-decision erasures

Context. Even in a quiet room a few percent of acoustic bytes are wrong: multipath nulls wipe out a tone. With only a checksum, nearly every frame would be discarded.

Decision.

Consequences. The hardest simulated room went from decoding nothing to passing. The parity costs 27–39% of each frame’s bytes, which the measured byte error rates justify.


ADR-12: Per-frame sync marker with leading-edge detection

Decision.

Why. Scoring a single full window normalised by its own energy is scale-invariant, so a window holding the back half of the marker plus silence scored as well as an aligned one, and every frame locked one frame early. The two-half minimum fixes that. Reflections can be louder than the direct sound, but the direct sound always arrives first, hence the leading edge.


ADR-13: Receiver auto-detects the Sound profile

Decision. All profiles share the same marker. The receiver runs one frame-sync per profile. The first profile whose frame passes RS, the CRC and the header’s blockLen check becomes the lock. After 4 markers with no good frame the lock goes stale and all profiles are heard again. After each completed message the receiver reopens to all profiles.

Consequences. The receiver never has to be configured, and the sender can switch to Rugged in a noisy room with no coordination. CPU use is about four frame-syncs while hunting, then one. (ADR-20 adds two Silent profiles with their own marker, making six.)


ADR-14: No back-channel: sender streams until stopped

Decision.

Consequences. The user decides when to stop, based on the receiver’s DONE. Receivers ignore the sender’s “tail” for sessions they have completed. Sound sends are also marked sent; earlier builds marked them delivered (see Known Issues 2.1).


ADR-15: Extend, don’t overwrite: legacy modems stay

Decision. The CSK light modem, the APCS1 text QR codec, two-tone FSK and on/off light keying remain in the codebase with their tests. New modems are added alongside them.

Consequences. There are more files, but earlier work stays reusable and testable and nothing regresses silently. The legacy modems aren’t selectable in the current UI.


ADR-16: Single controller plus fine-grained notifiers

Decision. One AppController (ChangeNotifier) owns application state. Values that change at frame or audio rate (QR bitmap, HUD metrics, meters) use dedicated singleton notifiers, so only the widgets displaying them rebuild. The QR painter repaints only when its bitmap changes and snaps modules to whole device pixels.

Consequences. The UI stays at display rate during streaming, the state flow stays simple, and no external state-management package is needed.


ADR-17: Headless simulators before hardware

Decision. Every modem parameter was chosen from simulations that drive the production decoders:

Consequences. Design choices are reproducible and regression-tested (flutter test). Real devices still vary, so rehearsal on the target phones remains essential.


ADR-18: Photos compressed to ≤ 120 KB; bundled demo samples

Decision.

Consequences. Demos need no files on the phone. Photo transfer time stays bounded (≤ ~1 minute over Light). Samples are re-compressed at send time (see Known Issues).


Decision. Received photos and videos are saved automatically to the album Adaptive Comm via the gal plugin. Saving is idempotent per message ID and exposes Save / Saving… / Saved / Retry states. It needs no permission on Android 10+, WRITE_EXTERNAL_STORAGE (maxSdk 29) on older Android, and photo-library add permission on iOS.

Consequences. Users keep what they receive without extra taps. WebM can’t be saved to iOS Photos.


ADR-20: A Silent (near-ultrasonic) band for Sound

Context. Audible Sound transfers are noticeable and get drowned out by talking. A comparable app sends two-tone FSK at 19 and 20 kHz (60 ms per bit, XOR checksum), which people can’t hear and voices don’t reach. Android’s CDD defines near-ultrasound support for 18.5–20 kHz on the voice-recognition microphone source.

Options.

  1. Copy the two-tone design: simple, but 60 ms per bit, a 255-byte limit and no error correction.
  2. Move the existing chord modem up unchanged: fast on paper, but two simultaneous near-ultrasonic tones produce an audible difference tone through a small speaker’s non-linearity, which defeats the point.
  3. Keep the whole stack (frame format, Reed-Solomon, fountain, auto-detection) and give it a second tone plan.

Decision. Option 3. AcousticBand.nearUltrasonic configures the same MtFskCodec:

Consequences. Silent is about as robust as Rugged at half its speed, can’t be heard by most adults, and shrugs off chatter. It is slower than the audible profiles and needs phones within about half a metre. It only works on phones whose speaker and microphone pass 19 kHz, which the simulator can’t predict, so the Receive screen shows a dedicated Silent band meter. The codec’s defaults keep the audible plan bit-for-bit, and receivers listen for both bands at once. For a three-letter text, two-tone FSK is 0.5 s quicker (4.3 s against one 4.8 s frame); from about five characters on, Silent is faster, and it has Reed-Solomon and an 8 KiB limit. Details: Sound Channel §14.


Context. Every text used to carry the name message.txt and MIME text/plain (links: link.url, text/uri-list): 21 bytes that the type byte already implies. On Sound a 17-character text became a 45-byte envelope, two Silent frames instead of one.

Decision. encodeText and encodeLink leave both fields empty. overheadBytes() and ComposePayload.envelopeBytes report the real on-air size without building the envelope.

Consequences. Text overhead drops from 28 to 7 bytes: “sos” is 10 B, a 17-character text 24 B. Short texts fit one Sound frame on every profile, and Vibration saves about 40 s per message. Older receivers already read empty fields as absent, so they decode the new envelopes unchanged. Received texts show no file name, as before.


ADR-22: Live frequency readout from the tone schedule and a throttled FFT

Context. Users and demo audiences want to see, in Hz or kHz, what the sender is playing and what the receiver hears. That matters most for the Silent band, where there’s nothing to hear. The existing meters show levels, not frequencies.

Options for the sender.

  1. Analyse the WAV (or the microphone) on the sending phone. This costs an FFT per tick, and self-listening is impossible anyway because the microphone is off during playback.
  2. Record the schedule of tones while the burst is rendered, and look it up against the playback clock.

Options for the receiver.

  1. Run Goertzel on every tone bin of every profile. This only sees the modem’s own frequencies and says nothing about noise or a whistle.
  2. Run a full FFT on every microphone chunk. That is wasted work, because the screen updates at most every 80 ms.
  3. Keep the latest 1 024 samples in a ring and run one FFT when the UI tick fires.

Decision. Sender option 2 and receiver option 3.

Consequences.