A record of the major technical decisions: the context, the options considered, what was chosen and why, and the consequences. Each record uses the Architecture Decision Record (ADR) format. When you change one of these decisions, add a new record rather than silently editing an old one.
Back to the documentation index.
| # | Decision | Status |
|---|---|---|
| ADR-01 | Physical channels only | Accepted |
| ADR-02 | Rateless fountain coding for Light and Sound | Accepted |
| ADR-03 | Replace robust soliton with cycled/dense/sparse rows and an exact decoder (APCF v3) | Accepted, supersedes the v2 mapping |
| ADR-04 | Raw QR byte mode instead of base64 text | Accepted |
| ADR-05 | Sparse QR codes and Auto density | Accepted, supersedes the 800 B default |
| ADR-06 | QR error-correction level L and fixed mask 0 | Accepted |
| ADR-07 | Camera tuned for screens: zoom, exposure, Y plane, single isolate | Accepted |
| ADR-08 | zxing2 decode sequence: GlobalHistogram → pureBarcode → Hybrid | Accepted |
| ADR-09 | Content-derived Light session IDs and Resume | Accepted |
| ADR-10 | Multi-tone FSK for sound | Accepted, supersedes two-tone FSK |
| ADR-11 | Reed-Solomon per frame with soft-decision erasures | Accepted |
| ADR-12 | Per-frame sync marker with leading-edge detection | Accepted |
| ADR-13 | Receiver auto-detects the Sound profile | Accepted |
| ADR-14 | No back-channel: sender streams until stopped | Accepted |
| ADR-15 | Extend, don’t overwrite: legacy modems stay | Accepted |
| ADR-16 | Single controller plus fine-grained notifiers | Accepted |
| ADR-17 | Headless simulators before hardware | Accepted |
| ADR-18 | Photos compressed to ≤ 120 KB; bundled demo samples | Accepted |
| ADR-19 | Automatic Gallery saving with gal |
Accepted |
| ADR-20 | A Silent (near-ultrasonic) band for Sound | Accepted, extends ADR-10 and ADR-13 |
| ADR-21 | Text and link envelopes without name or MIME type | Accepted |
| ADR-22 | Live frequency readout from the tone schedule and a throttled FFT | Accepted, applies ADR-15 and ADR-16 |
Context. The project’s goal is communication with no radio or network infrastructure: air-gapped, radio-silent, emergency and demonstration settings.
Decision. Only three channels exist: optical (screen → camera), acoustic (speaker → microphone) and vibration (motor → accelerometer). platform_capabilities.dart lists the excluded transports (Internet/IP, Wi-Fi, Bluetooth, NFC, cellular/SMS, cloud), and test/physical_only_test.dart checks that exactly three physical channels exist. The Android release manifest requests no INTERNET permission.
Consequences. Throughput is low compared with radio. The app can never “cheat” by falling back to a network, so every result is a genuine physical-channel result.
Context. A screen can’t hear a camera and a speaker has no reliable return path. With a “send each chunk once, then repeat” carousel, a receiver that misses chunk 7 of 40 must wait a whole cycle to see it again. The expected time to collect all N chunks at loss rate p grows like a coupon-collector problem.
Options.
Decision. Option 3, systematic LT, shared by Light and Sound (lt_codec.dart).
Consequences.
Context. The first fountain used the classic robust-soliton degree distribution with a peeling decoder. Field reports showed transfers “stuck at 1 / 3 symbols”. Measurement showed that at small K the robust soliton’s spike put about 47% of repair symbols on the same all-blocks equation (K = 3). Only 56% of random K+2 symbol sets for K = 1…10 were full rank, and a 3-block transfer completed from five distinct repair symbols only 69% of the time.
Decision (APCF v3).
min(K/2, ⌈2 ln K⌉ + 8).Consequences. Mean overhead is 0–2.2 symbols across K = 1…600 at 40% loss; K = 3 always completes from 5 consecutive repair symbols. The cost is O(K²/32) word operations per symbol in the dense regime, which is trivial at the K values used (≤ ~1 000). See Fountain Code.
Context. The first QR transport encoded binary chunks as base64 text, which costs 33% capacity and forces a denser QR for the same payload.
Decision. Encode APCF frames in true QR byte mode (QrCode.fromUint8List) and read the raw byte segment back from zxing2. A text form (FQR3: + base64url) remains only for the string-only web preview sampler.
Consequences. 25% fewer bytes inside each QR for the same payload (base64 turns 3 bytes into 4 characters), hence smaller QR versions and higher decode rates.
Context. The old default was 800 B per frame (QR v20, 97×97 modules). In the field it produced “DEC 2.0/s from CAP 18 fps” (≈11% decode rate).
Investigation. A camera simulator (test/optical_camera_sim.dart) rendered QR frames with perspective, blur, motion blur, washed-out contrast, glare, noise and a bezel. A density sweep with the production decoder gave, in the “typical” hand-held tier: 160 B 100%, 240 B 94%, 330 B 81%, 600 B 31%, 800 B 13%, 1 200 B 0%. The simulator reproduced the field report almost exactly.
Decision.
Consequences. Effective goodput rose even though each frame carries less, because far more frames decode. Worked example: 800 B × 12 fps × 13% ≈ 1.25 KB/s, against 160 B × 12 fps × 100% × a 0.7 refresh yield ≈ 1.34 KB/s. On top of that, the sparse code keeps working in the “hard” tier, where the dense one reads 0%.
Context. Higher QR EC levels (M, Q, H) add Reed-Solomon parity inside each code, making it denser. Mask selection evaluates 8 masks with penalty rules.
Decision.
Consequences. The steady 12 fps cadence stays within budget; the next symbol is built while the current one is displayed.
Decision and rationale.
| Choice | Why |
|---|---|
1280×720 (ResolutionPreset.high), YUV420 |
Enough pixels for v8–v12 at arm’s length; a cheap stream |
| 1.5× zoom by default (chips 1×/1.5×/2×/3×) | Zoom crops the sensor before it is downscaled, adding real pixels per module; the user can stay at 15–25 cm where every camera focuses |
| −0.7 EV exposure offset | Screens are bright. A shorter exposure means less motion blur and no bloom |
| Centre focus and metering, tap to refocus | The code is in the centre; lets the user recover focus |
| Use the Y plane directly | It is already luminance, so there is no colour conversion; the crop is a row memmove |
| Centred crop of 98% of the short side | Matches the on-screen aim brackets |
| One long-lived decode isolate, 1 in flight | Keeps the UI thread free; with fountain symbols, dropping a frame while busy costs nothing |
Context. The sweep compared binarizers and hints. Random payload bytes occasionally form fake 1:1:3:1:1 finder patterns that outrank the true corners.
Decision.
GlobalHistogramBinarizer first: the fastest, and it beat Hybrid in almost every sweep cell.pureBarcode, which reads the grid from the black bounding box. It rescued 100% of the fake-finder failures (3–8% of frames).HybridBinarizer last, for uneven lighting.tryHarder stays off: it’s cheaper to drop a frame and take the next symbol.
Decision. sessionId = (CRC32(envelope) ^ blockLen·0x9E3779B1) & 0xFFFFFFFF (0 → 1). The sender remembers the next unshown symbol index per session (last 16 sessions).
Consequences. Re-sending the same file at the same density continues the same fountain with new symbols. A receiver that got 80% of the way keeps its progress (it holds up to 3 partial sessions). Changing density creates a new session, as it must, because the block layout differs.
Context. The legacy two-tone FSK (1 800 / 3 200 Hz, 18 ms per bit) delivered about 55.6 b/s raw, had no error correction, and assumed symbol boundaries at fixed offsets from the buffer start. Microphone chunks arrive at arbitrary offsets, so windows straddled tones.
Options. The three M-ary keying families carry data in amplitude, phase or frequency:
The cost is lower spectral efficiency than PSK or QAM in a clean channel. OFDM remains on the roadmap for quiet rooms.
Decision. Multi-tone FSK in the style of ggwave: 6 or 8 simultaneous tones, each choosing 1 of 16 frequencies (4 bits). Every tone sits on an exact FFT bin of a 1 024-sample frame at 44.1 kHz (43.066 Hz spacing), so tones are orthogonal and phase-continuous. The band is about 1.2–7.2 kHz.
Consequences. Standard reaches 345 b/s raw and 27 B/s net after RS parity and sync, about 4× the old modem, with error correction. The cost is audible chords and a need for per-frame sync (ADR-12).
Context. Even in a quiet room a few percent of acoustic bytes are wrong: multipath nulls wipe out a tone. With only a checksum, nearly every frame would be discarded.
Decision.
0x11D) with P = 20–24 parity bytes per frame.2e + f ≤ P).Consequences. The hardest simulated room went from decoding nothing to passing. The parity costs 27–39% of each frame’s bytes, which the measured byte error rates justify.
Decision.
2·min(P₂₈, P₃₆)/E. An ideal aligned marker scores exactly 512; the threshold is 8.Why. Scoring a single full window normalised by its own energy is scale-invariant, so a window holding the back half of the marker plus silence scored as well as an aligned one, and every frame locked one frame early. The two-half minimum fixes that. Reflections can be louder than the direct sound, but the direct sound always arrives first, hence the leading edge.
Decision. All profiles share the same marker. The receiver runs one frame-sync per profile. The first profile whose frame passes RS, the CRC and the header’s blockLen check becomes the lock. After 4 markers with no good frame the lock goes stale and all profiles are heard again. After each completed message the receiver reopens to all profiles.
Consequences. The receiver never has to be configured, and the sender can switch to Rugged in a noisy room with no coordination. CPU use is about four frame-syncs while hunting, then one. (ADR-20 adds two Silent profiles with their own marker, making six.)
Decision.
max(6K, K+24), with a progress target of ⌈1.25K⌉ + 2.Consequences. The user decides when to stop, based on the receiver’s DONE. Receivers ignore the sender’s “tail” for sessions they have completed. Sound sends are also marked sent; earlier builds marked them delivered (see Known Issues 2.1).
Decision. The CSK light modem, the APCS1 text QR codec, two-tone FSK and on/off light keying remain in the codebase with their tests. New modems are added alongside them.
Consequences. There are more files, but earlier work stays reusable and testable and nothing regresses silently. The legacy modems aren’t selectable in the current UI.
Decision. One AppController (ChangeNotifier) owns application state. Values that change at frame or audio rate (QR bitmap, HUD metrics, meters) use dedicated singleton notifiers, so only the widgets displaying them rebuild. The QR painter repaints only when its bitmap changes and snaps modules to whole device pixels.
Consequences. The UI stays at display rate during streaming, the state flow stays simple, and no external state-management package is needed.
Decision. Every modem parameter was chosen from simulations that drive the production decoders:
Consequences. Design choices are reproducible and regression-tested (flutter test). Real devices still vary, so rehearsal on the target phones remains essential.
Decision.
assets/samples/, generated reproducibly by tool/ scripts with two-pass encoding to exact byte budgets.Consequences. Demos need no files on the phone. Photo transfer time stays bounded (≤ ~1 minute over Light). Samples are re-compressed at send time (see Known Issues).
galDecision. Received photos and videos are saved automatically to the album Adaptive Comm via the gal plugin. Saving is idempotent per message ID and exposes Save / Saving… / Saved / Retry states. It needs no permission on Android 10+, WRITE_EXTERNAL_STORAGE (maxSdk 29) on older Android, and photo-library add permission on iOS.
Consequences. Users keep what they receive without extra taps. WebM can’t be saved to iOS Photos.
Context. Audible Sound transfers are noticeable and get drowned out by talking. A comparable app sends two-tone FSK at 19 and 20 kHz (60 ms per bit, XOR checksum), which people can’t hear and voices don’t reach. Android’s CDD defines near-ultrasound support for 18.5–20 kHz on the voice-recognition microphone source.
Options.
Decision. Option 3. AcousticBand.nearUltrasonic configures the same MtFskCodec:
groups = 1) in 18.3–19.9 kHz, on every second bin (86 Hz apart) so hand-held Doppler doesn’t blur neighbours.Consequences. Silent is about as robust as Rugged at half its speed, can’t be heard by most adults, and shrugs off chatter. It is slower than the audible profiles and needs phones within about half a metre. It only works on phones whose speaker and microphone pass 19 kHz, which the simulator can’t predict, so the Receive screen shows a dedicated Silent band meter. The codec’s defaults keep the audible plan bit-for-bit, and receivers listen for both bands at once. For a three-letter text, two-tone FSK is 0.5 s quicker (4.3 s against one 4.8 s frame); from about five characters on, Silent is faster, and it has Reed-Solomon and an 8 KiB limit. Details: Sound Channel §14.
Context. Every text used to carry the name message.txt and MIME text/plain (links: link.url, text/uri-list): 21 bytes that the type byte already implies. On Sound a 17-character text became a 45-byte envelope, two Silent frames instead of one.
Decision. encodeText and encodeLink leave both fields empty. overheadBytes() and ComposePayload.envelopeBytes report the real on-air size without building the envelope.
Consequences. Text overhead drops from 28 to 7 bytes: “sos” is 10 B, a 17-character text 24 B. Short texts fit one Sound frame on every profile, and Vibration saves about 40 s per message. Older receivers already read empty fields as absent, so they decode the new envelopes unchanged. Received texts show no file name, as before.
Context. Users and demo audiences want to see, in Hz or kHz, what the sender is playing and what the receiver hears. That matters most for the Silent band, where there’s nothing to hear. The existing meters show levels, not frequencies.
Options for the sender.
Options for the receiver.
Decision. Sender option 2 and receiver option 3.
MtFskCodec.describe writes a ToneTimeline using the same _markerBins and _symbolBins helpers as encode, so the readout can’t drift from the audio. The modem builds a timeline only when transmit is given onBurst.SpectrumAnalyzer uses a Hann window, a radix-2 FFT with precomputed twiddles, peaks at least 18 dB above the median bin with leakage suppression, and parabolic interpolation.acousticSpectrumState notifier (ADR-16) to one reusable LiveToneMeter with .sending() and .hearing() variants.Consequences.
AudioPlayer.play() returns, so output latency makes it run slightly ahead of the audio (see Known Issues).