Project

sendon.live

A security-first, peer-to-peer file sharing platform

sendon, A Peer-to-Peer File Sharing Platform

sendon is available at sendon.live.

sendon is a security-first, peer-to-peer file sharing platform.

It is designed for both everyday file sharing and large-scale transfers, sendon handles everything from small files around 1 MB to files as large as 50 GB, without requiring cloud storage as an intermediary.

At its core, sendon is a binary data transfer system. Files, messages, frames, and protocol structures are represented and transferred as binary data, providing a consistent foundation for transfers of any size.

Sendon’s Internals

Protocol

sendon treats everything as a frame. Every message; a metadata announcement, a chunk of file bytes, an acknowledgement is simply a frame which is a byte array with a type tag, that gets encrypted before it’s sent and decrypted after it’s received.

Frame type  →  [type byte][JSON or binary payload]  →  encrypt  →  send over data channel

The receiver never knows what kind of frame arrived until it decrypts it.

Protocol Envelopes

There are two different encrypted formats used at different points in a transfer’s lifetime.

Bootstrap Envelope

The first one is the Bootstrap envelope which is constructed during handshake.

┌──────────┬────────────────────┬──────────────────────┬────────────┐
│ version  │  AES-GCM nonce     │  ciphertext          │  auth tag  │
│ (1 byte) │  (12 bytes,        │  ([type byte][JSON]) │ (16 bytes) │
│          │   random)          │                      │            │
└──────────┴────────────────────┴──────────────────────┴────────────┘
Transfer Envelope

The second one is the Transfer envelope which is constructed for every frame that has been sent.

┌──────────┬──────────────────────┬────────────┐
│ version  │  ciphertext          │  auth tag  │
│ (1 byte) │  ([type byte][data]) │ (16 bytes) │
└──────────┴──────────────────────┴────────────┘

Since the underlying WebRTC Data Channel is configured with ordered=true frame N is guarenteed to arrive before frame N+1 resulting the receiver will decrypt the respective frame N. This ordering guarantee is also what lets the transfer envelope leave its sequence number implicit instead of transmitting it. Therefore each side keeps its own counter and their counters holds the same value without ever exchanging the counter value itself.

To encrypt frame N, the sender builds nonce N and increments its counter.

To decrypt, the receiver builds the nonce it expects the same way, and only advances its counter once decryption succeeds.

A frame that was replayed, dropped, or reordered will fail AEAD authentication rather than silently decrypting with the wrong nonce, because the receiver is always trying a specific, predicted nonce, not reading one out of the message.

Per-transfer cryptographic state

The protocol uses several independent values to keep cryptographic state separated at different scopes.

MechanismSeparatesScopePrevents
transferIdOne transfer from another and identifies the handshake being acknowledgedOnce per transferAccidental association of an acknowledgement with the wrong transfer; contributes to deriving distinct transfer keys
receiverTransferSalt (feeds HKDF)This transfer’s derived keys from other transfers using the same linkKeyOnce per transferAccidental transfer-key reuse, including defense against a reused transferId
direction string (feeds HKDF)Sender → receiver traffic from receiver → sender trafficTwo derivations per transferReusing the same AES-GCM key/nonce pair across opposite directions when both sequence counters begin from the same value
sequence counter (feeds AES-GCM nonce)One encrypted frame from every other frame under the same directional keyOnce per frameAES-GCM nonce reuse within a directional transfer key; also causes replayed or unexpectedly reordered frames to fail authentication
transferId replay cacheA newly proposed transfer from bootstrap metadata already accepted on the same connectionOnce per accepted transferReplaying an old bootstrap metadata frame within the lifetime of the WebRTC connection

The important distinction is that receiverTransferSalt and the AES-GCM frame nonce are not the same thing.

The receiver checks the canonicalized transferId against the replay cache before opening the incoming destination. It records the ID only after the destination and directional transfer ciphers are successfully established, so a transient local setup failure does not permanently reject a legitimate retry. The cache is cleared when the WebRtcDirectConnection closes.


Frames

There are seven frame types.

ENUM FrameType
    metadata = 1
    chunk = 2
    completion = 3
    acknowledgement = 4
    metadataAcknowledgement = 5
    checkpoint = 6
    checkpointAcknowledgement = 7
END ENUM

metadata is used during the handshake to tell the receiver what the sender intends to transfer, including file name, size, media type, protocol version, and transferId.

metadataAcknowledgement tells the sender that the receiver accepted the transfer proposal and returns the receiver’s handshake contribution and negotiated capabilities.

chunk carries the actual file bytes after the handshake has completed.

checkpoint tells the receiver that the sender has transmitted up to a particular byte position.

checkpointAcknowledgement tells the sender that the receiver has successfully processed the corresponding checkpoint.

completion tells the receiver that the sender has finished transmitting the payload and includes the final size and optional SHA-256 digest.

acknowledgement is sent by the receiver after final verification to confirm the completed transfer.

TL;DR:

  • metadata / metadataAcknowledgement are used for the transfer handshake and key derivation negotiation.
  • checkpoint / checkpointAcknowledgement are used for flow control and pacing so the sender does not run too far ahead of the receiver.
  • completion / acknowledgement are used for end-to-end transfer verification, confirming that the received plaintext matches what the sender transmitted.

That is the entire protocol vocabulary: seven frame types.

Except for chunk, every frame payload is JSON serialized and UTF-8 encoded before encryption.

After encryption, every frame is transmitted as binary data over WebRTC.

                      WebRTC Data Channel
                             │
             ┌───────────────┴────────────────┐
             │                                │
   STRUCTURED (JSON)                        chunk
   metadata, metadataAck,              (raw file bytes,
   checkpoint, checkpointAck,           no JSON step)
   completion, acknowledgement                │
             │                                │
       JSON encode                            │
             │                                │
       UTF-8 encode                           │
             │                                │
             └───────────────┬────────────────┘
                             │
                           frame
                             │
                           encrypt
                             │
                           binary

Ciphers

Sendon uses two kinds of ciphers: Bootstrap Cipher and Transfer Cipher.

Bootstrap Cipher

Bootstrap Cipher is used during the handshake process and uses the shared link key generated by sender before sending a file. It used for bootstrap frames these are metadata frame and its metadataAcknowledgement reply.

Characteristics:

  • Algorithm: AES-256-GCM.
  • Key: the raw link key from the share link fragment, used as-is, with no derivation.
  • Scope: bootstrap frames only. Once the acknowledgement is received and verified, the bootstrap cipher is retired and never used again for the remainder of the transfer.
  • Purpose: it exists solely to protect the metadata handshake (transfer ID, file info, and the salt exchange) long enough for both sides to derive the keys that will protect everything else.

Transfer Cipher

This cipher is what replaces the bootstrap cipher for everything after the metadata handshake. It uses the shared link key as an input to generate new key pairs. It generates two keys using HKDF-SHA256 function; sender and receiver do not encrypt with the same key.

  • Algorithm: AES-256-GCM, same as the bootstrap cipher.
  • Key: not the link key directly. Each side derives its own pair of 256-bit keys via HKDF-SHA256, using the link key as input material, salted with transferId + transferSalt, and tagged by direction (sender→receiver / receiver→sender).
  • Two keys, not one: because there’s a separate key per direction, sender and receiver never encrypt with the same key, eliminating any chance of both sides’ nonce counters colliding.
  • Nonce: a locally maintained counter, starting at zero, incremented once per encryption. It is never transmitted, both sides derive it implicitly from their position in the stream, which the reliable/ordered data channel guarantees stays in sync. This is safe specifically because each key is used only within one transfer and one direction, so the counter never needs to persist or resume.
  • Scope: everything after the bootstrap exchange, for the lifetime of this transfer only. A new transfer (even between the same two peers) always gets a fresh transfer ID and salt, and therefore entirely new transfer cipher keys.

Handshake

  • The sender creates a “room” via a signaling server, optionally protected by a password.
  • The sender generates a random 256-bit transfer key (the “link key”) locally. This key never touches the signaling server.
  • The sender builds a shareable link: https://sendon.live/receive/{roomId}#key={base64url(transferKey)}.
  • The key is in the URL fragment (#), not a query parameter, fragments are never sent to servers by browsers, so the signaling server only ever sees the room ID, never the key.
  • The receiver opens the link, extracts the room ID (to join the room via signaling) and the key (from the fragment, entirely client-side).
Transport Establishment
  • Sender and receiver use the signaling server to exchange WebRTC connection info (SDP/ICE) and establish a direct peer-to-peer data channel.
  • The channel is reliable and ordered, this matters later, since the encryption layer relies on both sides staying in lockstep on an implicit sequence counter without transmitting it on the wire.

Sender side

The sender generates a transfer ID (random 16 bytes, unique per transfer), builds the metadata frame, it typically looks like the JSON below before encryption:

{
  "protocol": "<protocol version>",
  "transferId": "<encoded 16-byte transfer ID>",
  "connectionPath": "<connection path>",
  "name": "<file/data name>",
  "size": "<payload size>",
  "mediaType": "<media type>"
}

then prefixes it with a one-byte frame type marker to form the plaintext frame:

┌──────────────┬──────────────────────────────┐
│ Frame Type   │ JSON Payload                 │
│ 1 byte       │ UTF-8 bytes                  │
├──────────────┼──────────────────────────────┤
│    0x01      │ 7B 22 70 72 6F ...           │
└──────────────┴──────────────────────────────┘

and finally encrypts the framed payload with the bootstrap cipher and sends it to the corresponding receiver.

connectionPath is included so each side can compare its own view of how the connection is routed (e.g. direct vs. relayed) against the peer’s, since neither side alone has full visibility into the negotiated path, and this reconciled value later helps interpret performance/telemetry for the transfer.

Receiver side

Receiving and validating the metadata

The receiver decrypts the incoming metadata frame using the bootstrap cipher, strips the one-byte frame prefix, and decodes it. It validates name, size, mediaType, and protocol. The receiver expects the sender to send a metadata frame first, so nothing else is a legal first message; anything else at this stage is treated as a protocol violation, since the receiver hasn’t yet derived any per-transfer keys to make sense of anything else.

Once decrypted, the metadata (file name, size, media type, protocol version, transfer ID) is treated as untrusted input from the sender and validated before anything is acted on: the protocol version must match, the size must be within bounds, and the transfer ID must not be one already used in this session. This last check exists because the bootstrap cipher can prove a message came from the holder of the link key, but it can’t by itself prove that message hasn’t been replayed.

Committing to the transfer

After validation, the receiver commits local resources: it opens wherever the incoming data will be written, and prepares to track integrity of the stream as it arrives.

Contributing to key derivation

The bootstrap key was only ever meant to authenticate the handshake itself; it’s never used to protect the actual payload. Instead, once the receiver has accepted the transfer, both sides derive a fresh pair of keys specifically for that transfer, using HKDF.

The secret input to that derivation is still the link key. The receiver has no other shared secret with the sender, so everything downstream has to trace back to it. But deriving directly from the link key alone would mean every transfer over the same link produces identical keys, which is undesirable: it would let anyone comparing ciphertexts across transfers notice they share key material, and it would mean a key-recovery on one transfer compromises every other transfer ever made over that link.

To prevent that, the derivation is salted with two values unique to this specific transfer: the transfer ID the sender chose, and a random value the receiver generates itself.

salt = transferId || receiverTransferSalt

Combining a value from each side means neither participant unilaterally controls the salt. The sender can’t force key reuse by replaying a transfer ID, because the receiver’s contribution is fresh and unpredictable each time, and the receiver’s contribution alone isn’t enough either, since it’s paired with the sender’s ID to scope it to this exact transfer.

The derivation produces two distinct outputs rather than one, distinguished by an info label naming the direction of travel:

key_sender_to_receiver = HKDF(secret = linkKey, salt = salt, info = "sender-to-receiver")
key_receiver_to_sender = HKDF(secret = linkKey, salt = salt, info = "receiver-to-sender")

This means a message encrypted going one way could never be replayed back in the other direction and still decrypt correctly; the two data streams are cryptographically independent even though they’re derived from the same transfer secret and the same salt. Each side then assigns these two keys to its own “outgoing” and “incoming” roles based on whether it’s acting as sender or receiver.

The local sequence counter

Each derived key still needs a per-message nonce for the AEAD cipher to remain secure: reusing a nonce under the same key is what breaks these ciphers. Rather than transmitting a nonce or sequence number alongside every encrypted frame, each side simply keeps an internal counter: it starts at zero for a freshly derived key, and increments by exactly one for every message that key encrypts or decrypts, in order.

nonce = pad(counter)      // counter starts at 0, increments per message
counter += 1

This works only because the underlying channel is reliable and ordered, messages can’t arrive out of sequence or go missing silently, so both sides’ counters are guaranteed to stay in lockstep without ever needing to say the number out loud. The payoff is that every encrypted frame is a few bytes smaller, and there’s no explicit sequence field for an attacker to tamper with. The cost is that this design is only safe on a transport that guarantees order and delivery, if messages could be dropped or reordered, the two sides’ counters would drift apart and legitimate frames would fail to decrypt.

Acknowledging

The receiver sends its acknowledgement back: still under the bootstrap key, since the sender hasn’t yet learned the receiver’s contribution and thus can’t have derived the new transfer keys itself. This acknowledgement is the vehicle that carries the receiver’s random contribution back to the sender, along with confirmation of the transfer ID, protocol version, and acceptance, so the sender can perform the same key-derivation step and arrive at matching keys independently. It typically looks like the JSON below before encryption:

{
  "accepted": true,
  "protocol": "<protocol version>",
  "transferId": "<echoed 16-byte transfer ID>",
  "receiverTransferSalt": "<encoded 16-byte receiver salt>",
  "connectionPath": "<connection path>",
  "blockBytes": "<optional negotiated block size>"
}

then prefixes it with a one-byte frame type marker to form the plaintext frame:

┌──────────────┬──────────────────────────────┐
│ Frame Type   │ JSON Payload                 │
│ 1 byte       │ UTF-8 bytes                  │
├──────────────┼──────────────────────────────┤
│    0x05      │ 7B 22 61 63 63 ...           │
└──────────────┴──────────────────────────────┘

and finally encrypts the framed payload with the bootstrap cipher and sends it back to the sender.

From that acknowledgement onward, the bootstrap key’s job is done for this transfer, every subsequent message in both directions is encrypted with the transfer cipher.

Sender receives the acknowledgement

The sender has been waiting on this acknowledgement since it sent the metadata frame, under a timeout (30 seconds), if the receiver never responds (rejects, disconnects, or simply takes too long), the sender gives up on the transfer rather than hanging indefinitely.

Once it arrives, the sender treats the acknowledgement’s transferId as untrusted and checks it against the one it originally generated. This guards against a stale or mismatched acknowledgement being applied to the wrong in-flight transfer.

With that confirmed, the sender now has everything it needs to derive the same pair of transfer keys the receiver did, its own transferId plus the receiverTransferSalt it just received, and performs the identical HKDF derivation independently:

key_sender_to_receiver = HKDF(secret = linkKey, salt = transferId || receiverTransferSalt, info = "sender-to-receiver")
key_receiver_to_sender = HKDF(secret = linkKey, salt = transferId || receiverTransferSalt, info = "receiver-to-sender")

The sender assigns these to the mirror image of the receiver’s roles: its outgoing cipher is sender-to-receiver, its incoming cipher is receiver-to-sender, the opposite pairing from what the receiver assigned. Because both sides fed the derivation identical inputs, they arrive at identical keys without ever transmitting them.

From here the bootstrap phase is over. The sender adopts whatever block size the receiver negotiated in the acknowledgement, sets up a bulk data channel if the receiver indicated it supports one, and begins streaming the payload, now entirely under the freshly derived, direction-specific transfer keys.

Transfer Phase

Streaming the payload

Once the sender has accepted the metadata acknowledgement and derived its transfer ciphers, it starts reading the payload as a stream. It does not load the entire file into memory. Instead, it reads a bounded amount, updates the running SHA-256 digest with those plaintext bytes, and turns the bytes into a chunk frame.

Unlike the control frames, a chunk payload is not JSON encoded. The bytes read from the source are placed directly after the one-byte frame type marker:

┌──────────────┬──────────────────────────────────────┐
│ Frame Type   │ Raw Payload Bytes                    │
│ 1 byte       │ No JSON or UTF-8 conversion          │
├──────────────┼──────────────────────────────────────┤
│    0x02      │ 00 4A FF 18 2C ...                   │
└──────────────┴──────────────────────────────────────┘

That entire plaintext frame, including the 0x02 type byte, is encrypted with the sender-to-receiver transfer cipher. The AES-GCM nonce is produced from that cipher’s current local sequence counter, so no nonce or sequence field is added to the transmitted transfer envelope.

The block size used to read and encrypt the payload is the size agreed during the metadata acknowledgement. The baseline path uses 12 KiB blocks. A receiver that supports the larger-block path can advertise a 256 KiB block size and establish a separate reliable, ordered bulk data channel for the transfer.

A larger logical block is still not handed to the underlying WebRTC implementation as one oversized message. After encryption, sendon splits it into physical fragments small enough to stay below conservative Data Channel message limits. Each fragment carries a one-byte fragment marker which says whether it begins or ends the encrypted block:

                 one encrypted chunk frame
                            │
                            ▼
             ┌────────┬────────┬────────┐
             │ start  │ middle │  end   │
             │fragment│fragment│fragment│
             └────────┴────────┴────────┘
                  WebRTC Data Channel

These fragment markers are transport framing, not new protocol frame types. The receiver collects them in order, rejects orphaned, nested, or oversized fragment sequences, and only passes the reconstructed ciphertext to the transfer cipher after the final fragment arrives. Authentication therefore applies to the original complete encrypted frame, not to each transport fragment independently.

Whether a logical block travels as one Data Channel message or several physical fragments does not change what the transfer layer sees after reassembly:

raw payload bytes
      │
      ▼
[0x02][payload bytes]
      │
      ▼
AES-256-GCM with sender-to-receiver key and next sequence nonce
      │
      ▼
transfer envelope
      │
      ├── sent directly when already small enough
      └── physically fragmented and reassembled when necessary
Receiving each chunk

The receiver processes encrypted frames strictly in the order delivered by the reliable, ordered channel. For each reconstructed transfer envelope, it builds the nonce expected from its own incoming sequence counter and attempts AES-GCM decryption with the sender-to-receiver key.

Only authenticated plaintext is interpreted. If authentication succeeds, the receiver advances its incoming counter, reads the first plaintext byte as the frame type, and treats the remaining bytes as the raw chunk. If authentication fails, the frame is not accepted and the counter is not advanced.

Before writing anything, the receiver checks that appending the chunk would not exceed the size declared in the metadata. It then feeds the plaintext bytes into its own running SHA-256 calculation and adds them to a bounded write buffer. Compatible small network chunks are coalesced before being written to the destination, avoiding one filesystem or browser-storage operation for every Data Channel message while still keeping memory use bounded.

At this point two independent transcripts are being built from the same plaintext stream:

sender                                         receiver
──────                                         ────────
read source bytes                              decrypt authenticated chunk
      │                                                   │
      ├── update sender SHA-256                           ├── update receiver SHA-256
      │                                                   │
      └── encrypt and transmit                            └── buffer and write to destination

The sender’s digest describes exactly what it read. The receiver’s digest describes exactly what it authenticated and accepted. They are not compared until the payload has completely arrived.

Memory and storage boundaries

Streaming at the protocol layer and storing the complete payload are two different decisions. sendon always moves data through bounded chunks, even when the complete source or destination happens to be backed by memory. It never hands an entire large file to WebRTC as one message.

On the sending side, the source presented to the transfer protocol can be either a stable byte array or a storage-backed stream:

  • A native app may snapshot a supported single file into memory when it is no larger than 128 MiB. This gives the transfer a stable source even if the original file changes after selection.
  • A native single file larger than 128 MiB stays on disk and is read incrementally.
  • A browser does not duplicate a selected File into a complete Sendon-owned byte array. It reads the browser-managed file in 1 MiB slices and feeds those slices into the bounded transfer pipeline.
  • A generated folder or multi-file ZIP is always storage-backed, regardless of its size. Native builds stream it from a temporary file; Web builds stream it from browser-managed staging storage.
  • The command-line client always opens the source file as a filesystem stream and never snapshots the complete file into memory.

Whichever source is used, the transfer layer reshapes the incoming source stream into the negotiated logical block size. At any one moment it needs only the current source slice, logical plaintext block, encrypted block, physical Data Channel fragments, running digest state, and bounded WebRTC queue. Those working buffers are temporary and do not grow with the total file size.

The receiver chooses its complete-payload backing before acknowledging the metadata, because the advertised size is already known at that point. The current boundaries are inclusive: a payload exactly equal to the threshold uses memory, while a payload one byte larger uses storage streaming.

ReceiverComplete payload kept in memoryAbove the boundary
Android or iOS appUp to 64 MiBStream to device storage
Web browserUp to 256 MiBStream to browser OPFS staging
Native desktop appUp to 512 MiBStream to device storage
Command-line clientNeverAlways stream to the selected directory

For an in-memory destination, the receiver allocates a byte array sized to the declared payload and copies authenticated plaintext chunks into it. The complete array remains in process memory until the size and SHA-256 checks succeed and the platform save flow takes ownership of it. If the preferred allocation fails, receivers that have a storage-backed path fall back to streaming rather than failing the transfer immediately.

For a storage-backed destination, authenticated chunks are not retained until completion. The receiver coalesces compatible small chunks into a bounded 1 MiB write buffer, writes that buffer to a hidden or pending destination, clears it, and repeats:

authenticated plaintext chunks
              │
              ▼
      bounded 1 MiB write buffer
              │ full or checkpoint
              ▼
      hidden/pending storage entry
              │ size + SHA-256 verified
              ▼
         visible final file

On Android, a large transfer is streamed into a pending MediaStore download and is published under Downloads/Sendon only after verification. On iOS and native desktop filesystems, it is streamed into a hidden .sendon-*.sendon-part file in the destination directory and atomically renamed to a collision-free final name after verification. The command-line receiver uses this same partial-file and atomic-rename model for every payload size.

On Web, a large transfer is streamed into a hidden entry in the browser’s Origin Private File System. After verification, sendon obtains a disk-backed browser File from that entry and starts the normal browser download. OPFS is private browser storage, not the user’s visible Downloads directory; the browser performs the final visible download itself.

“Streams to storage” does not mean that no memory is used. AES-GCM input and output, fragment reassembly, SHA-256 state, the receiver write buffer, platform bridge copies, and WebRTC’s internal queues still occupy memory temporarily. The important boundary is that these allocations remain bounded independently of total payload size: receiving a 50 GiB file through a storage-backed destination does not require a 50 GiB application buffer.

Keeping the sender bounded

Reading from disk is often faster than encrypting, transmitting, receiving, and writing the same bytes. If the sender kept submitting frames without observing the Data Channel, memory could grow with the entire unsent remainder of a large file.

sendon therefore applies backpressure before adding more encrypted data. It observes the Data Channel’s buffered amount and stops submitting frames when the bounded send window is full. It resumes after the channel reports that its buffered amount has drained below the configured threshold. This limits queued data without forcing the sender to wait for one network round trip after every chunk.

Backpressure only proves that the local WebRTC implementation has accepted or drained queued messages. It does not prove that the receiver has decrypted or persisted the corresponding plaintext. That stronger confirmation is the job of checkpoints.

Checkpoints

After each 8 MiB of payload, and again when the final non-empty portion has been sent, the sender pauses the stream and sends a checkpoint frame. Its JSON payload records the cumulative number of plaintext bytes submitted so far:

{
  "sent": 8388608
}

The frame is prefixed with the checkpoint type byte (0x06) and encrypted with the sender-to-receiver transfer cipher. It therefore consumes the next sequence value in the same outgoing cryptographic stream as the chunk frames around it.

When the receiver reaches that checkpoint, all preceding chunk frames have already been decrypted and processed because the channel is ordered. The receiver compares sent with its own cumulative plaintext byte count.

If the values match, the receiver flushes its pending write buffer to the destination and replies with a checkpointAcknowledgement frame:

{
  "received": 8388608
}

This reply is prefixed with type 0x07 and encrypted with the receiver-to-sender transfer cipher. That reverse-direction cipher has its own independent sequence counter; acknowledgement traffic can never reuse the sender-to-receiver key or nonce sequence.

The sender waits for this reply under a 30-second timeout and requires received to equal the byte position it announced. Only then does it continue reading and sending the next part of the payload. A mismatch means the peers disagree about the plaintext transcript and the transfer fails immediately.

The checkpoint exchange provides bounded end-to-end pacing:

Sender                                           Receiver
  │                                                  │
  │──── encrypted chunk frames up to 8 MiB ─────────▶│
  │                                                  │ decrypt, hash, buffer/write
  │──── checkpoint { sent: 8388608 } ───────────────▶│
  │                                                  │ compare byte count
  │                                                  │ flush pending writes
  │◀── checkpointAck { received: 8388608 } ─────────│
  │                                                  │
  │──── continue with the next payload bytes ───────▶│

Checkpoints do not replace the final SHA-256 comparison. They detect byte-count disagreement and prevent the sender from running too far ahead, while the digest later verifies the contents of the complete plaintext transcript.

Completing and verifying the transfer

After the source stream ends, the sender first verifies that the number of bytes it read equals the size it declared in the metadata. A source that ended early or produced extra data is rejected as a payload-size mismatch.

The sender then finalizes its SHA-256 digest and sends a completion frame. It typically looks like this before framing and encryption:

{
  "size": 12582912,
  "sha256": "<encoded SHA-256 digest>",
  "connectionPath": "<reconciled connection path>"
}

The JSON is UTF-8 encoded, prefixed with the completion type byte (0x03), and encrypted with the sender-to-receiver transfer cipher using the next sequence nonce.

Receiving this frame does not by itself make the transfer successful. The receiver checks all of the following before committing the destination:

  • the completion size equals the size declared in the original metadata;
  • the number of plaintext bytes actually received equals that same size;
  • all buffered bytes can be flushed to the destination;
  • the receiver’s finalized SHA-256 digest equals the digest sent by the sender.

Only after those checks succeed does the receiver complete the destination. For a filesystem destination, the incoming payload has been written to a temporary partial file and completion commits it to its final, non-conflicting file name. Existing files are not silently overwritten.

The receiver then sends the final acknowledgement frame back to the sender:

{
  "size": 12582912,
  "sha256": "<receiver's encoded SHA-256 digest>"
}

It prefixes the JSON with type 0x04 and encrypts it using the receiver-to-sender transfer cipher.

The sender waits for this acknowledgement under a 30-second timeout. It compares both the acknowledged size and digest with its own values. The transfer is reported as complete on the sender only after that comparison succeeds.

This last exchange distinguishes “all bytes were handed to WebRTC” from “the receiver authenticated, persisted, and verified the same plaintext bytes”:

Sender                                           Receiver
  │                                                  │
  │──── completion { size, sha256 } ────────────────▶│
  │                                                  │ flush final bytes
  │                                                  │ finalize SHA-256
  │                                                  │ compare size and digest
  │                                                  │ commit destination
  │◀── acknowledgement { size, sha256 } ────────────│
  │                                                  │
  │ compare acknowledgement                          │ report local completion
  │ report completion                                │

If the receiver has already committed the payload but the final acknowledgement cannot be delivered, the receiver still reports an honest local success: it possesses a verified file. The sender, however, cannot claim that the receiver confirmed completion and eventually times out. The two sides report only what each can prove from its own state.

Failure and cleanup

Any authentication failure, unexpected frame type, sequence disagreement, size mismatch, checkpoint mismatch, digest mismatch, stalled channel, write failure, or timeout terminates the active transfer.

On the receiver, an uncommitted destination is aborted, temporary partial data is removed, the running digest is discarded, buffered plaintext and fragment state are cleared, and the per-transfer ciphers are retired. On the sender, pending checkpoint or completion waits are cancelled and its per-transfer ciphers and digest state are also retired.

The connection-scoped replay cache is separate from that cleanup. Once a receiver has successfully opened the destination and established the transfer ciphers for a transfer ID, that ID remains accepted for the lifetime of the WebRTC connection even if the later data phase fails. Replaying the old bootstrap metadata therefore cannot reopen another destination for the same accepted transfer.

At the end of either a successful or failed transfer, the direction-specific keys and their sequence counters are never reused. A later transfer must begin again with a new metadata exchange, a new transfer ID, a fresh receiver transfer salt, and newly derived transfer ciphers.

Optimization

The transfer protocol was optimized by measuring where time was actually spent rather than assuming that encryption or the network was always the bottleneck. sendon records source reads, block assembly, hashing, AES-GCM work, Data Channel buffer queries, buffer-drain waits, send dispatch, fragment reassembly, destination writes, checkpoint waits, and final acknowledgement time separately.

That distinction matters because the slowest component changes between platforms. A browser-to-browser transfer on a fast local network can become limited by hashing and browser storage, while a transfer involving an older phone can be limited by cryptography, platform-bridge calls, event-loop scheduling, or WebRTC backpressure. A direct Internet transfer is often limited almost entirely by the available network path.

The optimizations below keep the same seven frame types, AES-GCM authentication, direction-specific keys, ordered delivery requirement, checkpoint validation, and final SHA-256 comparison. They change how efficiently the same protocol work reaches the transport and destination.

Larger logical blocks without oversized WebRTC messages

The conservative interoperability unit is a 12 KiB physical Data Channel message. Sending every 12 KiB as an independent logical operation is safe, but it makes a large transfer cross the Dart, browser, Flutter, or native boundary thousands of extra times.

Where both peers support it, sendon instead assembles 256 KiB of plaintext as one logical block. It hashes and encrypts that logical block once, producing one authenticated transfer envelope. The encrypted envelope is then divided back into sub-16 KiB physical fragments for WebRTC delivery.

before
──────
12 KiB plaintext → encrypt → one WebRTC send
12 KiB plaintext → encrypt → one WebRTC send
12 KiB plaintext → encrypt → one WebRTC send
...

optimized
─────────
256 KiB plaintext → hash/encrypt once → bounded physical fragment burst

The receiver reassembles the physical fragments before decryption, so the cryptographic transcript remains a sequence of complete logical frames. Fragmentation does not weaken AES-GCM authentication and does not expose unauthenticated plaintext.

Production peers can carry these fragments over the primary ordered channel in bounded bursts. The protocol also supports a negotiated ordered bulk channel where the platform requires that path. A peer that cannot safely preserve fragment ordering, including affected browser implementations, simply declines the larger block size and stays on the proven 12 KiB path.

Bounded fragment bursts

Even after a 256 KiB logical block has been fragmented, submitting each physical fragment through a separate serial wait adds scheduling and bridge overhead. sendon groups at most eight physical fragments into one dispatch burst.

The burst is deliberately bounded. It reduces repeated asynchronous calls without allowing an entire file, or even an arbitrary number of blocks, to accumulate in the native transport queue. After the current logical block has been submitted, the sender returns to the normal buffer and checkpoint controls before advancing indefinitely.

Event-driven Data Channel backpressure

The sender cannot assume that a successful send call means the bytes have reached the receiver. It only means that the local WebRTC implementation accepted them into its queue. sendon therefore limits how much encrypted data may be buffered locally:

Web sender:             1 MiB maximum buffered window
Native/Android sender:  256 KiB maximum buffered window

The Web window is larger because browsers handled the deeper queue efficiently in measured transfers. The native window remains more conservative because increasing it on the tested older Android phone produced longer drain and checkpoint stalls instead of higher throughput.

Earlier drain handling repeatedly polled the buffered amount with timers. That creates unnecessary calls and behaves poorly when a browser throttles background timers. The production path now installs the Data Channel’s buffered-amount-low callback and sleeps until the transport reports that the queue crossed the threshold.

There is a small race between reading the buffered amount and installing that callback: the queue could drain during that gap, before the listener exists. sendon closes the race by reading the amount once more immediately after installing the callback. If it has already drained, the sender resumes without waiting for a notification that already happened. A genuine failure to drain still ends with the normal 30-second stall timeout.

For 256 KiB logical blocks, production checks the buffered amount once before submitting the block rather than before every small physical fragment. This removes most bridge queries while preserving the bounded window.

Wider cumulative checkpoints

Waiting for a receiver acknowledgement after every 256 KiB made round-trip latency and destination flushing visible in nearly every block. sendon now places cumulative checkpoints every 8 MiB, plus the final checkpoint for a non-empty payload.

This reduces acknowledgement traffic by a factor of 32 compared with a checkpoint per 256 KiB block:

old pacing:  32 checkpoint round trips per 8 MiB
current:      1 checkpoint round trip per 8 MiB

The sender is still bounded by the much smaller local Data Channel window while those 8 MiB are transmitted. The wider checkpoint therefore does not mean 8 MiB must sit in application memory. It only allows multiple drained transport windows to progress before requiring end-to-end confirmation from the receiver.

Coalesced destination writes

The receiver commonly gets authenticated plaintext in 12 KiB or 256 KiB logical blocks, but writing each small block separately is expensive, especially when a browser write crosses into JavaScript storage or an Android write crosses a platform channel.

sendon coalesces compatible plaintext blocks into a bounded 1 MiB destination buffer. It flushes this buffer when it is full, when a checkpoint must be acknowledged, and before final verification and commit.

This reduces filesystem, OPFS, and platform-bridge calls without changing the stream contents or retaining the complete large payload in memory. The 1 MiB buffer is reused as the transfer advances, so its size is independent of the payload size.

Native Android AES-GCM

Initial Android measurements showed that the pure Dart AES-GCM implementation was expensive on the tested phone. sendon therefore moves per-transfer Android AES-256-GCM operations to the platform JCA provider using AES/GCM/NoPadding.

Each transfer direction opens a native session containing only its derived directional key. Dart still owns the protocol decisions: HKDF inputs, direction labels, sequence counters, nonce construction, additional authenticated data, frame types, and failure behavior remain unchanged. The native layer receives one logical block, nonce, and authenticated-data value and returns the same ciphertext || authentication tag layout as the Dart implementation.

The native operations run on a dedicated background executor instead of blocking Flutter’s UI thread. Processing one 256 KiB logical block per platform call also avoids paying the Dart-to-Android bridge cost for every 12 KiB physical fragment.

In the measured Web-to-Android smoke transfer, native AES reduced average decryption time from roughly 231 milliseconds per logical frame to about 10.5 milliseconds, around a 21-fold reduction in that specific crypto bucket. End-to-end throughput improved by a smaller amount because the bottleneck moved into WebRTC queue drainage; optimizing one stage does not make the other stages disappear.

Native Android streaming SHA-256

After AES-GCM became inexpensive, SHA-256 became the largest explicit Android CPU cost. sendon moved the Android transcript digest to the platform MessageDigest SHA-256 implementation on the same dedicated executor.

The hash remains streaming and covers exactly the same plaintext bytes in the same order. Only the implementation performing each update changed. The native digest receives one logical plaintext block per call and returns the same final 32-byte SHA-256 value used by the completion handshake.

In the measured 128 MiB Web-to-Android comparison, hashing time fell from about 28.3 seconds to 4.5 seconds. End-to-end time fell from about 173 seconds to 141 seconds in that run. The sender’s buffer-drain time also fell because the receiver returned to servicing incoming WebRTC work sooner.

Storage-backed large transfers

Keeping a complete received file in memory makes the transfer itself appear simple, but saving can require another full-sized copy across a platform boundary. That was especially harmful on Android: a 512 MiB transfer could authenticate successfully and then fail only when the completed byte array was handed to the save layer.

The receiver backing boundaries described in the Transfer Phase prevent that amplification. Large mobile transfers stream into pending or partial device-storage entries, large Web transfers stream into OPFS, and CLI transfers always stream into a partial file. Only the small bounded working buffers remain in memory.

This optimization improves reliability more than raw network throughput. It moves destination I/O into the transfer lifetime, where checkpoints and backpressure can account for it, instead of creating one large allocation and save operation after all network work has already completed.

Source streaming and archive staging

Large native sources, every browser source, every CLI source, and generated ZIP payloads remain storage-backed. The transfer pipeline reads them incrementally and begins encrypting without creating another complete in-memory copy.

Folders and multi-file selections are packaged before transmission because the wire protocol transfers one named binary payload. Native builds create a temporary ZIP and later stream it; Web builds create the ZIP in browser-managed storage and stream it from there. Separating archive preparation from transfer avoids holding both the uncompressed inputs and complete archive in application memory.