Protocol Specification
This specification describes the protocol as implemented in Ember 1.6.3
(EMBER_DHT_VERSION = 4, EMBER_DHT_MIN_VERSION = 4).
Two breaking changes have landed since wire version 2: v3 reshaped contact
lists and FOUND_VALUE, and v4 bound a frame signature to the
session it arrives on. Both are marked in place. Everything added since
without a bump — query constraints, record media, the callback and
endorsement pair, channel records, and the advertised version range — is
marked additive, meaning a peer that does not
speak it parses the frame correctly and ignores the addition. Remaining
work, including the dormant native transfer path, is called out where it
affects interoperability. Design notes that are not yet wire-stable live in
docs/ember-dht.md.
What EmberDHT is, what it is not, and the constraints that shaped it.
Ember is a modern eMule-compatible P2P client. Beside speaking KAD and eD2K, every Ember node also runs EmberDHT: its own encrypted Kademlia overlay, used to find other Ember nodes, publish shared files, and resolve download sources without a directory server, tracker, or shipped seed list.
EmberDHT answers three questions among Ember-capable peers:
Content still moves over the eMule client-to-client wire. EmberDHT discovers
the source; eD2K transfers the bytes. A 256 KiB chunk protocol with a BLAKE3
hash tree exists in ember/transfer.rs but is not yet wired in.
EmberDHT is a second network, not a replacement. It is always on
(ember_native_enabled). The settings switch remains visible but cannot
be turned off. The overlay rides the same UDP socket as KAD (port 4672 by
default) and is demultiplexed by a two-byte magic, so one forwarded port
serves both stacks.
KAD and eD2K are also how a cold Ember node finds the overlay. There is no Ember bootstrap server. Join paths are the KAD rendezvous key, peers noticed in ordinary KAD traffic, Ember-capable eD2K sessions, gossip, and a persisted contact file. The Friends rendezvous server has no role in DHT bootstrap: a server-hosted pool would hand the operator an identity-to-IP roster of every participant.
BLAKE3(Ed25519 public key)[..16]. Gossip that claims an ID
without the matching key is discarded.DhtMessage only exists in verified form.
Integers on the DHT frame header and most payload lengths are
little-endian. Socket ports in contact-list and PONG
addresses are big-endian; TCP/UDP ports inside a signed
source record are little-endian. Byte strings are concatenated in the order shown.
Hexadecimal constants use a 0x prefix. Time constants are wall-clock
seconds unless noted as Instant (process-monotonic).
src-tauri/src/network/ember/ win. Protocol constants live in
dht/mod.rs and dht/messages.rs; publish/search drivers
and maintenance live in network/mod.rs.
Ember frames share the KAD UDP socket. The network task demultiplexes on
magic, decrypts through the Noise session table, then offers the plaintext
to the control decoder first and the DHT decoder second. The two namespaces
cannot alias: control version 0xC1 sits far outside the DHT
version range.
network/mod.rs — keyword & source cycles, KAD bridge, rendezvous lookup, streaming results0x01) and source (0x02) blobs, Ed25519-signed by the publisher, stored under 16-byte BLAKE3 keys0xEB 0x3E · default port 4672 · max datagram 4096 B, replies packed to 1400 BFigure 1. EmberDHT stack. The engine is deliberately free of sockets so the protocol is unit-testable without a live network.
| Does | Does not |
|---|---|
| Find Ember nodes and keep a verified routing table | Move file bytes (that is still eD2K c2c) |
| Publish and look up keyword and source records | Replace KAD or eD2K; it runs beside them |
| Carry BLAKE3 integrity digests on records | Host a bootstrap roster or seed list |
Let firewalled nodes publish via buddy PROXY_STORE, consume via CALLBACK_REQ, and start the Ember broker for LowID↔LowID | Guarantee a working relay (that still needs admitted ERAT candidates) |
| Refuse incompatible wire versions at the version byte | Negotiate an upgrade or tell the user why a peer was dropped |
Profiles that still had the overlay off are turned on at load. Publishing starts once the node has contacts. Search, the Ember Network page, and the status bar wait for a verified contact before showing Connected — gossip in the table is not a join. A fresh node therefore reports Connecting for a minute or two while the first maintenance cycle runs the KAD bridge, rather than a fault. After the join timeout with still-zero verified peers, Search shows the muted no-peers hint.
Every Ember node holds two static keypairs, generated once and persisted:
| Key | Algorithm | Role |
|---|---|---|
| Ed25519 | ed25519-dalek, strict verification | Node identity, frame signatures, record signatures, node ID derivation |
| X25519 | static Noise key | Transport session (IK when known, XX when not) |
The 128-bit ID is compatible with Ember’s existing ember_hash
field and is cryptographically bound to the signing key. Since wire version 3
a contact list carries no node_id at all: the receiver derives it
from the Ed25519 bytes beside it and refuses the contact if the key is not a
valid curve point. Before v3 the ID was transmitted and then re-derived and
compared, so the only thing those 16 bytes could do was disagree with the key
they travelled with. Deriving rather than checking is the primary defence
against routing-table poisoning: a peer cannot name a contact under an ID it
does not control, because it does not get to name the ID.
encode_message serializes version, type, request id, sender id,
public key, payload length, and payload, then appends a 64-byte
Ed25519 signature over every preceding byte. decode_message
verifies that signature and the sender_id == BLAKE3(pubkey)[..16]
binding before constructing a DhtMessage. An invalid
frame never becomes a message object. This build always includes the public
key on the wire.
Without that binding a signed frame was a bearer token. It proved only that
its sender_id had once signed those bytes, not that whoever
delivered them was that sender — so anyone who had ever received a frame
from Alice could replay it verbatim inside their own Noise session, and the
receiver, which records a verified contact from the session’s address and
static key on every frame that decodes, would file Alice at the replayer’s
address with the replayer’s key. The noise_pub pin then held
that entry against the real Alice, and replaying kept it fresh so it never
aged out.
A v3 signature cannot verify here and ours cannot verify there, so the version byte had to move with it. This is the change that makes a refusal at the version check preferable to a stream of “signature verification failed”.
A stored record is a publisher-signed blob. The signature covers the record body (type, hashes, size, publisher key, timestamp, name, and for sources the contact block). Storers re-send the identical bytes; they do not re-sign. Expiry is computed from the signed creation timestamp, so replication cannot extend a record’s life.
Ember UDP payloads are told apart from KAD/eD2K by magic
0xEB 0x3E. Everything after the magic is encrypted. Patterns:
Noise_IK_25519_ChaChaPoly_BLAKE2s — peer’s static X25519 is already known (routing-table contact, previous session).Noise_XX_25519_ChaChaPoly_BLAKE2s — first contact, including the eD2K-bridge path where no static key is known.| Type | Name | Role |
|---|---|---|
| 0x01 | PKT_IK_INIT | IK handshake, initiator → responder |
| 0x02 | PKT_IK_RESP | IK handshake, responder → initiator |
| 0x03 | PKT_XX_MSG1 | XX handshake message 1 |
| 0x04 | PKT_XX_MSG2 | XX handshake message 2 |
| 0x05 | PKT_XX_MSG3 | XX handshake message 3 |
| 0x06 | PKT_XX_COOKIE | Stateless retry cookie (return-routability) |
| 0x10 | PKT_TRANSPORT | AEAD payload (DHT or control) |
Figure 2. Outer UDP layout. Transport packets add an 8-byte nonce and 16-byte Poly1305 tag around the inner frame (27 bytes of transport overhead including magic and type).
An unauthenticated XX msg1 can be forged off-path. The responder sends a
stateless retry cookie (PKT_XX_COOKIE) once the
unvalidated-handshake budget is spent, so a peer that does not know this
packet type still completes first contact whenever the node is not under a
flood. The cookie is only useful to a source address that can receive a UDP
reply.
Inbound XX handshakes may briefly queue outbound traffic for that peer (honest simultaneous-open). After a 3-second grace the node dials the identity itself; once an initiator handshake of ours is pending, further inbound XX msg1s from that address are refused so a forged handshake cannot be renewed.
Two nodes that dial each other with IK at the same moment are settled by
Noise static key rather than both refusing and waiting out the 30-second
handshake timeout. The node with the lower static key stays initiator and
refuses the inbound IK_INIT. The node with the higher key answers
it as responder, and when that handshake is from the identity it was dialling,
it drops its own initiator and re-sends that initiator's first message and
queued payloads over the resulting session. Each side decides from its own key
and the key it dialled, so both reach the same answer. A pending XX initiator
still refuses an inbound IK_INIT.
| Limit | Value | Why |
|---|---|---|
| SESSION_TIMEOUT | 900 s | Must outlive the 600 s liveness-ping interval; 300 s forced a fresh handshake on every ping |
| MAX_SESSIONS | 4096 | LRU eviction when full |
| MAX_SESSIONS_PER_ADDR | 4 | Claimants coexist; keyed on (address, static key) |
| Pending handshakes | 512 | Caps unauthenticated work |
| MAX_EMBER_DATAGRAM_BYTES | 4096 | Hard drop before decryption; sender sees a “successful” send |
Sessions are keyed on (address, static key), so claimants at one
address coexist instead of ranking for a single live slot. Four matches the
old 1-live-plus-3-shadow budget. A genuine first contact at an address already
full of spoof sessions is kept; a named outgoing identity no longer discards
another key at that address.
Control frames and DHT frames share one decrypted byte stream. The payload
is offered to the control decoder first. Control version is 0xC1
so it can never collide with EMBER_DHT_VERSION (which counts up
from 1). A historical bug: control version 1 made
CONTROL_KIND_EXCHANGE_DATA (4) alias MSG_FOUND_NODE (4),
and every iterative lookup stalled after its first hop.
| Kind | Value | Body |
|---|---|---|
| CONTROL_VERSION | 0xC1 | Leading byte of every control frame |
| PING / PONG | 1 / 2 | Transport keepalive (not DHT PING) |
| EXCHANGE_REQUEST | 3 | Empty; ask for UDP EPX payload |
| EXCHANGE_DATA | 4 | EPX wire payload, packed to the datagram budget |
0xF0 on the eMule extended protocol)
is a separate mechanism: source lists exchanged with peers you are already
transferring with. It works with the overlay switched off. This specification
covers the DHT overlay only.
A receiver drops anything over 4096 bytes before decryption. Replies whose
size we choose are packed to MAX_UNFRAGMENTED_DATAGRAM = 1400,
matching the spirit of KAD’s UDP_KAD_MAXFRAGMENT (1420). Fragmented
UDP is dropped by a fair number of consumer NATs; a shorter answer that
arrives beats a complete one that does not.
27 = transport overhead (magic + type + nonce + tag). 120 = frame overhead
(22-byte header + 32-byte public key + 2-byte length + 64-byte signature).
FOUND_VALUE record blobs then have 1253 − 22 = 1231 bytes after
its 22-byte header (key, next_position, total_available,
count) — about five typical keyword records per reply.
EmberDHT is a 128-bit XOR-metric Kademlia with a replacement cache, verified contacts, and network-size-adaptive diversity caps.
Bucket 127 holds contacts in the opposite half of the keyspace; bucket 0 holds the nearest neighbours. Node IDs are uniform, so occupancy is geometric: about half of all contacts fall in the last bucket and a quarter in the one before. Diversity caps are designed around that, not around a flat table.
A routing-table contact is:
| Field | Size | Notes |
|---|---|---|
| node_id | 16 | Derived from ed25519_pub; not carried on the wire since v3 |
| addr | IPv4 or IPv6 + port | Port 0 is not admitted |
| noise_pub | 32 | X25519 static key for IK |
| ed25519_pub | 32 | Signing / identity key |
| last_seen | i64 unix | > 0 means verified |
| failed_queries | u8 | Consecutive unanswered queries |
Verified means we have heard a signed frame from the contact
directly. Gossip from FOUND_NODE / PEER_LIST /
ANNOUNCE_PEER and entries loaded from disk arrive with
last_seen = 0. They are kept as leads but are not preferred for
seeding lookups, answering peers, persisting, or computing network scale.
Each bucket holds up to k contacts plus a k-entry replacement cache. When
the bucket is full, a new contact is cached and the caller is asked to ping
the oldest live entry. If that ping fails, evict_and_replace
swaps in the newest cache entry. KAD in this codebase has no replacement
cache.
A contact is refused if any of the following hold:
ipfilter.dat.NetworkScale is full (see §12).BLAKE3(key)[..16] disagrees with the advertised ID.A node accepts a STORE only if it is among the k closest contacts it knows of to the key — or if it knows fewer than k contacts, in which case it is among the k closest by definition. Distance is XOR against the local ID compared to the k-th closest contact. Gossip does not count: letting unverified leads into the comparison would let an attacker push the node out of its own neighbourhood.
| Constant | Value | Meaning |
|---|---|---|
| EMBER_DHT_VERSION | 4 | Wire version this build speaks |
| EMBER_DHT_MIN_VERSION | 4 | Oldest version this build can parse |
Version 0 is invalid. A frame outside [MIN, VERSION] is refused
at the version byte rather than becoming a malformed-payload counter that
reads like packet loss.
| Version | Shipped | What changed shape |
|---|---|---|
| 2 | 1.5.3 | Batched store, its ack, contact-list trimming and payload limits all changed while the byte still read 1 — two peers announced the same version and then misparsed each other. |
| 3 | 1.5.7 | node_id dropped from wire contacts (87 → 71 bytes); FIND_VALUE gained start_position and FOUND_VALUE gained next_position and total_available. A v2 peer reads both at fixed offsets, so neither could be additive. |
| 4 | 1.6.0 | The sender’s Noise static key joined the signed bytes without being transmitted (§3.3). No layout change; a v3 signature simply cannot verify here. |
A change that only adds to the format can lower
MIN_VERSION instead of raising both. Neither of the last two
could: v3 moved fields a v2 peer reads at fixed offsets, and v4 changed what
the signature covers.
EMBER_DHT_VERSION partitions the overlay regardless of where
MIN_VERSION sits — the peer that refuses the frame is the one
running the old range, and it cannot be reasoned with after the fact.
Since 1.6.3, PING and PONG carry the range their
sender can decode (§6.5), so a future encoder can ask per peer rather than
assume, and send the older shape to anyone who has not answered. That does
not help against peers predating the advertisement itself, which is why the
diagnostic counting how many contacts have advertised one is the number to
watch before any further bump.
A refused frame is counted and split by direction, so a node can tell peers that are behind from peers that are ahead, and the Ember page raises a banner while older peers are being turned away. The peer that needs to update is by definition the one that cannot decode anything we send.
The encoder always includes the sender’s 32-byte Ed25519 public key
(include_pub_key = true). The decoder always expects it
(has_pub_key = true) and refuses a frame that omits it, so a
peer who has never seen us can verify the signature and learn our identity
even inside an established Noise session. The size budget in §4.5 therefore
always counts those 32 bytes. The encode/decode flags still exist in the
API; production never takes the other branch.
Header with public key is 54 bytes before payload_len: version + type + request id + sender id + public key. Request id 0 is skipped by the allocator so it can never collide with “unset.”
| ID | Name | Direction | Payload |
|---|---|---|---|
| 0x01 | PING | → | optional advertised version range |
| 0x02 | PONG | ← | optional observed SocketAddr, then optional advertised version range |
| 0x03 | FIND_NODE | → | target node ID (16) |
| 0x04 | FOUND_NODE | ← | contact list |
| 0x05 | STORE_RECORD | → | key + record + record signature |
| 0x06 | STORE_ACK | ← | key (16) |
| 0x07 | FIND_VALUE | → | 1…8 keys + start_position, then an optional constraint block |
| 0x08 | FOUND_VALUE | ← | key + next_position + total_available + records (or FOUND_NODE on miss) |
| 0x09 | ANNOUNCE_PEER | → | contact list (gossip dump) |
| 0x0A | PEER_LIST | ← | contact list |
| 0x0B | PROXY_STORE | → | same body as STORE_RECORD |
| 0x0C | PROXY_STORE_ACK | ← | key (16) — buddy accepted, not a DHT STORE_ACK |
| 0x0D | STORE_BATCH | → | up to 64 records for one destination |
| 0x0E | STORE_BATCH_ACK | ← | u64 bitmap of accepted positions |
| 0x0F | CALLBACK_REQ | → | publisher id (16) + file hash (16) + searcher TCP port (2 LE) + crypt options (1) + searcher user hash (16) + callback token (16). Searcher → HighID buddy. The searcher's IP is the UDP source, not in the payload. |
| 0x10 | CALLBACK | → | file hash (16) + searcher IPv4 (4) + TCP port (2 LE) + crypt options (1) + searcher user hash (16) + callback token (16). Buddy → firewalled publisher: connect-and-serve this searcher. |
| 0x11 | CHANNEL_MSG | → | AEAD channel gossip body for overlay rooms. additive |
| 0x12 | CHANNEL_RELAY | → | opaque CHANNEL_MSG envelope for a HighID hop to forward to a LowID target. additive |
| 0x13 | BUDDY_ENDORSE_REQ | → | empty — the requester is the authenticated frame sender_id, and that is what the endorsement binds to. additive |
| 0x14 | BUDDY_ENDORSE | ← | buddy IPv4 (4) + UDP port (2) + noise_pub (32) + expiry (8) + signature (64). additive |
Every type from 0x0F up is additive: a peer that does not speak
one decodes it as Unknown and ignores it, because the header
layout is unchanged. None of them cost a version bump.
BUDDY_ENDORSE_REQ carries no payload and its
decoder ignores the body, so sharing a number with CHANNEL_MSG
would have had a channel frame decode cleanly on that arm and draw a signed
endorsement in reply. Two features developed in parallel must not both claim
the next free byte.
Used by FOUND_NODE, ANNOUNCE_PEER, and
PEER_LIST. Contacts are already ordered closest-first; the
encoder trims the tail to stay inside the unfragmented payload budget.
An IPv4 contact is 7 + 32 + 32 = 71 bytes. At 1253 bytes of
payload, 17 IPv4 contacts fit, so a FOUND_NODE
still never carries a full k-bucket of 20 —
MAX_CONTACTS_PER_RESPONSE is 20 but is not reachable on IPv4.
Version 2 also transmitted a 16-byte node_id per contact, at 87
bytes each and 14 per reply. Every decoder re-derived the ID from the Ed25519
key beside it and compared, so the only thing those bytes could do was
disagree with that key; dropping them bought most of a hop on a sparse table.
This is the wire format only — nodes_ember.dat still
stores an advisory ID per contact (§14.1), because changing the file format
would cost every user their bootstrap set for no bandwidth saving.
The asker is excluded from a FOUND_NODE reply: they were just
added to our table, their distance to themselves is zero, and leaving them
in spent a slot of a size-capped response on the one contact they already have.
PONG may carry the sender’s observed address (the address we
received the ping from), encoded the same way as a contact address:
0x04/0x06 + IP + port (big-endian). An empty PONG
payload decodes as “no observation” for backward compatibility. Observed-IP
voting is specified in §11.
Both frames may then carry a tag/len/value block advertising the wire versions their sender can decode. additive
This is what lets a future encoder pick a frame shape per peer instead of
assuming (§6.1). It is additive because no existing decoder looks here:
PING discards its payload entirely, and PONG reads
its address through a decoder that checks only a minimum length and reports
what it consumed. The version byte therefore must not move for this —
the peers it exists to reach are exactly the ones that would refuse the frame
carrying it.
Three properties are load-bearing. The block only ever trails a field a reader
parses first, so a PONG with no observed address carries none: with
nothing in front of it, an older build would read the tag byte as an address
type and reject the whole frame. Absent, truncated and nonsensical
(min = 0, or min > max) all read as “this peer
said nothing” rather than as a range or an error, because an advisory field
must not be able to drop a frame. And an unknown tag is skipped rather than
fatal, which is what lets a later capability share the same block.
The claim rides inside the signed bytes, so it is exactly as forgeable as the rest of the frame and no more: a relay cannot rewrite what a peer says it speaks. A receiver holds the range per peer and forgets it when the peer leaves the routing table — it is worth what the session that proved it is worth, and one ping relearns it.
The responder returns the closest contacts it has, excluding the asker, packed as a contact list. Iterative use is specified in §8.
MAX_STORE_RECORD_BYTES is the largest body that still fits in a
FOUND_VALUE as the only blob (pack budget minus the 2-byte
length prefix and 64-byte publisher signature). A larger body would store
but never be served.
STORE_ACK and PROXY_STORE_ACK carry the 16-byte key.
PROXY_STORE asks a HighID buddy to fan the same publisher-signed
source record out as ordinary STORE_RECORDs. The buddy’s ack
means “I accepted the proxy request,” not “the DHT stored it.” HighID
sources must not ride PROXY_STORE — that would bypass the
anti-reflection bind (§12.3).
Publishing a large library one datagram per (record, target) puts frame count in proportion to records, saturates the link, and trips the receiver’s per-peer rate limit. Records destined for the same peer travel together, so frame count scales with the number of peers.
Acceptance is per record: the storer re-evaluates proximity and capacity. A count alone forced the publisher to treat “some landed” as “all landed” and retire files that were never stored. The datagram budget caps a real batch at roughly twenty minimum-size records; 64 is a decode-side bound matching the width of the ack bitmap.
The searcher walks the primary (longest) keyword key. Additional keys ride
along so a peer that holds records can intersect on file_hash
locally. On a miss the responder sends FOUND_NODE with closer
contacts, as in classic Kademlia. A declared count the buffer cannot satisfy
is a framing error — the whole frame is refused, so a peer cannot smuggle a
truncated payload. Bytes after the last declared record are ignored, which
leaves room for an additive trailer.
Exhaustion is next_position ≥ total_available, never a zero: a
page that reaches the end of the key reports next_position equal
to total_available, and a searcher stops paging there.
Paging is the searcher’s, not the responder’s. Version 2 had the responder advance a rotation cursor per key, which could not tell “this searcher wants the next page” from “a different searcher wants the first”, so two searchers on one hot key advanced each other’s window and neither saw a contiguous run. Serving is now deterministic: the same request gets the same answer every time.
next_position is reported rather than inferred. The packer skips a
record too large for the budget left on a page and keeps scanning for
one that fits, so a page is not always a contiguous run and
start + len would step over what was skipped — the oversized
record would then never be first, never face an empty budget, and never be
served, while the responder went on reporting it as held. A page therefore
resumes at the earliest record it passed over, which re-sends the ones after
it; content-based dedup at the searcher absorbs that. Positions are advisory,
since the responder’s list shifts as records expire, so paging may repeat or
skip an entry.
Query constraints. additive
A FIND_VALUE may append a tag/len/value block carrying minimum
size, maximum size, file type, file extension, and any keyword hashes past the
eight the count-prefixed run can name (up to
MAX_FIND_VALUE_KEYS_TOTAL = 23 across both runs). The responder
applies it before packing, so total_available counts
matches and a searcher pages through matches rather than through positions it
would discard. A key where nothing matches is answered with contacts, exactly
as a key we do not hold is.
Additive because MSG_FIND_VALUE has always required only a
minimum length and read its fields at fixed offsets, so bytes past
start_position have always been valid and ignored. An
unconstrained query encodes byte-for-byte as it did before. The block is
advisory: a searcher re-applies the same filters at emit, because a peer
predating it will answer without having applied them.
Availability is deliberately not a constraint. KAD can filter on it because its keyword entries carry a publisher-claimed source count; an Ember record carries none, and the number a search page shows counts distinct publishers across the network, which no single responder can see. A constraint no responder could evaluate honestly is worse than none.
Two record types share a header. The DHT key is not a free field: it is
bound to content (keyword_hash or BLAKE3(file_hash)[..16])
and a STORE under the wrong key is rejected.
Then: UTF-8 file name (name_len bytes), then for source records only the contact block.
| Type | Value | DHT key | TTL |
|---|---|---|---|
| RECORD_TYPE_KEYWORD | 0x01 | BLAKE3(lowercase keyword)[..16] | 24 h |
| RECORD_TYPE_SOURCE | 0x02 | BLAKE3(eD2K file_hash)[..16] | 6 h |
| RECORD_TYPE_CHANNEL | 0x03 | Sharded room index / presence / moderation / handoff key | Varies by kind |
Keyword media. additive
A keyword record may carry an optional block after its name:
version[1] then tag/len/value triples for duration,
bitrate, codec, artist, album and title. It is parsed through to the search
row, so an Ember-only hit can fill the Length, Bitrate, Codec, Artist, Album
and Title columns that only a server result used to populate.
Additive because parse_unverified reads the name from its length
prefix and does not length-check a keyword record: an older build parses a
longer body correctly and ignores the tail, and a storer relays the bytes it
was handed. The name budget charges the block, and the publisher signature
covers it, so a relay cannot rewrite it. Publisher-supplied strings are capped
and required to be real UTF-8 — they decide sort keys and column widths, and
must not become a second name field that disagrees with the file name.
The reader shipped one release before the writer, on purpose: by the time anything published a block, the builds that would receive it already understood it. That is the right order for any additive wire change.
Publish, and the Search UI, tokenize with eMule’s extract_keywords
so Ember and KAD hash the same words: split on the eMule separators
()[]{}<>,._-!?:;\/" plus whitespace, keep tokens of
UTF-8 length ≥ 3, de-duplicate case-insensitively in first-seen
order, and drop a trailing 3-byte/3-character token when it is a file extension.
Stemming is not performed.
Each surviving token is hashed as above. The Search UI then packs those hashes
onto FIND_VALUE via compute_keyword_hashes: join the
tokens with spaces, split on whitespace, drop anything shorter than 2 characters
(already gone after extract_keywords), sort by length descending
(most selective first), and de-duplicate. The first hash is the primary walk
key; up to seven more ride along for peer-side file_hash
intersection.
Appended after the file name on source records and covered by the publisher signature:
| Field | Size | Notes |
|---|---|---|
| ip | 4 | IPv4, the address downloaders should dial |
| tcp_port | 2 LE | eD2K client-to-client |
| udp_port | 2 LE | Ember / KAD UDP |
| flags | 1 | bit 0 firewalled, bit 1 obfuscation, bit 2 relay-capable, bit 7 QUIC port present |
| noise_pub | 32 | Stashed for future native (Noise) dialing |
| user_hash | 16 | Optional trailer: publisher's eD2K user hash, present with the buddy block |
| buddy_ip | 4 | Optional trailer: HighID buddy the searcher should CALLBACK_REQ |
| buddy_udp | 2 LE | Optional trailer |
| buddy_noise_pub | 32 | Optional trailer: buddy's Noise static key |
| callback_token | 16 | Optional trailer: publisher-derived token copied into CALLBACK_REQ / CALLBACK |
| buddy endorsement | 104 | Optional trailer: buddy Ed25519 key (32), expiry (8 LE), signature (64) |
| quic_port | 2 LE | Only when flags bit 7 is set, and always the record's last two bytes: the publisher's advertised QUIC port, which a relay dials |
Firewalled records set bit 7 and end with the publisher's advertised QUIC
port. The relay dials that port, not tcp_port. From 1.7.1 QUIC
shares the UDP socket, so the value is the public UDP port. A 1.7.0 publisher's
QUIC endpoint sits on the TCP port number unless it could not bind it, in which
case it is on a neighbouring port, and a NAT may remap either.
Older parsers read these records unchanged. They infer the callback trailer
from the residual length being at least its size and ignore any bytes after
it, and two bytes alone fall short of that size. The decoder clears bit 7
before exposing flags, so it never reaches the EPX source flags
that share the byte.
additive
A downloader uses ip + tcp_port over the existing eD2K path
when the source is not firewalled. Firewalled records append a 70-byte
callback trailer (user hash + buddy address + buddy Noise key + token) after the
41-byte contact. HighID records stay 41 bytes. A parser that ignores extra
bytes still accepts the contact block; this build reads the trailer when it is
present. That is an additive record-body change, not a DHT-frame version bump.
additive
Naming a buddy requires that buddy’s signed endorsement
(BUDDY_ENDORSE, 0x14). It was once also possible to name a buddy
by quoting its Noise static key, which is on every signed frame that buddy
sends — and that was never consent: anyone who had ever heard from a node
could name it and spend a replica, a twenty-node fan-out and a callback slot
on its behalf. That compatibility path is gone from both sides. A firewalled
publisher with no endorsement yet publishes nothing for that file and retries
on the next tick, rather than placing a record no searcher would dial.
ember_file_hash is a 32-byte streaming BLAKE3 of the
file contents — the same digest blake3::hash would
produce for the whole file, computed in the same pass as the eD2K and AICH
hashes (Blake3FileHasher). Search hits, DHT keyword and source
records, and known.met / library entries can carry it. A download
verifies against it whenever one is available: a match shows a green Ember
badge (ember_verified); a mismatch is a permanent download failure
with a red Ember badge and does not reopen parts that already matched the
ed2k hash. Deep links without a digest still complete and are hashed for
future sharing.
That digest is not the dormant native-transfer hash tree. The unused
256 KiB chunk tree in ember/transfer.rs hashes each chunk,
then BLAKE3 of the concatenated chunk hashes; that root is a different value
and is not what DHT records carry.
A source record names an address to download from, so it stops being true the moment that peer goes offline. A keyword record only says a file exists under a word, which stays true whoever is online. Sharing one 24-hour TTL meant a departed peer kept being handed to downloaders for the rest of the day. Publishers re-announce sources every 2 hours, so 6 hours survives two missed republishes while clearing a departed peer four times sooner than a 24-hour source TTL would.
A record's signed timestamp may sit at most min(1 hour, TTL / 2) in the future: an hour for keyword, source and room-index records, 22.5 minutes for a 45-minute presence record. A flat hour was longer than a presence record's whole life. A record dated ahead of the storer's clock is aged from the storer's clock, not from its timestamp, so the tolerated skew cannot extend its life past the TTL. A record older than its TTL is expired.
A presence leave tombstone is the exception: it is aged from its own timestamp. It has to outlive every live copy dated before it, and a storer admits those until their timestamp plus the TTL, so a tombstone from a fast clock aged from the storer's clock lapsed first, and a harvested live copy could be stored again and served as present.
From 1.7.1 a storer refuses a presence record from a publisher whose clock runs 22.5 to 60 minutes fast; 1.7.0 accepted it. Such a node's room presence does not reach 1.7.1 storers until its clock is corrected.
A replayed older copy cannot replace a newer one: created_at is
kept and compared. The store does not acknowledge the refused copy
(STORE_ACK is withheld and its STORE_BATCH_ACK bit is
clear), so its sender does not count it as placed. One exception remains: a
byte-identical copy whose signature the storer saw within
STORE_SIG_REPLAY_TTL (§12.4) is still answered as a replay after
a newer copy superseded it.
| Cap | Value | Behaviour when hit |
|---|---|---|
| MAX_RECORDS_PER_KEY | 1000 | Refuse the newcomer (do not evict incumbents) |
| MAX_RECORDS_PER_PUBLISHER_PER_KEY | 150 | Refuse; ~15% of the key, and now KAD’s absolute numbers rather than only its ratio |
| MAX_KEYS | 50,000 | Evict the furthest key by XOR-distance when the incoming key is closer; refuse a newcomer that is no closer than the furthest held key |
| MAX_STORE_BYTES | 48 MiB | Evict least valuable (furthest from our ID, then nearest expiry) |
| Sources per IP (adaptive) | 8 / 5 / 3 | See NetworkScale; firewalled sources attributed to the forwarder |
The publisher cap is network-wide in effect: every storer applies the same cap
to the same publisher key, and the same allowance applies to their own local
store. It was 45 of 300 through wire version 2 — so a user sharing 200 files
with a word in common got 45 of them findable under that word anywhere,
30% of what KAD serves. Raising both to KAD’s own 150 of 1000 waited on
searcher-driven paging (§6.9): while a peer could only ever serve its first
window, the extra stored records had no way to reach a searcher and the one
certain effect would have been more replication traffic.
MAX_STORE_BYTES is unchanged, so per-key capacity tripled without
raising what the process may hold resident.
MAX_KEYS is not a hard admission freeze. When the map is full,
expired keys are dropped first; otherwise the furthest key from this node
(XOR-distance) is given up, and only when the incoming key is closer than
it. A flood of distant keys therefore cannot deny keys this node is
responsible for. Without a local ID there is no notion of responsibility,
and a full map refuses newcomers.
On the wire, each record in FOUND_VALUE is
record_body ‖ signature[64]. The searcher verifies the publisher
signature, checks that keyword_hash matches the queried key, and
drops junk. The per-node result budget is charged on blobs offered,
not only those accepted, so a peer that sends nothing but junk still hits
its own limit.
A responder need not have applied the storer's clock rule (§7.4), so the searcher drops a blob outside its life before it takes a result slot. Its own clock gets the skew tolerance again in each direction: a blob may be dated up to twice min(1 hour, TTL / 2) ahead, and be up to its TTL plus that tolerance old. Without the second allowance the searcher's clock error stacked on the publisher's, and a searcher half an hour slow refused most presence records.
Both FIND_NODE and FIND_VALUE use an iterative
shortlist, α = 5 outstanding queries, up to k = 20 closest contacts as the
frontier.
| Limit | Value |
|---|---|
| MAX_ACTIVE_SEARCHES | 64 |
| SEARCH_TIMEOUT_SECS | 60 |
| MAX_SEARCH_RESULTS | 300 |
| MAX_LOCAL_SEED_RESULTS | 150 (half the budget) |
| MAX_RESULTS_PER_NODE | 75 (quarter of the budget) |
| MAX_QUERY_ATTEMPTS | 2 |
| Queued-query timeout (cold handshake) | 12 s |
The 12-second queued-query timeout covers a Noise_XX 2-RTT setup plus the query round trip. Charging a first-contact hop the ordinary budget failed the contact before it could answer — and on a cold table every contact is a first contact.
Intersection is sparse: missing secondary keys are skipped,
plus a filename match at emit time. It is not a strict worldwide AND of
every keyword. Product search already tokenizes with
extract_keywords (§7.1), so 1- and 2-byte tokens never become
DHT keys. Successive queries rotate which window a storer serves (§8.4).
The Search page exposes an Ember Only method. Global queries Ember alongside KAD and servers, merging into one de-duplicated list. With KAD and eD2K both offline, Global falls back to Ember alone. Downloads also resolve sources on EmberDHT in addition to KAD and servers.
Ember used to buffer everything and emit on completion, so on a cold table a user waited most of the 60-second cap while KAD hits were already on screen. As of 1.5.3 the keyword search carries a cursor into an append-only result list. A 1-second sweep emits everything past it on the same cadence as KAD: the first record immediately, then every 20. Locally seeded records therefore reach the UI on the first tick.
Batches run through hash de-duplication against what KAD already streamed,
so a hash KAD found arrives as an availability update rather than a
duplicate row. Only the batch flagged final clears ember_pending,
and the completion batch is still queued when empty so that happens exactly
once. The timeout backstop that reaps an expired search runs before
the streaming and emit steps in the same sweep, which is what stops a reaped
search from being streamed after its search-complete.
A peer answers a keyword query with whatever fits the 1231-byte record budget: about five records for a bare keyword record, four for one carrying media, two in the worst case. Through wire version 2 packing filled from the front of insertion order, so the oldest five were the only ones a node would ever serve; a rotation cursor was then added, but the responder owned it and could not tell one searcher from another (§6.9).
Since v3 the searcher owns the offset, so a walk pages a well-stocked node
until the key is exhausted. Paging is the one mechanism here where a
responder influences how many queries we send, so the searcher bounds
it independently of what total_available claims: each follow-up
must name an offset strictly past the one it answered, and both the per-node
result allowance and the page ceiling are two-tier.
While the shortlist still holds an unqueried hop — or any query is outstanding
— one peer may offer MAX_RESULTS_PER_NODE (75, a quarter of the
300-file budget) over up to MAX_PAGES_PER_NODE (25) follow-ups.
Once neither is true there is no hop left for extra records to crowd out, so a
lone storer may spend what remains of the whole budget, over up to
MAX_PAGES_PER_NODE_EXHAUSTED (100) pages. That second page tier
has to be earned: past the base ceiling a node keeps paging only
while it sustains MIN_RECORDS_PER_PAGE_TO_CONTINUE (2) records
per page on average, so a peer answering one record at a time while claiming a
huge total stops at the base ceiling.
Blob de-duplication on the searcher is content-based, so a re-sent record costs bandwidth but not a slot of the per-node allowance, which is charged per distinct blob. Local seeding does no packing at all.
Truncation is still counted:
ember_dht_found_value_truncated and
ember_dht_found_value_withheld, which is how to read how far the
datagram ceiling binds on real keys. Read withheld as records past
this page’s window that it has not served. It once counted from the rewound
resume point, which included records the same page had just put on the wire;
a page that rewinds but still reaches the end of its key now reports zero
withheld and does not increment truncated, because it truncated
nothing.
Shared files are published as keyword records (one per
extract_keywords token) and one source record, and republished
on a cycle to stay alive. Publishing starts once the node has contacts. The
Library marks a file with an
Ember badge once a source record is placed — that is the
point at which other Ember users can actually fetch it.
| Cycle | Interval | Notes |
|---|---|---|
| Source republish | 2 h | Persisted in known.met as FT_EMBER_SOURCE_PUBLISH (0xE4) |
| Keyword republish | 12 h | Persisted in known.met as FT_EMBER_KEYWORD_PUBLISH (0xE5) |
| Rendezvous advertise | 5 h | Only while reachable and already publishing something |
| Publish timeout | 30 s | Wait for STORE_ACK |
| MIN_STORE_NODES | 5 | Minimum replicas the publisher aims for |
| MAX_ACTIVE_PUBLISHES | 128 | In-flight publish operations |
Keyword publish rate per maintenance tick scales with library size, clamped between 2 and 96 files per tick (not records), aiming to finish a full pass inside the 12-hour cycle at an estimate of 8 keywords per file. Source publish is similarly paced (5–256 files per tick) so a large library is not slammed in one tick after a restart.
A keyword round that placed some of a file’s keys but lost one on every replica counts as published and comes back after 30 minutes, up to three times. After that the file waits out the 12-hour cycle, and gets no further short retries until a round places every key: a key storers refuse for as long as the file is shared, such as a word past the per-publisher cap (§7.5), would otherwise bring the file’s whole keyword set back three times every cycle.
The last successful source-publish and keyword-publish are written to
known.met (0xE4 / 0xE5). On start, a
stamp still inside its interval (2 hours for sources, 12 hours for keywords)
is hydrated back to an Instant; a stamp older than the interval,
or one that cannot be represented because the process has not been up that
long, is omitted and the file is due immediately — republish-too-eager, the
safe direction. Previously both schedules were keyed on Instant
alone, so every restart marked the whole library as never-published.
Nodes that hold a record re-send the identical signed bytes to the current
closest nodes every 2 hours (EMBER_RECORD_REPUBLISH_SECS = 7200),
up to 200 records per 60-second maintenance tick, using STORE_BATCH
where possible.
Replication cannot extend a record’s lifetime. Expiry is derived from the publisher’s signed creation timestamp; every recipient computes the same absolute death time. What it buys is churn coverage: copies reaching nodes that joined since the publisher’s last round. Each record already has up to 20 replicas and lives at most 24 hours (6 for sources). Hourly replication was Ember’s single largest traffic item — roughly 48,000 frames an hour at 200 records × 20 replicas — about double the entire publish load. Two-hourly halves that.
Firewalled source records are not republished by storers: their declared
address is unverified, and repeating it would amplify a lie. A
republish_due flag, rather than an old Instant,
marks records whose previous attempt never made it onto the wire —
Instant::now() − 24h is not representable on a machine that has
been up for less than a day, and the saturating fallback stamped the record
as just republished.
Each maintenance cycle logs an Ember publish cycle heartbeat
and an Ember replication cycle heartbeat (due, selected, queued,
re-armed, leftover backlog).
There is no central bootstrap and no shipped address list. A cold node gets in through the following, in rough order of who arrives first.
ember-dht-rendezvous-v1, hashed with MD4 into a KAD ID. The derivation is versioned so the key space can be sharded later without a flag day. Health is surfaced on the Ember page as ember_dht_rendezvous_* (listed / lookups / empty).nodes_ember.dat. FOUND_NODE / PEER_LIST / ANNOUNCE_PEER, plus up to 200 persisted contacts, once the node has been online before. Persisted contacts load as unverified leads.nodes_ember.dat presupposes either a live KAD
connection or an eD2K transfer with an Ember-capable peer. A first-run user
with KAD off and no servers has no way in. Ember rides eMule’s bootstrap by
design. Seed lists are deliberately not planned. For a local two-node test
where neither side can reach KAD, the harness command
add_ember_dht_contact is the only way to introduce two nodes
directly.
seeds.txt or DNS SRV seed lists.The Friends rendezvous server is still used for friend NAT traversal and relay. That is a different protocol and a different trust boundary.
A PONG may echo the address the ping was received from. Votes
are counted per reporter /24, expire after 15 minutes, and
require 3 distinct public nets before the address is
confirmed. Private/loopback reporters and private reported addresses are
ignored so LAN Sybils cannot vote. At most 64 distinct reported addresses
are tracked; the least-recently-updated is dropped.
Without expiry the three-vote threshold was cumulative over the whole process lifetime, so a stale address stayed qualified indefinitely and a genuine address change could not displace it.
A rival address displaces a confirmed one only with strictly more distinct nets. When the confirmed address lapses because its votes aged out, then for one vote lifetime a rival must beat the most nets that ever backed it at once, or else have a quorum of its own nets among those that voted for the lapsed address. The first rule stops three quiet /24s from taking the address the moment honest votes age out. The second admits a genuine address change: the same peers now seeing us somewhere else. On a network with no more nets than the old peak, the change could not otherwise confirm until the lapse aged out. The backers remembered for that test are the 256 nets that voted most recently, so on a long uptime they are still the peers talking to us when the address moves.
A confirmed address becomes the node’s external address only when it has none, and only if STUN has no reading or agrees. Votes never replace an address already held: a confirmation that differs from it triggers a STUN re-probe, no sooner than a minute after the last one, even on an idle node that would otherwise not probe, and the STUN result decides. A live HighID is left as it is.
A LowID / firewalled node sets SOURCE_FLAG_FIREWALLED on its
source record, names one HighID PROXY_STORE target in the
callback trailer, and asks that contact (only) to fan the record out.
Overlay STORE_BATCH of that firewalled record waits for the
buddy's matching PROXY_STORE_ACK, so FIND_VALUE
cannot advertise a buddy that has not yet accepted the record. Storers
attribute the record to the forwarder for the per-IP source cap, because the
declared address is unverified and an attacker could otherwise invent a new
address per record — each getting a fresh quota.
Consume is the KAD callback analogue. A reachable searcher that finds a
firewalled record with a buddy trailer sends CALLBACK_REQ to
that buddy over the Ember Noise session, carrying the token from the signed
trailer. The buddy forwards
CALLBACK only if it recently accepted PROXY_STORE
from the named publisher, copying the token through. The publisher accepts CALLBACK only
from a buddy that answered PROXY_STORE_ACK with the request id
it was sent for that file, and only when the token matches the one
it published, then connects eD2K TCP back to the
searcher's observed address (the UDP source of
CALLBACK_REQ, copied by the buddy — never a claimed IP in the
request). Firewalled Ember DHT contacts are not registered in SourceManager
(pending promotion would otherwise TCP-dial the claimed NAT IP).
WaitCallbackKad rows are kept out of TCP reask and pause/resume
seeding; only CALLBACK_REQ retries them. Consume uses the same TCP-firewalled predicate as publish
(LowID or KAD/server Firewalled), not the UPnP-pessimistic
startup flag. A named buddy that is banned, filtered, or special-use parks
the source — including Searching-only pending downloads — rather than dropping it.
When the searcher is also TCP-firewalled, ingest does not send
CALLBACK_REQ; firewalled Ember records set
SOURCE_FLAG_RELAY_CAPABLE, and that path starts the same Ember
punch/relay broker KAD uses for Ember-capable LowID sources rather than
leaving both sides parked. The broker still needs admitted ERAT candidates.
LowID publishing through buddy PROXY_STORE is implemented but
not yet exercised end to end on a live network. Consume is implemented on
the same path: the buddy table is populated when a PROXY_STORE_ACK
matches a request we sent, so a network that is already forwarding publishes can bounce
callbacks without a second handshake.
When neither side can accept a connection, the broker asks a relay admitted
from an ERAT attestation to bridge a QUIC stream. The request is
RELAY_REQUEST (relay message type 0x01), framed as
type (1) | session_id (4 LE) | payload_len (2 LE) | payload.
Version 2 has a 183-byte payload. Version 3 appends the target's Ember node id
(16 bytes) before the signature, for a 199-byte payload. The relay then
refuses to bridge to any endpoint whose QUIC certificate does not prove that
node id.
| Bytes | Field | Notes |
|---|---|---|
| 0 | version | 2 or 3 |
| 1..5 | target_ip | IPv4 the relay dials |
| 5..7 | target_port | LE; the port the relay dials over QUIC (for Ember DHT sources, the record's quic_port) |
| 7..23 | file_hash | eD2K file hash |
| 23..55 | attestation_hash | The relay's ERAT attestation hash |
| 55..87 | requester_pubkey | Requester's Ed25519 key |
| 87..103 | requester_ember_hash | Must bind to requester_pubkey |
| 103..119 | nonce | Fresh random; replays within 10 minutes are refused |
| 119..135 | target_node_id | Version 3 only |
| 135..199 | signature | Version 3; at 119..183 in version 2 |
The signature is Ed25519 by requester_pubkey over
domain | session_id (4 LE) | attestation_hash | target_ip | target_port (2 LE) | file_hash | requester_pubkey | requester_ember_hash | nonce,
followed by target_node_id in version 3. The domain is
"ember-relay-request-v3\0" for version 3 and
"ember-relay-request-v2\0" for version 2, so a version 2
signature cannot be replayed with a node id attached. The signed order is not
the payload order: the attestation hash comes before the target.
A relay advertises version 3 support with ERAT capability bit
0x02 (RELAY_ATTESTATION_CAP_PINNED_TARGET) beside
the required relay bit 0x01. Verifiers require only
0x01, so older peers accept attestations carrying the new bit. A
requester sends version 3 only to a relay with bit 0x02, because a
1.7.0 relay accepts no payload length but version 2's.
additive
The relay answers on the same stream, echoing session_id, and
only once it has reached the target: it opens a QUIC stream to the target,
sends RELAY_CONNECT (type 0x03, payload the 16-byte
file hash), and then answers RELAY_ACCEPT (type 0x02,
empty payload), after which the stream carries eD2K bytes in both directions.
Any refusal is RELAY_REJECT (type 0x06) with a
one-byte reason, so a requester never takes an unreachable target for a
working relay. Type 0x05 was RELAY_CLOSE, which a
1.7.0 relay sends after ACCEPT when the target cannot be reached; it is
retired, not free.
| Reason | Meaning | Requester |
|---|---|---|
| 0x01 | No session free, overall or for this requester, or relaying is turned off | Skips the relay for 60 s |
| 0x02 | Target refused: port 0, non-public, the requester's own address, or filtered or banned | Attempt fails |
| 0x03 | Payload length of no accepted version, a requester that is not a friend or not the connection's identity, or an unknown or expired attestation hash | Skips the relay for an hour; a friend's relay only while that attestation is still the one held |
| 0x04 | Unsupported version or length mismatch, a key that does not bind to the Ember hash, a bad signature, or a replayed nonce | Attempt fails |
| 0x05 | Target unreachable, without an Ember identity, not the pinned node id, or the requester itself | Attempt fails |
No REJECT counts against the relay: each is a working relay answering. A
relay with no room for the connection itself turns it away before any request
is read, by refusing the handshake (CONNECTION_REFUSED) or by
closing it with reason no session capacity,
active principal session cap reached or
friend reserve requires pinned identity. The requester skips such
a relay for 60 s. A close is not held against the relay, since the handshake
proved who sent it; a refused handshake is, since anyone at the address could
send one. A requester has at most two requests in flight to one relay, the
sessions a relay holds per requester, and keeps the attestation with the
latest expiry when an older copy arrives again.
A bridge passes each direction's end of stream on as soon as that direction ends, closes after 120 s in which neither direction moves a byte, and when it ends waits up to 10 s for each side to acknowledge what is still on its way before the connections close.
Every anti-abuse limit has the same tension: strict values keep one host from occupying the table or the store, but on a small network a legitimate peer looks a lot like an attacker. Limits are derived from verified contacts — gossip is cheap and proves nothing. Counting leads would let anyone who can talk at us drive the limits to their strict tier while we had reached almost nobody.
| Tier | Verified contacts | Per IP | /24 global | /24 per bucket | Sources/IP | Stores/min |
|---|---|---|---|---|---|---|
| Bootstrap | < 10 | 8 | 20 | 5 | 8 | 300 |
| Small | 10…79 | 4 | 16 | 4 | 5 | 200 |
| Established | ≥ 80 | 2 | 10 | 3 | 3 | 120 |
Established per-IP is 2 rather than KAD’s 1 because two genuine instances behind one NAT is ordinary, and unlike KAD, Ember can tell them apart cryptographically. Established per-subnet matches KAD (10 global, and the per-bucket cap converges on 3). Store rate is counted in records, not frames: every record costs two signature verifications whether it arrives alone or batched. Charging per frame let one dense batch buy many times the admitted work of a well-behaved publisher.
Tightening never evicts contacts admitted under a looser tier, so whatever a cold start allows, an adversary present at that moment keeps for the life of the process — in the bucket that matters most, since geometric occupancy puts half of everyone there. The Bootstrap per-bucket cap is therefore kept ≤ k/4.
Applied after the Noise session is up. A two-bucket sliding window approximates a trailing window so a sender cannot straddle a reset and spend twice the limit. Effective sustained rate is about twice the named constant because the window is enforced over a trailing half.
| Budget | Window | Named cap | Keyed on |
|---|---|---|---|
| All DHT frames | 1 s | 40 | IP |
| FIND_NODE / FIND_VALUE / CALLBACK_REQ | 10 s | 60 | IP (shared behind NAT) |
| STORE records | 60 s | NetworkScale | Verified node ID when known |
Lookup budgets stay on the address because resolving identity for every datagram would put a routing-table scan in front of the gate that exists to make junk cheap to reject. STORE budgets use the verified node ID so peers sharing a NAT do not throttle each other.
HighID / direct source records must claim the observed Noise sender IP. Otherwise a peer could point downloaders at a third-party victim (and Ember would be a traffic reflector). Firewalled sources are exempt: they may be stored by a HighID buddy, or from a NAT mapping that differs from the STUN hint. Authorship is still bound by the publisher Ed25519 signature.
Return-routability is required before a node spends anything substantial on a request (the XX cookie is the handshake-level form of the same idea).
Identical STORE frames (same publisher signature) collapse for 60 seconds
so a retransmit storm cannot re-verify the same blob forever. A replay only
counts if the store still holds that record byte for byte, or a newer copy
of it from the same publisher that superseded it (§7.4): the cache is never
cleared by eviction, and treating an evicted record as a replay made
STORE_BATCH set the accepted bit for a replica that was gone,
after which the publisher retired the file. A superseded copy is still
answered as a replay, since the key holds that publisher's newer record for
the file; it is not stored again and does not re-arm the cache.
Cache size 50,000; swept at most once per TTL. An unconditional size check meant a single 64-record batch could walk a 25,000-entry map 64 times.
A 60-second tick drives bucket refresh, liveness pings, gossip, publisher cycles, and storer replication. Each task is internally gated on a longer interval, so the cadence itself is cheap.
| Task | Interval / budget | Notes |
|---|---|---|
| Tick | 60 s | EMBER_MAINT_INTERVAL |
| Disconnect / re-bootstrap | inbound silence > 20 min | Re-arms the KAD bridge, rendezvous lookup, and self-lookup rather than sitting on a dead table |
| Stale-contact purge | verified unheard > 2 h | Contacts still needed by an in-flight search are spared. Unverified leads are not purged this way (last_seen is 0) |
| Self-lookup | 180 s cold / 20 s if ≥ 10 verified; repeats every 4 h | FIND_NODE for our own ID, to fill close-to-home buckets |
| Bucket refresh | idle > 3600 s, max 3 FIND_NODE / tick | Random target in the idle bucket |
| Liveness ping | unheard > 600 s | 8/tick steady; 32/tick while verified < 20. Scales to cover the table inside the stale window, capped at 32. 8 s ping timeout. |
| Lead ping reserve | 1 in 4 of the budget | When verified contacts could spend the whole budget, keep promoting gossip |
| ANNOUNCE_PEER | 2/tick; 16/tick while starved | Highest-yield cold frame; self-limits above 20 verified |
| Storer republish | 2 h, 200 records/tick | STORE_BATCH to current closest |
| KAD bridge | while verified < 20 | Same ping budget as starved liveness |
| Rendezvous lookup | every 10 min while verified < 20 | Stops at one k-bucket |
A ping is the only way an unverified lead becomes usable, so the starved budget is also the join rate. Eight a minute is a sensible steady-state trickle and a poor join. Verification is cheap — one small frame each.
last_seen is discarded and set to 0. Restoring it would make every entry look verified the moment it loads, so the staleness purge would delete the whole bootstrap set before a single ping went out. Zero also sorts them first for liveness probing.data and signature are load-bearing on the way back in. Key, publisher, and expiry are re-derived from the signed body. Believing the unsigned fields would let anything that can write the file place a genuine record under the wrong key or with a life it was never granted.| Location | What |
|---|---|
| known.met tag 0xE4 | Last successful Ember source-publish per file |
| known.met tag 0xE5 | Last successful Ember keyword-publish per file |
| ember_dht_highwater.json | Daily and all-time high-water of verified contacts |
| ember_source_address.json | External address our source records carry, and since when; source-publish stamps older than that do not hold a file back from republishing |
EmberDHT was compared against this repository’s own KAD stack, constant for constant. On design it already matches or exceeds KAD operationally in the rows below. Remaining gaps are in §16.
CALLBACK_REQ for firewalled consumeFIND_VALUE paging
Roughly half the round trips per lookup come from α = 5 against KAD’s 1 for
FindNode. Replica count is double. The diagnostic surface is
already richer. Firewalled consume (§11.2) now has an Ember callback path
matching KAD's buddy relay, and Ember DHT ingest starts the same LowID↔LowID
broker KAD uses for Ember-capable sources. The largest
product gap is that bytes still move over eD2K.
Both items this section previously listed as waiting on a bump have landed, as wire version 3, together — one bump paid for both, and the capacity raise was only worth having once paging could serve the extra records.
start_position on FIND_VALUE / FOUND_VALUE — searcher-owned paging, replacing the responder-side rotation cursor (§6.9, §8.4).node_id dropped from wire contacts — 87 → 71 bytes, 14 → 17 contacts per FOUND_NODE (§6.4).MAX_RECORDS_PER_PUBLISHER_PER_KEY raised to KAD’s 150 of 1000, unblocked by the above (§7.5).Everything else added since needed no bump, because each went where an existing decoder does not look — after the fields a payload’s parser reads at fixed offsets, or after a record’s length-prefixed name: query constraints and keyword media, the callback and endorsement message pair, channel records, and the advertised version range. A change that must alter an existing field still has no path but a bump.
ember/transfer.rs).Gossip is now scored rather than merely rate-limited: an introducer’s probes are rationed by how many of the leads it gave us went on to answer, though it is never consulted while the table is starved and never refuses a contact outright. Sybil pressure on the store remains bounded by address diversity rather than by identity, since every per-publisher share is keyed on a free keypair; the per-address STORE ceiling is where a rotating-identity flood becomes visible, and it is counted.
ember/transfer.rs specifies a stream of
REQUEST_CHUNKS, CHUNK_DATA, FILE_INFO,
and TRANSFER_COMPLETE over 256 KiB chunks with a BLAKE3 hash
tree whose root is a separate digest from the streaming BLAKE3 already
carried on DHT records. Nothing in the tree imports this module. Wiring it
is the largest remaining piece if the goal is a network that does not need
the eMule wire at all.
Values as compiled in Ember 1.6.3. Source files in parentheses.
| Constant | Value | File |
|---|---|---|
| EMBER_DHT_VERSION / MIN | 4 / 4 | dht/mod.rs |
| K_BUCKET_SIZE | 20 | dht/mod.rs |
| ALPHA | 5 | dht/mod.rs |
| ID_BITS | 128 | dht/mod.rs |
| MAX_CONTACTS_PER_RESPONSE | 20 (17 IPv4 fit) | dht/mod.rs |
| CONTACT_TIMEOUT_SECS | 600 | dht/mod.rs |
| MAX_FAILED_QUERIES | 3 | dht/mod.rs |
| EMBER_MAGIC | 0xEB 0x3E | transport.rs |
| CONTROL_VERSION | 0xC1 | transport.rs |
| Constant | Value | File |
|---|---|---|
| MAX_EMBER_DATAGRAM_BYTES | 4096 | transport.rs |
| MAX_UNFRAGMENTED_DATAGRAM | 1400 | dht/messages.rs |
| FRAME_OVERHEAD | 120 | dht/messages.rs |
| TRANSPORT_OVERHEAD | 27 | dht/messages.rs |
| MAX_UNFRAGMENTED_PAYLOAD | 1253 | dht/messages.rs |
| MAX_FOUND_VALUE_RECORD_BYTES | 1231 | dht/messages.rs |
| MAX_STORE_RECORD_BYTES | 1165 | dht/messages.rs — FOUND_VALUE singleton pack budget (1231 − 2 − 64) |
| MAX_FOUND_VALUE_BLOB_BYTES | 1229 | dht/messages.rs — body + 64-byte signature |
| MAX_FIND_VALUE_KEYS | 8 | dht/messages.rs |
| MAX_FIND_VALUE_EXTRA_KEYS | 15 (constraint block) | dht/messages.rs |
| MAX_FIND_VALUE_KEYS_TOTAL | 23 | dht/messages.rs |
| CONTACT_WIRE_LEN_IPV4 | 71 | dht/messages.rs |
| MAX_STORE_BATCH_RECORDS | 64 | dht/messages.rs |
| MAX_FOUND_VALUE_RECORDS | 300 | dht/messages.rs |
| Constant | Value | File |
|---|---|---|
| EMBER_MAINT_INTERVAL | 60 s | network/mod.rs |
| EMBER_BUCKET_REFRESH_SECS | 3600 | network/mod.rs |
| EMBER_RECORD_REPUBLISH_SECS | 7200 | network/mod.rs |
| EMBER_SOURCE_REPUBLISH | 2 h | network/mod.rs |
| EMBER_KEYWORD_REPUBLISH | 12 h | network/mod.rs |
| EMBER_KEYWORD_PUBLISH_MIN/MAX_PER_TICK | 2 / 96 files | network/mod.rs |
| EMBER_SOURCE_PUBLISH_MIN/MAX_PER_TICK | 5 / 256 files | network/mod.rs |
| KEYWORD_RECORD_TTL | 24 h | dht/store.rs |
| SOURCE_RECORD_TTL | 6 h | dht/store.rs |
| CLOCK_SKEW_TOLERANCE_SECS | 1 h, capped at TTL / 2 per record | dht/store.rs |
| EMBER_RENDEZVOUS_REPUBLISH_SECS | 5 h | network/mod.rs |
| EMBER_RENDEZVOUS_LOOKUP_INTERVAL_SECS | 10 min | network/mod.rs |
| EMBER_SEARCH_QUEUED_QUERY_TIMEOUT | 12 s | network/mod.rs |
| EMBER_SELF_LOOKUP_FIRST_DELAY_SECS | 180 s (20 s if ≥ 10 verified) | network/mod.rs |
| EMBER_SELF_LOOKUP_REPEAT_SECS | 4 h | network/mod.rs |
| EMBER_DISCONNECT_SECS | 20 min | network/mod.rs |
| EMBER_CONTACT_STALE_SECS | 2 h | network/mod.rs |
| SESSION_TIMEOUT | 900 s | transport.rs |
| MAX_SESSIONS_PER_ADDR | 4 | transport.rs |
| STORE_SIG_REPLAY_TTL | 60 s | dht/engine.rs |
| MIN_OBSERVED_IP_VOTES | 3 / 15 min | dht/observed.rs |
| Constant | Value | File |
|---|---|---|
| EMBER_PERSIST_MAX_CONTACTS | 200 | network/mod.rs |
| EMBER_PERSIST_MAX_RECORDS | 20,000 | network/mod.rs |
| NODES_EMBER_MAGIC | EMB3 | dht/bootstrap.rs |
| STORE_EMBER_MAGIC | EMS1 | dht/bootstrap.rs |
| FT_EMBER_SOURCE_PUBLISH | 0xE4 | known.met |
| FT_EMBER_KEYWORD_PUBLISH | 0xE5 | known.met |
| EMBER_RENDEZVOUS_SEED | ember-dht-rendezvous-v1 | kad/publish.rs |
| Area | Location |
|---|---|
| DHT engine, routing, store, search, publish, wire | src-tauri/src/network/ember/dht/ |
| Protocol constants (k, α, version, IDs) | dht/mod.rs |
| Frame encode/decode, message types, size budgets | dht/messages.rs |
| IO-free inbound handling, proximity, PROXY_STORE | dht/engine.rs |
| 128-bucket table, replacement cache, admission | dht/routing.rs |
| Signed records, capacity, TTL, serve cursor | dht/store.rs |
| Iterative lookup, FIND_VALUE key packing | dht/search.rs |
| Keyword/source record builders | dht/publish.rs |
| Filename / query tokenization | kad/publish.rs · extract_keywords() |
| Streaming BLAKE3 file digest | ember/crypto.rs · Blake3FileHasher |
| Adaptive abuse limits | dht/scale.rs |
| Per-IP / per-peer rate limits | dht/protection.rs |
| Observed-IP voting | dht/observed.rs |
| nodes_ember.dat / store_ember.dat | dht/bootstrap.rs |
| Noise transport, magic, sessions | ember/transport.rs |
| Node ID, sign/verify, hash binding | ember/crypto.rs |
| Dormant native transfer | ember/transfer.rs |
| Publish/search drivers, maintenance, KAD bridge | network/mod.rs |
| Rendezvous KAD key | kad/publish.rs · ember_rendezvous_key() |
| User-facing status | Ember Network page (/ember), status bar |
| Settings | Settings → Network · ember_native_enabled (always on) |
| Dev harness | add_ember_dht_contact, ember_dht_ping_peer, ember_dht_find_node, ember_dht_iterative_find_node, ember_dht_publish_keyword, ember_dht_find_value, ember_dht_run_maintenance |
| Standing work log | docs/ember-dht.md |