Summary
A gean aggregator produced one aggregated-attestation proof that every other node on the devnet rejected, including the three other gean nodes. It happened once in five hours and did not recur, but it is a client emitting an unverifiable proof, so it is worth understanding before it appears somewhere that matters.
ethlambda names the failure precisely:
Aggregated signature verification failed: verification failed: Snark(Bus(Gkr(LayerMismatch { layer: 2 })))
That is a leanVM proof-structure error, not a signer-set or claim mismatch.
Observed
Mixed devnet, 4 gean (all four aggregators) + 6 ethlambda + 4 lantern, 4 committees, 14 validators. All clients on leanVM 48a904208d682848dac0e18ef8b01ebfc40df9ad. gean 14da17a (PR #441 head). Five hours of runtime, chain healthy throughout: head 4158-4161, finalized 4150-4153, identical across all three clients, zero restarts.
At one instant, all four aggregators built an aggregate for the same slot from the same shape of input:
18:47:31.099 gean_0 aggregate: slot=3786 raw=10 children=0 total=10 proof=160696 bytes duration=1.208s
18:47:31.207 gean_2 aggregate: slot=3786 raw=10 children=0 total=10 proof=159864 bytes duration=1.321s
18:47:31.208 gean_3 aggregate: slot=3786 raw=10 children=0 total=10 proof=159864 bytes duration=1.319s
18:47:31.307 gean_1 aggregate: slot=3786 raw=10 children=0 total=10 proof=159864 bytes duration=1.419s
Three proofs are byte-identical in length. gean_0's is 832 bytes larger.
44 ms after gean_0 published, every other node rejected an aggregate:
| node |
rejections |
| gean_0 |
0 — the producer |
| gean_1, gean_2, gean_3 |
1 each |
| ethlambda_0 … ethlambda_5 |
1 each |
| lantern_0 … lantern_3 |
0 — did not receive it |
Nine rejections, all within 18:47:31.143-18:47:31.190, all naming the same error. One bad aggregate, rejected by everyone that saw it.
gean's own log line is generic:
ERROR [signature] aggregated attestation verification failed: aggregated signature verification failed
The four lantern nodes logged nothing, which is most likely gossip fan-out rather than immunity.
Why this is not the same-slot bug
Zero occurrences of two claims at one epoch / two different messages in one aggregate across the whole run on any client. This is a different failure: the proof itself is structurally wrong, not the claim grouping.
What is not known
- Whether gean_0's inputs differed from the other three. All four report
raw=10 children=0, but that does not prove the same ten signatures were selected. If the inputs were identical, a differing proof size points at non-determinism in proving; if they differed, the question is why one input set produces an invalid proof.
- What
LayerMismatch { layer: 2 } corresponds to in leanVM's GKR layering, and whether the 832-byte size difference is diagnostic of it.
- Whether this correlates with aggregate shape. This one was
raw=10 children=0, the largest raw count seen in the run; most aggregates in the same period were raw=2..9. Worth checking whether it only appears at the upper end.
- Whether gean_0 being the sole aggregator that also holds subnet 0 matters.
Reproducing
It occurred once in five hours across four aggregators, so it is rare. A targeted attempt would be to drive aggregation with raw=10 repeatedly and verify each proof locally after building it, rather than waiting for a peer to reject one.
A cheap mitigation regardless of root cause: have an aggregator verify its own proof before gossiping it, and log the inputs when that fails. That turns a silent bad publish into a local diagnostic with the input set attached.
Environment
- gean
14da17a, image ghcr.io/geanlabs/gean:perf
- ethlambda image
006453f4, lantern image built from 9ce166c4
- leanVM
48a904208d682848dac0e18ef8b01ebfc40df9ad on all three
- 14 validators, 4 committees, gean holding all four aggregator slots
Summary
A gean aggregator produced one aggregated-attestation proof that every other node on the devnet rejected, including the three other gean nodes. It happened once in five hours and did not recur, but it is a client emitting an unverifiable proof, so it is worth understanding before it appears somewhere that matters.
ethlambda names the failure precisely:
That is a leanVM proof-structure error, not a signer-set or claim mismatch.
Observed
Mixed devnet, 4 gean (all four aggregators) + 6 ethlambda + 4 lantern, 4 committees, 14 validators. All clients on leanVM
48a904208d682848dac0e18ef8b01ebfc40df9ad. gean14da17a(PR #441 head). Five hours of runtime, chain healthy throughout: head 4158-4161, finalized 4150-4153, identical across all three clients, zero restarts.At one instant, all four aggregators built an aggregate for the same slot from the same shape of input:
Three proofs are byte-identical in length. gean_0's is 832 bytes larger.
44 ms after gean_0 published, every other node rejected an aggregate:
Nine rejections, all within 18:47:31.143-18:47:31.190, all naming the same error. One bad aggregate, rejected by everyone that saw it.
gean's own log line is generic:
The four lantern nodes logged nothing, which is most likely gossip fan-out rather than immunity.
Why this is not the same-slot bug
Zero occurrences of
two claims at one epoch/two different messages in one aggregateacross the whole run on any client. This is a different failure: the proof itself is structurally wrong, not the claim grouping.What is not known
raw=10 children=0, but that does not prove the same ten signatures were selected. If the inputs were identical, a differing proof size points at non-determinism in proving; if they differed, the question is why one input set produces an invalid proof.LayerMismatch { layer: 2 }corresponds to in leanVM's GKR layering, and whether the 832-byte size difference is diagnostic of it.raw=10 children=0, the largest raw count seen in the run; most aggregates in the same period wereraw=2..9. Worth checking whether it only appears at the upper end.Reproducing
It occurred once in five hours across four aggregators, so it is rare. A targeted attempt would be to drive aggregation with
raw=10repeatedly and verify each proof locally after building it, rather than waiting for a peer to reject one.A cheap mitigation regardless of root cause: have an aggregator verify its own proof before gossiping it, and log the inputs when that fails. That turns a silent bad publish into a local diagnostic with the input set attached.
Environment
14da17a, imageghcr.io/geanlabs/gean:perf006453f4, lantern image built from9ce166c448a904208d682848dac0e18ef8b01ebfc40df9adon all three