Retina Agent executes coordinated network probes to infer forwarding behavior across distributed vantage points, producing forwarding information elements (FIEs) for topology analysis.
The agent connects to an orchestrator via TCP, receives probing directives, executes network probes, and returns forwarding information elements (FIEs).
Part of the Retina system:
- Generator: Creates probing directives
- Orchestrator: Distributes directives to agents, collects FIEs
- Agent: Executes network probes (this component)
┌─────────────┐
│Orchestrator │
└──────┬──────┘
│ TCP (JSON over newline-delimited stream)
│
┌──────▼──────────────────────────────┐
│ Retina Agent │
│ │
│ ┌────────┐ ┌──────────┐ ┌──────┐ │
│ │ Reader │─▶│Processor │─▶│Writer│ │
│ └────────┘ └─────┬────┘ └──────┘ │
│ │ │
│ ┌─────▼─────┐ │
│ │ Prober │ │
│ │ (caracal) │ │
│ └───────────┘ │
└─────────────────────────────────────┘
Three-stage pipeline:
- Reader: Receives
ProbingDirectivemessages from orchestrator - Processor: Executes two probes per directive (near TTL, far TTL) in parallel, sends FIE once both complete or time out
- Writer: Sends
ForwardingInfoElementresults back to orchestrator
Key features:
- Non-blocking probe execution (thousands of concurrent probes, bounded by
--write-queue-sizeand OS limits) - Automatic reconnection with exponential backoff
- Graceful shutdown on SIGINT/SIGTERM
- Go 1.24.4
- For production: caracal and raw socket privileges
git clone https://github.com/dioptra-io/retina-agent
cd retina-agent
go build -o retina-agent ./cmd/retina-agent./retina-agent --id agent-1 --address localhost:50050 --prober-type mockUse the mock orchestrator to test the complete pipeline:
# Terminal 1: Start mock orchestrator
go run test/mock_orchestrator.go
# Terminal 2: Start agent with mock prober
./retina-agent --id agent-1 --address localhost:50050 --prober-type mock| Flag | Default | Description |
|---|---|---|
--id |
agent-1 |
Agent identifier |
--address |
localhost:50050 |
Orchestrator address (host:port) |
--prober-type |
caracal |
Prober: caracal or mock |
--prober-path |
(searches PATH) | Path to prober executable |
--prober-arg |
Additional argument to pass to the prober (repeatable) | |
--write-queue-size |
1000 |
Prober write queue buffer size |
--cleanup-interval |
10s |
Prober stale probe cleanup interval |
--pds-buffer |
100 |
Directives channel buffer size |
--fies-buffer |
100 |
FIEs channel buffer size |
--read-deadline |
10s |
Read timeout for orchestrator connection |
--write-deadline |
5s |
Write timeout for orchestrator connection |
--probe-timeout |
5s |
Timeout for individual probe responses |
--max-reconnect-backoff |
5m |
Maximum wait time between reconnection attempts |
--max-consecutive-decode-errors |
3 |
Max consecutive decode errors before reconnecting (0 to disable) |
--log-level |
info |
Log level (debug, info, warn, error) |
--metrics-addr |
:9312 |
Address to expose Prometheus metrics on |
See --help for all options.
All flags can be configured via environment variables. These act as defaults and are overridden by CLI flags.
Precedence:
CLI flags > environment variables > hardcoded defaults
| Variable | Default | Description |
|---|---|---|
RETINA_SECRET |
* | Shared secret for orchestrator authentication, required |
RETINA_ID |
agent-1 |
Agent identifier |
RETINA_ADDRESS |
localhost:50050 |
Orchestrator address (host:port) |
RETINA_PROBER_TYPE |
caracal |
Prober to use (caracal or mock) |
RETINA_PROBER_PATH |
(searches PATH) | Path to prober executable |
RETINA_WRITE_QUEUE_SIZE |
1000 |
Prober write queue buffer size |
RETINA_CLEANUP_INTERVAL |
10s |
Prober stale probe cleanup interval |
RETINA_PDS_BUFFER |
100 |
Directives channel buffer size |
RETINA_FIES_BUFFER |
100 |
FIEs channel buffer size |
RETINA_READ_DEADLINE |
10s |
Read timeout for orchestrator connection |
RETINA_WRITE_DEADLINE |
5s |
Write timeout for orchestrator connection |
RETINA_PROBE_TIMEOUT |
5s |
Timeout for individual probe responses |
RETINA_MAX_RECONNECT_BACKOFF |
5m |
Maximum wait between reconnection attempts |
RETINA_MAX_CONSECUTIVE_DECODE_ERRORS |
3 |
Max consecutive decode errors before reconnecting (0 to disable) |
RETINA_LOG_LEVEL |
info |
Log level (debug, info, warn, error) |
RETINA_METRICS_ADDR |
:9312 |
Address to expose Prometheus metrics on |
For each ProbingDirective:
- Launch two probes concurrently:
- Near probe: TTL =
directive.NearTTL - Far probe: TTL =
directive.NearTTL + 1
- Near probe: TTL =
- Correlate results by destination, protocol, header fields, TTL, and send timestamp
- Always send a FIE: if a probe times out, the corresponding
NearInfoorFarInfofield is nil
The caracal prober uses a high-throughput pipeline:
- Multiple goroutines queue probe requests (non-blocking)
- Single writer goroutine sends to caracal stdin (CSV format)
- Single reader goroutine receives from caracal stdout (CSV format)
- Results correlated back to waiting goroutines via shared map
- Network errors: Trigger reconnection with exponential backoff
- Decode errors: Log and skip (reconnect after
--max-consecutive-decode-errorsconsecutive) - Probe timeouts: Expected behavior, FIE sent with nil NearInfo/FarInfo
- Context cancellation: Clean shutdown
- Implement the
Proberinterface:
type Prober interface {
Probe(ctx context.Context, pd *api.ProbingDirective, ttl uint8) (*ProbeResult, error)
Close() error
}- Add to
createProber()inagent.go:
case "myprober":
return NewMyProber(cfg), nil- Use it:
./retina-agent --prober-type myproberMetrics are exposed at --metrics-addr (default :9312) in Prometheus format, covering pipeline throughput (directives received, FIEs sent), probe outcomes (success/timeout/error rates, RTT distribution), connectivity (reconnections, decode errors), and caracal internals (queue depth, in-flight probes, correlation failures). See internal/agent/metrics.go for the full list.
MIT License - see LICENSE for details