Both rule89 (a10_session_opened) and rule91 (a10_session_timeout)
failed to compile live: \\" (two backslashes before a quote) is
read by Graylog's rule DSL as an escaped backslash followed by an
unescaped string terminator, not an escaped quote - so the regex()
string literal ended early and everything after it parsed as
garbage ("Unknown function S", "mismatched input '('", etc).
Fix: exactly one backslash before each quote (\") so the DSL treats
it as an escaped quote character, matching the \S/\d/\. occurrences
elsewhere in the same pattern which correctly use two backslashes
(DSL-decodes to one, which is what the regex engine needs). The
other 4 new A10 rules didn't have this issue and already compiled
successfully on the user's first live run.
8 new rules (vendor=a10) wired into Network Equipment Parsing: LSN
TCP/Session/ICMP per-user quota exceeded (critical - real service
impact, drops new connections for that subscriber), BGP-4-MAXPFX
prefix-limit warning, and admin session open/close/timeout/auth-success
(aXAPI and CLI both covered by one pattern each).
Unlike the CSV-report-derived rules, these are built directly from
real captured A10 log output the user provided, so confidence is
higher - closer to the accel-ppp rules' provenance. Multi-entry LSN
lines (several 'ip(count)' pairs in one quota-exceeded message) only
have their first pair extracted into fields; the full list stays in
the raw message.
Not live-verified - the user is bringing the target system up
themselves this time rather than through the test container used
earlier in this branch of work.
Closes the parsing gap the README explicitly called out (no D-Link
parsing, no ZTE ONU alarms) plus adds BDCOM GPON and expands Juniper
coverage (DDoS, PSU/memory/ASIC hardware faults, LACP/BGP/SNMP, config
commit). 58 new rules across 5 vendors, wired into Network Equipment
Parsing's stage 0 ahead of the generic_critical_severity fallback.
Where the same real-world event is reported by multiple vendors
(dying_gasp, onu_offline, optical_low_power, cli_login/cli_logout,
config_saved, interface_link_state, lag_state_change), rules share one
event_type value so dashboards can aggregate across vendors, same
normalization approach as accelppp_interface.
Built directly from the user's CSV signature report, not from real
device log samples - each rule's description says so explicitly. `when`
conditions use plain substring/contains matching on the report's own
pattern text to keep classification robust; regex field extraction is
only added where the source format is unambiguous. Passed offline
checks (JSON validity, every pipeline-referenced rule resolves to a
file, all regex patterns compile). Live compilation against a running
Graylog instance - which caught 2 real bugs during the dashboard/stream
fixes earlier this session - could NOT be completed: the test container
went unreachable mid-session. Re-run install-graylog.sh once it's back
up to confirm these compile before relying on them.
Analyzed a real 1GB accel-ppp log (2026-07-23): only 2,819 of its lines
were error:/warn:, and one pattern - "can't determine router address" -
repeated 2,746 times over ~4 hours for two specific subscriber interfaces
before self-resolving, completely unalerted since no alert covered it.
- New pipeline rules for 3 previously-unclassified real message types
(radius:dm_coa session not found, mac change detected, dhcpv4 short
packet) plus a text-based fallback pair (accelppp_unclassified_error/
warn) for anything not yet specifically classified - needed because
Vector-shipped accel-ppp lines have no real syslog PRI header, so the
numeric-severity generic_critical_severity rule never fires for this
source.
- New alert7: group by gl2_remote_ip + event_type, fires on >5 occurrences
in 5 minutes - low enough to have caught the real incident within its
first cycle, high enough to tolerate a single transient warning.
- vector-accel-ppp-setup.md gained a "keep only error/warn" filter option
(Step 2b), explicitly documented as a deliberate tradeoff: it also drops
RADIUS accounting, so the Servers & Sessions dashboard and session
correlation go empty for any server that applies it. Both the
with-filter and without-filter full configs are included so the choice
is per-server, not global.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Two flood-detection alerts (per-source message volume, calibrated live
against real traffic) grouped by gl2_remote_ip
- Session correlation: accelppp_interface fallback tagging plus
radius_session_id/calling_station_id/radius_username extraction, so a
subscriber's full session lifecycle is searchable by one key
- Replace the single combined dashboard with three focused ones (Overview
& Alerts, Network Equipment, Servers & Sessions)
- Propagate GRAYLOG_ROOT_TIMEZONE and IP-in-alerts fixes into the reusable
install script and templates
- Add a Forgejo Actions workflow (manual trigger) that re-runs
install-graylog.sh on a self-hosted runner living in the container,
automating the deploy step this project has done by hand all along
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>