Same class of bug just fixed in step_pipeline_rules()/step_pipelines():
title-only existence check meant an alert's config (query, threshold,
group_by) could never actually change on a re-run once created. Confirmed
live: alert7's query was corrected in git (vendor:accel-ppp ->
vendor:accel-ppp AND event_type:*) but a subsequent re-run still reported
"already exists" with the old, buggy query untouched. Now compares by the
'config' object (not just title) and PUT-updates + re-enables the
schedule if it differs, same pattern as the other two fixes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- step_pipeline_rules() and step_pipelines() only ever created new
rules/pipelines by title, never updated existing ones whose content
changed under the same title. Confirmed live this session: 5 new rules
were added to rules/*.json and created fine in Graylog, but "Servers
Parsing"'s stage list was never updated to actually call them, since
the pipeline already existed. Both functions now compare by 'source'
content and PUT-update if it differs, matching the pattern already used
elsewhere (e.g. step_inputs()'s timezone self-heal).
- alert7's query (vendor:accel-ppp) was too broad: it also matches
routine traffic tagged only by the generic accelppp_interface_tag
fallback rule, which sets vendor but never event_type - producing false
"(Empty Value)" group_by matches on ordinary volume. Confirmed live.
Fixed to vendor:accel-ppp AND event_type:* - severity_tag was
considered instead but rejected, since accelppp_router_address_error
(the rule that inspired this alert) never set it either.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Analyzed a real 1GB accel-ppp log (2026-07-23): only 2,819 of its lines
were error:/warn:, and one pattern - "can't determine router address" -
repeated 2,746 times over ~4 hours for two specific subscriber interfaces
before self-resolving, completely unalerted since no alert covered it.
- New pipeline rules for 3 previously-unclassified real message types
(radius:dm_coa session not found, mac change detected, dhcpv4 short
packet) plus a text-based fallback pair (accelppp_unclassified_error/
warn) for anything not yet specifically classified - needed because
Vector-shipped accel-ppp lines have no real syslog PRI header, so the
numeric-severity generic_critical_severity rule never fires for this
source.
- New alert7: group by gl2_remote_ip + event_type, fires on >5 occurrences
in 5 minutes - low enough to have caught the real incident within its
first cycle, high enough to tolerate a single transient warning.
- vector-accel-ppp-setup.md gained a "keep only error/warn" filter option
(Step 2b), explicitly documented as a deliberate tradeoff: it also drops
RADIUS accounting, so the Servers & Sessions dashboard and session
correlation go empty for any server that applies it. Both the
with-filter and without-filter full configs are included so the choice
is per-server, not global.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Dropping every "[DHCPv4 " line unconditionally would also swallow a
genuine DHCP-side problem (pool exhaustion, lease timeout) if accel-ppp
ever logs one at error/warn level. Verified the refined condition against
sample lines before writing it: plain info-level DHCP still drops, RADIUS
lines are unaffected, and DHCP lines tagged error:/warn: now pass through.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Motivated by live capacity math: 2 active NAS servers measured at ~634
bytes/document in Graylog and ~75 msg/s each; projected to 10+ servers of
similar profile that's ~38.5 GiB/day, which the current 50GB disk / 14-21
day retention can't hold. DHCP request/response chatter (Discover/Request/
Offer/Ack) is the bulk of that volume and isn't needed for RADIUS/
accounting analytics, so dropping it via a Vector `filter` transform
(before the packet ever leaves the NAS server) is the cheapest place to
cut it - no wasted network/disk/CPU on data nobody queries.
Also cross-referenced two things diagnosed live earlier: the
read_from: beginning cold-start replay burst, and where to look if a
filtered-out message type stops appearing (intentional, not a bug).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The alert5 description and both READMEs said "~4,600-4,800 msgs/10min" as
the calibrated baseline - a transcription error, a digit short of the real
figure (~46,000-48,000/10min, matching the ~278k/hour number right next to
it, which was correct). The configured threshold (150,000) was actually
computed from the correct number, so no behavior changed - only the
description text was wrong.
Also documents a real event from today: EX-NAS-1-2's first-ever Vector
startup tripped this alert (430,982 msgs/10min) - Vector's file source
replays existing log content from the beginning with no checkpoint yet, so
a busy log's backlog shows up as a single burst rather than a real ongoing
issue. Confirmed the sample messages were all routine DHCP/session churn
timestamped around the same moment.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- step_compose_files() now uses GRAYLOG_ADMIN_PASSWORD as the actual admin
password on a fresh install if set, instead of always generating a
random one. Solves the CI secret-staleness problem at the root: pin a
password once and it's correct both at creation time and on every later
re-run, instead of a fresh install randomly generating a password the
stored secret then has to be manually kept in sync with.
- deploy-from-scratch.yml gained a `cores` workflow_dispatch input
(default 4), passed through to create-graylog-lxc.sh's --cores flag.
Only takes effect when the container is actually created fresh, same as
every other --cores usage in this project.
- Documented both in README.md/README.uk.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
deploy.yml already forwarded this secret; deploy-from-scratch.yml didn't,
so a re-run against an already-initialized Graylog (one-time credentials
file already consumed by an earlier successful run) died at
resolve_admin_password with "Cannot find admin password" - confirmed
live.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- New "Ways to run this" comparison table up top
- Step 9 split into 9a (in-container runner) / 9b (host-level runner,
with the RUNNER_USER=claude-deploy warning and single-line paste-safe
command) / 9c (secrets) / removal instructions
- CI/CD section split into deploy.yml and deploy-from-scratch.yml
subsections, each explaining its own runner/scope, plus what's common
to both
- File layout updated with every script and workflow file added this
session (bootstrap-host.sh, cleanup-host.sh, fix-lxc-apparmor.sh,
setup-forgejo-runner.sh, alerts/, dashboards/, .forgejo/workflows/)
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Introduced by an earlier edit to .admin_credentials_ONE_TIME's own
instructional text ("...you've stored the password: rm <path>"), which
itself contains "password: " - the same substring resolve_admin_password()
greps for. grep -oP matched BOTH occurrences; command substitution joined
them with a real newline, producing a corrupted two-line "password" that
never matched .env's GRAYLOG_ROOT_PASSWORD_SHA2, causing every gcurl call
to silently 401 and wait_for_api_ready to time out no matter how generous
the timeout was (confirmed by re-verifying live: the real password,
extracted correctly, hashes to exactly what .env already has).
Fixed both ends: reworded the instructional text to not repeat "password:",
and hardened the regex itself (^Graylog admin password: anchor, grep -m1)
so a future wording change can't reintroduce the same class of bug.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Confirmed live: on a CI-triggered run, the REST API genuinely took longer
than 60s to accept authenticated requests even though the container had
already reported healthy - verified after the fact that both the API and
the admin credentials were fine, this was purely insufficient margin
(likely due to concurrent load during the run), not a logic bug. 180s
matches the same order of magnitude as the existing 300s docker-health
wait elsewhere in this script.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
"the repo's Actions" inside FORGEJO_RUNNER_TOKEN's :? error message threw
off bash's parser even though the whole expression sits inside double
quotes - confirmed live via bisection (bash -n on truncated line ranges
pinpointed line 20 exactly, "unexpected EOF while looking for matching
`''`"). Single quotes inside \${VAR:?message} aren't neutralized by the
outer double quotes the way they would be in a plain string. Reworded to
avoid the apostrophe entirely.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Docker's healthcheck can report graylog-server "healthy" a few seconds
before the REST API is actually ready to serve authenticated requests -
confirmed live: the first gcurl call (step_index_retention) intermittently
got an empty response body, crashing the downstream `python3 -c
"json.load(sys.stdin)"` with "Expecting value: line 1 column 1".
Adds wait_for_api_ready(), polling the same endpoint step_index_retention
already needs (up to 60s) before proceeding, mirroring the existing
retry-loop pattern already used for docker compose pull.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
setup-forgejo-runner.sh gained RUNNER_LABEL/RUNNER_DIR/SERVICE_NAME/
RUNNER_USER params so the same script can register either kind of runner:
- inside the container (unchanged defaults, root - already scoped to just
that container)
- on the Proxmox host itself, where RUNNER_USER=claude-deploy is required:
a root-owned systemd service with no User= would hand every CI job
unrestricted root on the host, defeating the whole point of
claude-deploy's narrowly-scoped sudoers rules.
deploy-from-scratch.yml runs create-graylog-lxc.sh on the host-level
runner. Deliberately does NOT run pct destroy - that stays a manual,
deliberate human step. The idempotent create+install path is safe to
trigger any time: repairs an existing container in place, or fully
recreates one if it was destroyed beforehand.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- export LC_ALL=C.UTF-8 LANG=C.UTF-8 at the top of the script: the
container inherits LANG=en_US.UTF-8 from pct exec's calling shell but
never generates that locale, so every apt-get call printed "Setting
locale failed" warnings from perl/apt-listchanges. C.UTF-8 is glibc-
builtin, no locale-gen needed.
- The final summary now prints the generated admin user/password directly
and deletes .admin_credentials_ONE_TIME right after, instead of just
pointing at the file and leaving it for the operator to read and clean
up by hand.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Mirrors bootstrap-host.sh in reverse: removes the sudoers rule and the
fix-lxc-apparmor.sh script it installed, so the elevated grant only stands
for the duration of an active deployment instead of indefinitely.
create-graylog-lxc.sh already degrades cleanly to its manual fallback when
the automation isn't present, so this is a safe no-op-adjacent revoke.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds bootstrap-host.sh (one-time, run as root on a fresh Proxmox host) and
fix-lxc-apparmor.sh, the fixed-content script it installs. The sudoers
rule it wires up only ever invokes that one root-owned script with a VMID
argument - deliberately not a broader rule like `tee -a <conf>` or
`sh -c '...'`, since those only restrict the command's own argv, not
stdin/heredoc content, letting the caller write arbitrary lines to any
200-299 container's config instead of just this one fixed line.
create-graylog-lxc.sh now tries `sudo -n fix-lxc-apparmor.sh` first and
falls back to the existing manual instructions if that sudoers rule isn't
present yet - fully backward compatible with hosts that haven't run
bootstrap-host.sh.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The CI runner setup was done by hand this session (download, verify,
register, systemd unit) - script it the same idempotent way as
install-graylog.sh so a fresh deployment can reproduce it instead of
requiring manual SSH archaeology. Wire it into create-graylog-lxc.sh's
payload copy, and add step 9 to both READMEs covering the one-time setup.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Forgejo auto-issues a per-job, repo-scoped secrets.GITHUB_TOKEN (created at
workflow start, destroyed at completion, usable only against this repo);
use it in the clone URL instead of an anonymous HTTPS clone, which would
start failing with 401/403 the moment the repo's visibility changes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
actions/checkout@v4 is a Node.js-based action; the self-hosted runner
executes in host mode directly on the Graylog appliance container, which
has no Node.js and shouldn't need one just to check out a repo. A plain
git clone avoids the dependency entirely.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Two flood-detection alerts (per-source message volume, calibrated live
against real traffic) grouped by gl2_remote_ip
- Session correlation: accelppp_interface fallback tagging plus
radius_session_id/calling_station_id/radius_username extraction, so a
subscriber's full session lifecycle is searchable by one key
- Replace the single combined dashboard with three focused ones (Overview
& Alerts, Network Equipment, Servers & Sessions)
- Propagate GRAYLOG_ROOT_TIMEZONE and IP-in-alerts fixes into the reusable
install script and templates
- Add a Forgejo Actions workflow (manual trigger) that re-runs
install-graylog.sh on a self-hosted runner living in the container,
automating the deploy step this project has done by hand all along
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>