Fix wrong baseline figure in flood alert docs; document cold-start replay
The alert5 description and both READMEs said "~4,600-4,800 msgs/10min" as the calibrated baseline - a transcription error, a digit short of the real figure (~46,000-48,000/10min, matching the ~278k/hour number right next to it, which was correct). The configured threshold (150,000) was actually computed from the correct number, so no behavior changed - only the description text was wrong. Also documents a real event from today: EX-NAS-1-2's first-ever Vector startup tripped this alert (430,982 msgs/10min) - Vector's file source replays existing log content from the beginning with no checkpoint yet, so a busy log's backlog shows up as a single burst rather than a real ongoing issue. Confirmed the sample messages were all routine DHCP/session churn timestamped around the same moment. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
parent
eb6225f2f8
commit
a24defe596
3 changed files with 34 additions and 6 deletions
19
README.md
19
README.md
|
|
@ -469,7 +469,7 @@ compared against its own volume, not the whole stream's total.
|
||||||
The thresholds are not guessed - they were calibrated live on 2026-07-22
|
The thresholds are not guessed - they were calibrated live on 2026-07-22
|
||||||
against real traffic, via a Views API pivot search grouped by
|
against real traffic, via a Views API pivot search grouped by
|
||||||
`gl2_remote_ip`:
|
`gl2_remote_ip`:
|
||||||
- The one active accel-ppp/RADIUS server was steadily sending **~4,600-4,800
|
- The one active accel-ppp/RADIUS server was steadily sending **~46,000-48,000
|
||||||
messages per 10 minutes** (~278k/hour) under normal load. The Servers
|
messages per 10 minutes** (~278k/hour) under normal load. The Servers
|
||||||
stream threshold (150,000/10min) gives roughly 3x headroom above that.
|
stream threshold (150,000/10min) gives roughly 3x headroom above that.
|
||||||
- The one active network device was steadily sending **~30-60 messages per
|
- The one active network device was steadily sending **~30-60 messages per
|
||||||
|
|
@ -480,8 +480,21 @@ These are starting points based on partial rollout (1 server + 1 device
|
||||||
active at calibration time). Revisit both thresholds once more of the
|
active at calibration time). Revisit both thresholds once more of the
|
||||||
planned ~10-15 servers and ~10-20 switches/OLTs are sending real traffic -
|
planned ~10-15 servers and ~10-20 switches/OLTs are sending real traffic -
|
||||||
what looks like 3x headroom today could be too tight or too loose once
|
what looks like 3x headroom today could be too tight or too loose once
|
||||||
every server's individual baseline is known. Check current per-source
|
every server's individual baseline is known.
|
||||||
volume any time with a query like:
|
|
||||||
|
**A server's first-ever Vector startup can trip this once, harmlessly** -
|
||||||
|
confirmed live on 2026-07-23 when a second NAS server (EX-NAS-1-2) came
|
||||||
|
online: Vector's file source reads existing log content from the
|
||||||
|
beginning on a fresh run (no checkpoint yet), so a busy accel-ppp log
|
||||||
|
already containing hours of routine DHCP/session churn gets replayed as a
|
||||||
|
single burst - 430,982 messages in one 10-minute window in that case, all
|
||||||
|
ordinary (`DHCPv4 Ack`, `ipoe stop/remove`), all timestamped right around
|
||||||
|
the same moment. Not a loop, not a misconfiguration - just backlog replay
|
||||||
|
on cold start. Check whether it's settled with the same per-source query
|
||||||
|
below; if volume stays elevated past the first window or two, treat it as
|
||||||
|
a real incident instead.
|
||||||
|
|
||||||
|
Check current per-source volume any time with a query like:
|
||||||
```
|
```
|
||||||
gl2_remote_ip:<ip>
|
gl2_remote_ip:<ip>
|
||||||
```
|
```
|
||||||
|
|
|
||||||
19
README.uk.md
19
README.uk.md
|
|
@ -485,7 +485,7 @@ Graylog відхиляє його з помилкою Jackson-поліморфі
|
||||||
|
|
||||||
Пороги не вигадані — вони каліброві наживо 2026-07-22 на реальному
|
Пороги не вигадані — вони каліброві наживо 2026-07-22 на реальному
|
||||||
трафіку через pivot-пошук Views API, згрупований за `gl2_remote_ip`:
|
трафіку через pivot-пошук Views API, згрупований за `gl2_remote_ip`:
|
||||||
- Один активний accel-ppp/RADIUS сервер стабільно слав **~4 600-4 800
|
- Один активний accel-ppp/RADIUS сервер стабільно слав **~46 000-48 000
|
||||||
повідомлень за 10 хвилин** (~278 тис/год) за нормального навантаження.
|
повідомлень за 10 хвилин** (~278 тис/год) за нормального навантаження.
|
||||||
Поріг для стріму Servers (150 000/10хв) дає приблизно 3x запасу над цим.
|
Поріг для стріму Servers (150 000/10хв) дає приблизно 3x запасу над цим.
|
||||||
- Один активний мережевий пристрій стабільно слав **~30-60 повідомлень за
|
- Один активний мережевий пристрій стабільно слав **~30-60 повідомлень за
|
||||||
|
|
@ -497,7 +497,22 @@ Graylog відхиляє його з помилкою Jackson-поліморфі
|
||||||
по мірі підключення решти запланованих ~10-15 серверів і ~10-20
|
по мірі підключення решти запланованих ~10-15 серверів і ~10-20
|
||||||
комутаторів/OLT — те, що сьогодні виглядає як 3x запасу, може виявитись
|
комутаторів/OLT — те, що сьогодні виглядає як 3x запасу, може виявитись
|
||||||
затісним або занадто вільним, коли буде відомий власний baseline кожного
|
затісним або занадто вільним, коли буде відомий власний baseline кожного
|
||||||
сервера. Перевірити поточний обсяг по джерелу можна будь-коли запитом
|
сервера.
|
||||||
|
|
||||||
|
**Перший запуск Vector на новому сервері може спрацювати цей алерт один
|
||||||
|
раз, безпечно** — підтверджено наживо 2026-07-23, коли другий NAS-сервер
|
||||||
|
(EX-NAS-1-2) вийшов онлайн: file-джерело Vector при свіжому запуску (без
|
||||||
|
чекпоінта) читає наявний вміст логу від самого початку, тож завантажений
|
||||||
|
accel-ppp-лог із годинами рутинного DHCP/session-трафіку відтворюється
|
||||||
|
одним сплеском — у тому випадку 430 982 повідомлення за одне 10-хвилинне
|
||||||
|
вікно, усі звичайнісінькі (`DHCPv4 Ack`, `ipoe stop/remove`), усі з
|
||||||
|
міткою часу приблизно в один момент. Не цикл, не помилка конфігурації —
|
||||||
|
просто відтворення накопиченого при холодному старті. Перевірте, чи
|
||||||
|
вляглося, тим самим запитом по джерелу нижче; якщо обсяг лишається
|
||||||
|
підвищеним довше одного-двох вікон — тоді вже трактуйте як реальний
|
||||||
|
інцидент.
|
||||||
|
|
||||||
|
Перевірити поточний обсяг по джерелу можна будь-коли запитом
|
||||||
на кшталт:
|
на кшталт:
|
||||||
```
|
```
|
||||||
gl2_remote_ip:<ip>
|
gl2_remote_ip:<ip>
|
||||||
|
|
|
||||||
|
|
@ -1,6 +1,6 @@
|
||||||
{
|
{
|
||||||
"title": "WARNING: abnormal message volume from one server",
|
"title": "WARNING: abnormal message volume from one server",
|
||||||
"description": "A single server is sending far more syslog messages than its established baseline - possible log loop, misconfiguration, or a genuine incident generating excessive events. Left unchecked this can fill the retention window's disk budget early. Threshold calibrated live on 2026-07-22 against real accel-ppp/RADIUS traffic: the one active NAS server was steadily sending ~4,600-4,800 msgs/10min (~278k/hour) under totally normal load, so the threshold is set to roughly 3x that baseline.",
|
"description": "A single server is sending far more syslog messages than its established baseline - possible log loop, misconfiguration, or a genuine incident generating excessive events. Left unchecked this can fill the retention window's disk budget early. Threshold calibrated live on 2026-07-22 against real accel-ppp/RADIUS traffic: the one active NAS server was steadily sending ~46,000-48,000 msgs/10min (~278k/hour) under totally normal load, so the threshold is set to roughly 3x that baseline. Note: a server's first-ever Vector startup can trigger this once, harmlessly - Vector's file source reads existing log content from the start on a fresh run, producing a one-time burst of backlog rather than a real ongoing issue.",
|
||||||
"priority": 2,
|
"priority": 2,
|
||||||
"alert": true,
|
"alert": true,
|
||||||
"config": {
|
"config": {
|
||||||
|
|
|
||||||
Loading…
Add table
Reference in a new issue