A must-gather image for collecting diagnostic information from Red Hat Advanced Cluster Security (RHACS / StackRox) deployments on OpenShift clusters.
oc adm must-gather --image=quay.io/rhn_support_shaising/acs-must-gather:latestThis collects data related to ACS components only. For general cluster diagnostics, run oc adm must-gather without a custom image.
The image collects three complementary layers of data:
- acs-must-gather collects platform-specific data about the OpenShift cluster where RHACS is running (operator, workloads, cluster-scoped resources, RBAC, logs).
- acs-diagnostic-bundle collects deep RHACS-related data — the full diagnostic bundle produced by
roxctl central debug download-diagnostics— from Central and every connected Secured Cluster. - acs-debug-dump collects Central's debug dump — the deepest, Central-focused profiling data produced by
roxctl central debug dump, including a 30-second CPU profile.
# Collect logs from the last 8 hours
oc adm must-gather --image=quay.io/rhn_support_shaising/acs-must-gather:latest -- /usr/bin/gather MUST_GATHER_SINCE=8h
# Collect logs since a specific time
oc adm must-gather --image=quay.io/rhn_support_shaising/acs-must-gather:latest -- /usr/bin/gather MUST_GATHER_SINCE_TIME=2024-01-15T10:00:00Z- RHACS operator namespace (pods, logs, deployments, events)
- OLM resources (ClusterServiceVersion, Subscription, InstallPlan)
- Central deployment, Central DB
- Scanner V4 (indexer, matcher, database)
- Legacy Scanner (if present)
- ConfigMaps, Services, Routes, NetworkPolicies, PVCs, HPAs
- Sensor deployment
- Collector DaemonSet
- Admission Controller deployment
- NetworkPolicies
- ACS CRDs and CR instances (Central, SecuredCluster, SecurityPolicy)
- ValidatingWebhookConfigurations and MutatingWebhookConfigurations
- ClusterRoles and ClusterRoleBindings
- SecurityContextConstraints (OpenShift)
- Node summary and a node-to-kernel matrix (OS image + kernel version per node)
- ClusterVersion (OpenShift), StorageClasses, and PersistentVolumes
/v1/metadata— Central version and build info/v1/clusters— connected cluster list/v1/centralhealth/upgradestatus— upgrade/rollback status/v1/database/status— database health/debug/goroutine— goroutine stack dump/debug/heap— heap memory profile
Extracted into the acs-diagnostic-bundle/ folder. This is the full diagnostic
bundle produced by roxctl central debug download-diagnostics, downloaded from
Central's /api/extensions/diagnostics endpoint and unpacked so it is browsable
within the must-gather. It contains deep RHACS data from Central and every
connected Secured Cluster, including:
- Build versions and Central-DB (PostgreSQL) diagnostics (
pg_statstatistics) - Prometheus metrics and heap / goroutine / mutex profiles
- System configuration, scrubbed auth providers, roles, and notifiers
- Telemetry data
- Kubernetes introspection (resource manifests, pod logs, events) from Central and each Secured Cluster
Central serves this endpoint with admin authentication only. The password is read
from the central-htpasswd secret (falling back to stackrox-admin-password), and
Central is reached over an oc port-forward since Central container images do not
ship curl.
Extracted into the acs-debug-dump/ folder. This is Central's debug dump,
produced by roxctl central debug dump and downloaded from Central's
/debug/dump endpoint, then unpacked so it is browsable within the must-gather.
It is the deepest, Central-focused diagnostic layer and includes:
- A 30-second CPU profile, plus heap, goroutine, and mutex profiles
- Two Prometheus metrics passes and Central-DB (PostgreSQL) data
- Central version, system access control, notifiers, and log-imbue data
Because the CPU profile briefly adds load to Central, this collector can be
turned off with GATHER_DEBUG_DUMP=false. Like the diagnostic bundle, it
authenticates with the admin password from the central-htpasswd secret
(falling back to stackrox-admin-password) over an oc port-forward, since
Central container images do not ship curl.
Extracted into the advanced-acs-diagnostics/ folder. This is an advanced
layer collected in addition to the officially-supported diagnostic bundle and
debug dump above. It targets data those layers cannot provide — most importantly
data that does not depend on Central being reachable. Each sub-collector is
best-effort and isolated; a failure in this layer never affects the rest of the
must-gather. Disable the whole layer with GATHER_ADVANCED=false.
secured-cluster-local/— Sensor, Collector, and Admission Controller data collected directly from the pods, without requiring Central or an admin login — exactly what is missing when Sensor cannot reach Central. Includes Sensor's pprof heap/goroutine dumps and its cluster-entities store, Prometheus/metricsfrom Sensor / Admission Controller / Collector, the Collector probe/driver type (COLLECTION_METHOD) and pod state, a grep'd connectivity/certificate summary from the Sensor log (includingcertificate signed by unknown authority, the CA-rotation signature), and a count of thebelongs to 2 or more deploymentsduplicate-IP warning (a known Sensor memory-growth driver). Reached overoc port-forward, which can also read the loopback-only debug servers.tls-certs/— an X.509 expiry report (cert-expiry-summary.txt) for the RHACS service certificates.oc adm inspectredacts secrets, so expired or mismatched certs are otherwise invisible. Only public certificate material is read (private keys are never decoded), and certs expiring within 30 days are flagged.crash-upgrade-forensics/— previous-container logs for restarted containers,oc describe podoutput (OOMKilled / FailedScheduling), thesensor-upgraderdeployment / logs / RBAC, and Central's Administration Events feed (best-effort; needs Central).scanner-v4/— Central-independent Scanner V4 triage, collected directly from the indexer / matcher / db pods. Includes a pod-status table (phase / ready / restarts / last-state / image), each component's/health/readiness(HTTPS 9443) and Prometheus/metrics(9091, best-effort — secure metrics may need a client cert, in which case an.erroris written), vulnerability-updater / definitions markers grep'd from the component logs,oc describefor the indexer / matcher / db deployments, a memory-tuning summary (requests / limits +GOMEMLIMIT, to correlate matcher OOMs during VEX feed updates), and the componentService/Endpoints(a matcher or indexer that never becomes ready shows up as a Service with no ready endpoints). Surfaces stuck vulnerability updates and never-ready or crash-looping Scanner V4 components that the Central-focused bundle misses. Reached overoc port-forward.vuln-report/— an image-scan / CVE and policy-violation snapshot pulled from Central's REST API (admin basic-auth overoc port-forward, the same mechanism as the debug dump). The officially-supported bundle describes Central's health; it does not carry the per-image CVE findings or the violation list a support case usually turns on. Includesvuln-mgmt-workloads.json(streaming/v1/export/vuln-mgmt/workloads— every deployment joined to its images with full CVE data, the machine-readable dataset the analyzer filters),image-cves.csv(human-readable image CVEs, opens in any spreadsheet),violations.json(policy violations, paged), andalerts-summary-counts-*.json(violation rollups by cluster / category). Violations default toACTIVE,ATTEMPTEDto keep the bundle bounded on long-lived clusters — setADV_VULN_ALERT_STATESfor a fuller history.platform/— platform scoping, storage, and startup forensics that map to recurring support cases but thatoc adm inspectdoes not surface cleanly: theinit-dbinit-container log (current + previous) for every RHACS PostgreSQL database —central-db,scanner-db, andscanner-v4-db— since a permission error on the data volume is the classicdb-initCrashLoopBackOff and the log is lost once the pod is recreated, the Sensorcrsinit-container log (CRS-based cluster registration and cert setup — a failure there stops the Secured Cluster from registering or connecting), per-database Postgres migration / lock / slow-upgrade markers,oc describe pvc(binding/provisioning events the PVC yaml does not spell out), the OpenShift internal-registry CA configmap (image-registry-certificates, which lives outside the ACS namespaces and is needed to debugx509: certificate signed by unknown authorityafter a CA rotation), and a single chronologically-sorted events file per namespace.
| Variable | Description | Default |
|---|---|---|
MUST_GATHER_SINCE |
Duration filter for logs (e.g., 8h, 30m) |
(all logs) |
MUST_GATHER_SINCE_TIME |
ISO 8601 timestamp for log start | (all logs) |
MUST_GATHER_DIR |
Output base directory | /must-gather |
GATHER_DIAGNOSTICS |
Enable Central diagnostic endpoint collection | true |
GATHER_DIAGNOSTIC_BUNDLE |
Enable RHACS diagnostic bundle collection | true |
INSPECT_TIMEOUT |
Timeout per oc adm inspect call (seconds) |
120 |
DIAG_TIMEOUT |
Timeout per diagnostic endpoint call (seconds) | 30 |
DIAG_BUNDLE_TIMEOUT |
Timeout for the diagnostic bundle download (seconds) | 300 |
DIAG_BUNDLE_SINCE |
RFC3339 log start time for the bundle (--since) |
MUST_GATHER_SINCE_TIME |
DIAG_BUNDLE_CLUSTERS |
Comma-separated Secured Cluster names to include (--clusters) |
(all clusters) |
DIAG_BUNDLE_COMPLIANCE_OPERATOR |
Include Compliance Operator resources (--with-compliance-operator) |
false |
DIAG_BUNDLE_DATABASE_ONLY |
Collect only Central-DB diagnostics (--with-database-only) |
false |
GATHER_DEBUG_DUMP |
Enable Central debug dump collection | true |
DEBUG_DUMP_TIMEOUT |
Timeout for the debug dump download (seconds) | 300 |
DEBUG_DUMP_LOGS |
Include Central logs in the dump (roxctl central debug dump --logs) |
false |
GATHER_ADVANCED |
Enable the advanced ACS diagnostics layer | true |
GATHER_ADV_SECURED_CLUSTER |
Enable Central-independent secured-cluster collection | true |
GATHER_ADV_TLS_CERTS |
Enable the TLS certificate expiry report | true |
GATHER_ADV_FORENSICS |
Enable crash & upgrade forensics collection | true |
GATHER_ADV_SCANNER_V4 |
Enable Central-independent Scanner V4 health collection | true |
GATHER_ADV_VULN_REPORT |
Enable the image-CVE / violation snapshot from Central's API | true |
GATHER_ADV_PLATFORM |
Enable platform scoping, storage & startup forensics (db-init log, PVC describe, registry CA, sorted events) | true |
ADV_SC_TIMEOUT |
Timeout per secured-cluster endpoint call (seconds) | DIAG_TIMEOUT (30) |
ADV_FORENSICS_TIMEOUT |
Timeout per forensics call (seconds) | DIAG_TIMEOUT (30) |
ADV_SCANNER_V4_TIMEOUT |
Timeout per Scanner V4 endpoint call (seconds) | DIAG_TIMEOUT (30) |
ADV_VULN_TIMEOUT |
Timeout per vuln-report API call (seconds) | 300 |
ADV_VULN_ALERT_STATES |
Violation states to snapshot (comma-separated) | ACTIVE,ATTEMPTED |
ADV_PLATFORM_TIMEOUT |
Timeout per platform-forensics log call (seconds) | DIAG_TIMEOUT (30) |
analysis/acs-analyze reads an extracted must-gather and prints an automated
health report — the "is this cluster healthy?" checklist, scored. It runs
entirely offline against the extracted files and never contacts a cluster.
# point it at the extracted must-gather (root, or the image sub-directory)
analysis/acs-analyze path/to/must-gather.local.XXXX
# or via make
make analyze BUNDLE=path/to/must-gather.local.XXXXIt reports Central version / license, Central-DB availability, connected
clusters (collection method + version skew), distinct node kernel versions, and
— when the advanced layer is present — collection errors, Sensor↔Central
connectivity (including x509 CA-rotation breakage), TLS-cert expiry,
Sensor/Central heap and goroutine counts, Collector status, Scanner V4
readiness / restart churn, db-init permission/startup failures on any RHACS
database (central-db / scanner-db / scanner-v4-db), the
Sensor duplicate-IP warning count, admin events, and OOMKilled / restarting
pods. The heap check reads each component's pod memory
limit from its manifest and warns on the real percentage used (WARN ≥75%,
FAIL ≥90%). When the vuln-report/ layer is present it also summarizes
image-scan CVEs (flagging images with fixable critical / important findings) and
policy violations by severity; pass --image <name> to drill into the per-CVE
list for a specific image. Each check is OK / WARN / FAIL / SKIP; the
process exits non-zero if any check FAILs, so it is CI/scripting friendly.
Requires python3 (stdlib only). The heap checks use go tool pprof when go
is installed; without it they SKIP.
This is a step-by-step guide to reading the advanced-acs-diagnostics/ folder and
deciding whether a Secured Cluster is healthy. It assumes you have already
extracted a must-gather and are sitting in the image sub-directory (the one named
quay-io-...-acs-must-gather-sha256-...).
go— providesgo tool pprof, used to read the*.pb.gzprofiles. (Or install the standalonepprof:go install github.com/google/pprof@latest.)graphviz— only needed for the graphical/flame-graph views of pprof (brew install graphviz/dnf install graphviz). Text views work without it.jq— for the*.jsonfiles.
Set a shell variable to the folder so the examples below copy-paste cleanly:
ADV="advanced-acs-diagnostics"advanced-acs-diagnostics/
├── README.txt # what each collector is
├── secured-cluster-local/<ns>/ # Sensor/Collector/Admission, collected WITHOUT Central
├── tls-certs/ # certificate expiry report
├── crash-upgrade-forensics/ # crash + upgrade root-cause artifacts
├── scanner-v4/ # Central-independent Scanner V4 triage
├── vuln-report/ # image-CVE & policy-violation snapshot
└── platform/ # db-init logs, PVC/registry-CA, sorted events
First, sanity-check the run itself. A healthy collection has no .error files:
find "$ADV" -name '*.error' # expect: no outputA .error file is not fatal — each collector is best-effort — but it tells you
what was unreachable (often itself a finding). Note the two shapes: a single
sensor.error means no running Sensor pod was found at all, whereas per-endpoint
files like sensor-heap.pb.gz.error mean the pod is up but that debug/metrics
server did not answer.
A sensor-clusterentities-state.json.skipped file is normal — that store is
only served when Sensor runs with ROX_DEBUG_CLUSTER_ENTITIES_STORE=true.
What it stores: a sampled Go heap profile (a gzip-compressed protobuf) — a
snapshot of Sensor's live memory at collection time, broken down by the call
stack that allocated it. It holds four measurements per allocation site:
inuse_space/inuse_objects (memory live right now) and
alloc_space/alloc_objects (memory ever allocated, i.e. GC churn). It does
not contain object contents — no secrets or payloads.
Simplest examples (start here). The .pb.gz is just a gzipped protobuf — you
never unzip it yourself; go tool pprof reads it directly, fully offline (nothing
here touches the cluster). If you only remember three commands:
HEAP="$ADV/secured-cluster-local/stackrox/sensor-heap.pb.gz"
# 1. How much memory is live in total, and what is using the most? (top 10)
go tool pprof -inuse_space -top -nodecount=10 "$HEAP"
# 2. Show the same thing as a picture (needs graphviz)
go tool pprof -inuse_space -web "$HEAP"
# 3. Zoom in on one function by name (e.g. anything matching "Unmarshal")
go tool pprof -inuse_space -peek Unmarshal "$HEAP"(Command 3 uses -peek, not -list: -peek needs only the profile, while the
source-line -list view also needs the matching Sensor binary — see the note at
the end of this step.)
Reading command 1: the header line (... of 189.15MB total) is the live heap; the
first column (flat) is memory held by that function itself, cum is it plus
everything it calls. Biggest flat at the top = your top memory consumer. The
detailed steps below build on exactly these commands.
Step 1a — who is using memory right now:
go tool pprof -inuse_space -top -nodecount=15 "$ADV/secured-cluster-local/stackrox/sensor-heap.pb.gz"Example output from a healthy Sensor:
Type: inuse_space
Showing nodes accounting for 137.81MB, 72.86% of 189.15MB total
flat flat% sum% cum cum%
59.27MB 31.34% 31.34% 59.27MB 31.34% go.uber.org/zap/zapcore.newCounters
15.50MB 8.20% 39.53% 23.51MB 12.43% ...json.(*decodeState).objectInterface
9.50MB 5.02% 44.56% 25MB 13.22% storage.(*ImageScan).UnmarshalVT
9.50MB 5.02% 49.58% 15MB 7.93% storage.(*NetworkEntityInfo).UnmarshalVT
7.75MB 4.10% 53.68% 7.75MB 4.10% sensor/common/externalsrcs.(*handlerImpl).saveEntitiesNoLock
How to read it:
189 MB totallive is normal for a Sensor. Compare against the pod's memory limit (seecrash-upgrade-forensics/<ns>/describe-pods.txt): if live heap is near the limit, OOMKills are imminent.- The top consumers here are logging counters and protobuf/JSON unmarshaling of
ImageScan/NetworkEntityInfo— expected for a Sensor processing image scans and network flows. A single unexpected function holding a large, ever-growing share is the red flag.
Step 1b — interactive exploration (drill into a suspect function):
go tool pprof "$ADV/secured-cluster-local/stackrox/sensor-heap.pb.gz"
(pprof) top20 # biggest consumers
(pprof) tree saveEntities # callers/callees of a function
(pprof) peek Unmarshal # who calls anything matching "Unmarshal"
(pprof) web # opens a graph in the browser (needs graphviz)Step 1c — leak detection (the most valuable use). A single snapshot cannot prove a leak; a trend can. Collect two must-gathers a few minutes apart and diff them:
go tool pprof -inuse_space -diff_base OLD/sensor-heap.pb.gz NEW/sensor-heap.pb.gzAny node whose memory keeps growing across snapshots (especially in StackRox code) is a leak candidate. Flat or fluctuating = healthy churn.
Function names are embedded, so
top/treework out of the box. Source-line view (list <fn>) needs the matching Sensor binary — identify it by the Build ID printed at the top of the profile.
What it stores: a full text dump of every goroutine (Go's lightweight
threads) and its current stack, from /debug/goroutine?debug=2.
Step 2a — count goroutines (the single best liveness number):
grep -c '^goroutine ' "$ADV/secured-cluster-local/stackrox/sensor-goroutine.txt"A healthy Sensor is typically a few hundred to low thousands. Tens of thousands, or a number that keeps climbing across snapshots, means a goroutine leak — usually a blocked channel send/receive or an un-cancelled context.
Step 2b — find what they are stuck on (group by top-of-stack function):
grep -A1 '^goroutine ' "$ADV/secured-cluster-local/stackrox/sensor-goroutine.txt" \
| grep -v '^goroutine' | grep -v '^--' | sort | uniq -c | sort -rn | headStep 2c — look for long-blocked goroutines. The dump annotates how long a goroutine has been blocked; anything stuck for many minutes is suspicious:
grep -oE '[0-9]+ minutes' "$ADV/secured-cluster-local/stackrox/sensor-goroutine.txt" | sort -rn | headWhat they store: a one-shot scrape of each component's Prometheus /metrics
endpoint (plain text, # HELP / # TYPE / metric{labels} value).
Step 3a — is Sensor actually talking to Central? Look at Sensor's own counters and its gRPC client totals; these should be non-zero and, across two snapshots, increasing. Exact metric names vary by version, so grep broadly rather than for a fixed name:
grep -iE '^rox_sensor_|grpc' "$ADV/secured-cluster-local/stackrox/sensor-metrics.txt" | head -30Step 3b — Go runtime health (leak corroboration):
grep -E '^go_goroutines|^go_memstats_heap_inuse_bytes|^process_resident_memory_bytes' \
"$ADV/secured-cluster-local/stackrox/sensor-metrics.txt"go_goroutines should match Step 2a; process_resident_memory_bytes is the RSS
you compare against the pod limit.
Step 3c — error/drop counters (should be low and stable):
grep -iE 'error|dropped|failed|reject' "$ADV/secured-cluster-local/stackrox/sensor-metrics.txt" | grep -vE ' 0$'Anything here with a large or growing value is a live problem.
What they store: the collection driver each Collector pod is using and its
restart state (collector-status.txt), plus oc get pods -o wide for the
DaemonSet (collector-pods.txt).
cat "$ADV/secured-cluster-local/stackrox/collector-status.txt"Example (healthy):
=== collector-bpdfv ===
collectionMethod: CORE_BPF
restartCount: 0
lastState:
How to read it:
collectionMethod: CORE_BPF(orEBPF) on every node is what you want.restartCountclimbing, or alastStateofTerminated/OOMKilled, points to kernel-compatibility or memory problems on that specific node.- Confirm there is one Collector pod per node in
collector-pods.txt; a missing node means that node's runtime activity is not being observed.
What it stores: the connection/certificate-relevant lines grep'd out of the Sensor log, so you can judge the Sensor↔Central link without reading full logs.
Healthy looks like a clean connect sequence:
Info: Connecting to Central server central.stackrox.svc:443
Info: Established connection to Central.
Info: Communication with central started.
Unhealthy — watch for these and act accordingly:
| Line contains | Likely cause |
|---|---|
not trusted / invalid trust info signature |
cert/CA mismatch → check Step 6 |
different Central installation |
Sensor pointed at a re-installed Central → re-issue init bundle |
checking central status ... failed / repeated Connecting to Central |
network/DNS/Central-down |
offline mode |
Sensor lost Central and is buffering |
What it stores: subject/issuer/validity for every RHACS service certificate,
decoded from the TLS secrets (public cert material only — private keys are
never read). This is otherwise invisible because oc adm inspect redacts
secrets.
Find anything not healthy in one command:
grep -vE 'OK \(>30d' "$ADV/tls-certs/cert-expiry-summary.txt" | grep -E 'status:|==='status: OK (>30d remaining)on all entries → certs are fine.status: WARNING - expires within 30 days→ plan a rotation now.status: EXPIRED→ this is very likely your root cause; expired mTLS certs are the classic reason Sensor/Scanner "suddenly can't connect". Cross-reference with thenot trustedline in Step 5.
Step 7a — why did a container die? previous-logs/ holds the log of the
previous (crashed) instance of every container that has restarted — the single
most useful crash artifact, and one a normal must-gather does not isolate:
ls "$ADV/crash-upgrade-forensics/stackrox/previous-logs/"
tail -50 "$ADV/crash-upgrade-forensics/stackrox/previous-logs/<pod>-<container>.log"Step 7b — OOMKilled / scheduling failures: describe-pods.txt carries the
container Last State, exit codes, and events:
grep -E 'OOMKilled|Reason|Exit Code|FailedScheduling|Back-off' \
"$ADV/crash-upgrade-forensics/stackrox/describe-pods.txt"OOMKilled here + a high inuse_space in Step 1 = raise the memory limit or
find the leak.
Step 7c — RHACS's own view of problems: administration-events.json is what
Central surfaces to admins (scan failures, integration errors, expiring tokens):
jq -r '.events[] | "\(.level) \(.type) \(.hint // .message)"' \
"$ADV/crash-upgrade-forensics/administration-events.json"Step 7d — upgrade problems: if a Sensor upgrade is stuck, look at the
per-namespace sensor-upgrader.log / sensor-upgrader-deployment.txt /
sensor-upgrader-sa.yaml under crash-upgrade-forensics/<ns>/, plus the
cluster-scoped crash-upgrade-forensics/upgrade-sensors-rbac.yaml; their absence
simply means no upgrade was in progress.
| Check | Healthy | Where |
|---|---|---|
| No collection errors | no .error files |
find "$ADV" -name '*.error' |
| Sensor↔Central connected | Established connection to Central |
Step 5 |
| Certs valid | all OK (>30d) |
Step 6 |
| Sensor memory sane | live heap well under pod limit | Steps 1 + 7b |
| No goroutine leak | count stable, low thousands | Step 2 |
| No crash loops | restartCount low, no OOMKilled |
Steps 4 + 7b |
| Collector on every node | one CORE_BPF/EBPF pod per node |
Step 4 |
| No admin-visible errors | few/no ERROR events |
Step 7c |
If every row is green, the Secured Cluster is functioning well. The most common real failures this bundle exposes are, in order: expired certificates (Step 6), OOMKilled components (Steps 1 + 7b), and Sensor↔Central connectivity breaks (Step 5).
make build
make pushmake lintRequires shellcheck (and python3 for the
analyzer syntax check).
make testRuns the analyzer test suite (tests/test_analyze.py, stdlib unittest).
This project is licensed under the Apache License 2.0.