---
name: Datadog Agent
slug: datadog-agent-releases
type: github
source_url: https://github.com/DataDog/datadog-agent
changelog_url: https://github.com/DataDog/datadog-agent/blob/HEAD/CHANGELOG.rst
organization: Datadog
organization_slug: datadog
total_releases: 129
latest_version: 7.84.2
latest_date: 2026-10-07
last_updated: 2026-10-08
tracking_since: 2023-04-20
canonical: https://releases.sh/datadog/datadog-agent-releases
organization_url: https://releases.sh/datadog
---

<Release version="7.84.2" date="October 7, 2026" published="2026-10-07T13:53:20.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.84.2">
# Agent

### Prelude

Released on: 2026-10-07

-   Please refer to the [7.84.2 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7842) for the list of changes on the Core Checks

### New Features

-   Add Fleet-managed upgrades and rollback for Windows FIPS Agents.

### Enhancement Notes

-   Allow enabling Private Action Runner split mode during a Windows Agent installation with the `DD_PRIVATE_ACTION_RUNNER_SPLIT_ENABLED` MSI property.

### Security Notes

-   Bumped pip to 26.2.1 in the embedded Python distribution.
-   Patch the embedded MIT krb5 library so NegoEx GSS token parsing rejects truncated headers and missing extension vectors instead of reading past the buffer.

### Bug Fixes

-   Fix Kueue pods sometimes missing the `kueue_workload`, `kueue_workload_uid` and `kueue_resource_flavor` tags after an Agent restart.
-   Windows Fleet installer setup now rejects standard/FIPS Agent mismatches before stopping the installed Agent, with a message to use the matching installer.

# Datadog Cluster Agent

### Prelude

Released on: 2026-10-07 Pinned to datadog-agent v7.84.2: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7842).

### Bug Fixes

-   The Cluster Agent now collects namespace metadata when `apm_config.instrumentation.on_demand` is enabled, even if `apm_config.instrumentation.enabled` is false. Before this fix, Remote Configuration instrumentation policies that match on namespace labels never matched, and injected pods in namespaces that enforce the `restricted` Pod Security Standard were rejected because the init containers did not get a restricted security context.

</Release>

<Release version="7.84.1" date="October 2, 2026" published="2026-10-02T13:42:48.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.84.1">
# Agent

### Prelude

Released on: 2026-10-02

-   Please refer to the [7.84.1 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7841) for the list of changes on the Core Checks

### Security Notes

-   Update `golang.org/x/crypto` to v0.56.0.

# Datadog Cluster Agent

### Prelude

Released on: 2026-10-02 Pinned to datadog-agent v7.84.1: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7841).

</Release>

<Release version="7.84.0" date="September 30, 2026" published="2026-09-30T13:47:20.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.84.0">
# Agent

### Prelude

Released on: 2026-09-30

-   Please refer to the [7.84.0 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7840) for the list of changes on the Core Checks

### Upgrade Notes

-   On Linux, the Datadog process manager (`dd-procmgrd`) is now a first-class service manager alongside systemd, upstart and sysvinit, and it is the default when the `dd-procmgrd` binary ships with the Agent. The OpenTelemetry Collector distribution is supervised through `processes.d` and the `datadog-agent-ddot` systemd unit is no longer installed.

    Each service manager now installs its own self-consistent set of units, so a host runs either the `dd-procmgrd` supervision path or the per-payload systemd units, never a mix of the two.

-   Linux Agent packages no longer include the static archive `/opt/datadog-agent/embedded/lib/libpcap.a`. This affects users whose custom integrations or build scripts link directly against that archive. Before upgrading, check for references to this path and instead install the platform's libpcap development package or provide a separate libpcap build. The libpcap headers under `/opt/datadog-agent/embedded/include` remain available.

-   The `com.datadoghq.remoteaction.agent` Private Action Runner bundle (Preview) has been renamed to `com.datadoghq.remoteaction.datadogagent`.

-   The logs Agent now sends log payloads at a fixed higher concurrency by default instead of scaling the number of concurrent senders dynamically with intake latency (the previous "RTT fairness" behavior). This improves throughput for high-latency and high-volume workloads without manual tuning.

    The Agent now uses a static send concurrency of `logs_config.pipelines` x 10 (for example, 40 concurrent senders on a host with the default 4 pipelines).

    **Connection impact — please review before upgrading:** because the Agent now runs more senders concurrently by default, it may open more simultaneous connections to the logs intake (up to the static concurrency above). If you need to limit connections — for example behind a proxy or on a connection-constrained network — set `logs_config.batch_max_concurrent_send` to a lower value (`1` keeps a single sender per pipeline).

### New Features

-   Data security PostgreSQL scans can now connect over TLS. The instance option `ssl` accepts `disable`, `allow`, `prefer` and `require`. The `verify-ca` and `verify-full` modes are not supported yet, as server certificates are not validated.
-   Data security PostgreSQL scans now support the `verify-ca` and `verify-full` SSL modes, and send a client certificate and key on TLS connections when `ssl_cert` and `ssl_key` are set.
-   On Linux and macOS, the Agent now reports an Agent Health issue when the log file tailer loses log data because a file was rotated before the tailer finished reading it. The issue is reported once per host and covers every affected log source, reporting the total bytes lost and rotations involved over the last 24 hours together with a breakdown per source and service. Loss recorded before an Agent restart is not carried across the restart.
-   On macOS, the `thermal` check now reports hardware temperatures and the system thermal pressure level. It submits `system.thermal.temperature.cpu`, `system.thermal.temperature.gpu`, `system.thermal.temperature.ssd` and `system.thermal.temperature.battery` in degrees Celsius, along with `system.thermal.pressure_level` (0 nominal, 1 moderate, 2 heavy, 3 trapping, 4 sleeping) tagged with the level name. Both Apple Silicon and Intel Macs are supported. Sensors that the host does not expose are omitted rather than reported as zero, so the set of available metrics varies by hardware model.
-   OTLP ingestion: Adds a new `otlp_config.metrics.infra_attributes.as_tags` option. When enabled, custom tagger-derived tags (for example, tags configured via `kubernetesResourcesLabelsAsTags`/`kubernetesResourcesAnnotationsAsTags`) are promoted so they survive the metrics translator's allowlist and are emitted as metric tags for OTLP metrics ingested directly by the Agent. Without it, custom tags that are not known Datadog or OpenTelemetry conventions are dropped. Default behavior is unchanged.
-   DDOT: The `infraattributes` processor now supports a new `metrics_attributes_as_tags` option for the metrics pipeline. When enabled, custom tagger-derived tags (for example, tags configured via `kubernetesResourcesLabelsAsTags`/`kubernetesResourcesAnnotationsAsTags`) are promoted so they survive the metrics translator's allowlist and are emitted as metric tags. Default behavior is unchanged.
-   Rshell privileged helper process is introduced and allows specific, allow-listed commands to run with elevated privileges. This is enabled with private\_action\_runner: enabled: true restricted\_shell: privileged: enabled: true it requires `seccomp` and `landlock` and is linux-only.
-   Added the `device_tags_source` option to the SNMP check, controlling where the device tags on metrics come from. `resource` (default) sends only the device resource tag and lets the backend attach the device tags from the metadata payload. `agent` has the Agent attach the device tags to every metric and omits the resource tag, so no backend enrichment happens. `both` sends the device tags and the resource tag. Has no effect when `collect_device_metadata` is disabled.

### Enhancement Notes

-   The Datadog installer's setup script now accepts a `DD_LOG_LEVEL` environment variable to set the `log_level` option in `datadog.yaml` at install time. On Windows, the Agent MSI installer also exposes a `DD_LOG_LEVEL` property and passes it through to the Datadog installer.
-   Network Path Synthetic tests now report an NDM `namespace` on emitted paths. By default the Agent uses its configured `network_devices.namespace` (matching the `network_path` integration); individual tests can override it via a `namespace` field in their configuration. Previously synthetic paths were always emitted with an empty namespace.
-   Adds `data_plane.preflight_mode_duration`, which sets how long the Agent Data Plane (ADP) pre-flight leaves ADP running. It defaults to `90s` and is clamped to that as a minimum, so it can only extend the window: a shorter one stops ADP while its startup is still in progress and reports a healthy host as a failure. It is intended for benchmarking harnesses that need ADP resident for the whole of a fixed-length run; on a real host the setting to reach for remains `data_plane.preflight_mode`.
-   APM : The trace-agent EVP proxy now forwards the `DD-EVP-ORIGIN` and `DD-EVP-ORIGIN-VERSION` request headers to Event Platform intake. Previously these headers were stripped by the proxy allowlist, so server SDK metadata (SDK name and version) was lost at the Agent even when the SDK sent it. Values are forwarded unchanged; the Agent does not synthesize defaults for requests that omit them.
-   APM : Use the OpenTelemetry `url.template` attribute in HTTP client resource names when available, while retaining method-only resource names as the fallback.
-   APM: Add `apm_config.traces_send_to_main_endpoint` (`DD_APM_TRACES_SEND_TO_MAIN_ENDPOINT`, default `true`). This setting is for internal use, and we expect to remove it in a future version. When set to `false`, the trace-agent trace and stats writers stop sending to the main endpoint (`api_key` + `apm_config.apm_dd_url`) and forward traces and APM stats only to `apm_config.additional_endpoints`. The main endpoint's API key is still used by the other trace-agent proxies, and the configuration is rejected when no additional endpoint would remain. This is the traces/stats counterpart of `apm_config.profiling_send_to_main_endpoint`.
-   APM : Add support for installing APM injection on AMD64 systems that also run 32-bit binaries. 32-bit executables run without injection, while supported 64-bit executables continue to be instrumented.
-   Agents are now built with Go `1.26.6`.
-   Agents are now built with Go `1.26.7`.
-   The DDOT configuration converter now reuses a user-defined `pprof`, `zpages`, `health_check` or `ddflare` extension that was declared under `extensions` but not wired into `service.extensions`, instead of adding a default `<name>/dd-autoconfigured` copy. This matches the behavior already in place for the `datadog` and `dogtel` extensions, so a user's custom extension configuration is honored even when they forget to add it to the service's extension list.
-   The Datadog Distribution of the OpenTelemetry Collector (DDOT) now sends series metrics to Datadog using the v3 metrics intake, reducing metrics egress bandwidth. This aligns DDOT with the core Agent: series follow the `use_v3_api.series.enabled` setting, whose default (`datadog_only`) uses the v3 intake for Datadog destinations while other destinations continue to use the v2 intake. To keep using the v2 intake, set `use_v3_api.series.enabled` to `false` (or the `DD_USE_V3_API_SERIES_ENABLED` environment variable).
-   DogStatsD diagnostic commands (`dogstatsd-stats`, `dogstatsd-capture`, `dogstatsd-replay`, `dogstatsd top`, and `dogstatsd dump-contexts`) now print a clear error message and exit when invoked against the Core Agent while the Agent Data Plane is configured to handle DogStatsD traffic. In that mode the Core Agent's DogStatsD pipeline is dormant and the commands would silently produce empty or misleading results. Use the equivalent commands through the `agent-data-plane` binary instead.
-   Added retry-with-backoff to startup of the External Metrics Provider. The Cluster Agent now retries the external metrics server setup on transient APIServer failures instead of failing on the first attempt.
-   gpum: Add `gpu.legacy_sm_active` config toggle. When enabled on GPM-capable NVIDIA GPUs, `gpu.sm_active` reports the GPM SM utilization value while `gpu.sm_utilization` remains available.
-   gpu: add `gpu_nvlink_capable` and `gpu_nvlink_version` tags to GPU metrics, to allow filtering GPUs that have NVLink enabled and the version supported.
-   Health Platform: the `health_platform.issues_detected` telemetry counter is now also tagged with `severity`, in addition to the existing `issue_type` tag. This allows filtering or grouping detected health issues by their severity level.
-   Tag Kubernetes Job events emitted by the `kubernetes_apiserver` check with `kube_cronjob` when the Job's name matches the pattern generated by a CronJob, similar to the tagging already applied to Pod events.
-   `agent status` now reports why a file matched by a log configuration is not being tailed when its fingerprint cannot be used, under the log source that matched it. The most common cause is a file that holds less data than `logs_config.fingerprint_config.count`, for instance right after a log rotation replaced it with a smaller file, and the message names the setting to change. This previously showed up only as a shortfall in the `N files tailed out of M files matching` line, with no explanation.
-   The Logs Agent now logs a warning when a file matched by a log configuration is not tailed because its fingerprint cannot be used, which previously happened without any log line at any level. The most common cause is a file that holds less data than `logs_config.fingerprint_config.count`, for instance right after a log rotation replaced it with a smaller file. The warning reports the path, the current size of the file and how much data fingerprinting requires, and distinguishes a file that is simply too short from a file whose fingerprint could not be computed at all. It is emitted once when the file starts being skipped rather than once per check, and a closing message is logged when it stops being skipped, whether the file started being tailed again or was never tailed at all, so the collection gap can be measured from the Agent log alone.
-   The logon duration event on Windows now breaks each boot Group Policy pass down into the individual client-side extension invocations that ran during it, reporting the offset, duration, and outcome of each along with the Group Policy objects that fed it. The breakdown appears as a new `group_policy_details` block in the event's `custom` attributes, split into `computer` and `user` arrays.
-   Notable Events on macOS now reports the cause of the previous shutdown on Apple silicon: the Agent classifies the power management unit's boot-fault record after each boot and emits a `System shutdown fault` event when the previous shutdown was caused by a power fault, a processor crash signal, a watchdog timeout, a hardware fault or a thermal fault. The event reports the fault classification and the underlying fault tokens, and is emitted at most once per boot.
-   The Private Action Runner can now validate MongoDB connections
-   The Network Configuration Management (NCM) check now emits a `datadog.ncm.check_failure` metric, tagged with the failure reason, whenever the check encounters an error.
-   Network Config Management (NCM) local configuration store is now bounded by the network\_devices.config\_management.store.min\_configs\_per\_device, network\_devices.config\_management.store.max\_configs\_per\_device, and network\_devices.config\_management.store.max\_raw\_config\_store\_bytes configurations. Once these limits are exceeded, the least-recently-used configs are evicted. min\_configs\_per\_device and max\_configs\_per\_device is floored at 2. min\_configs\_per\_device and max\_configs\_per\_device are hard limits and will be enforced even if the size of the database file is greater than or less than max\_raw\_config\_store\_bytes. Default values are as follows: min\_configs\_per\_device = 3, max\_configs\_per\_device = 50, max\_raw\_config\_store\_bytes = 2000000000. Updates to store configurations become active upon agent restart.
-   The Datadog Distribution of OpenTelemetry (DDOT) Collector now compresses all telemetry signals with `zstd`, providing a consistent compression algorithm across metrics, logs, and traces. Metrics and logs default to `zstd` level 3 and the level is configurable through the `serializer_zstd_compressor_level` and `logs_config.zstd_compression_level` settings. Previously, metrics were compressed with `zlib` and traces with `gzip`.
-   OTLP ingest and DDOT: Logs received through the OTLP receiver now map the instrumentation scope name and version to `otel.scope.name` and `otel.scope.version`. Incoming `otel.library.name`/`otel.library.version` attributes (the deprecated OpenTelemetry predecessors) are remapped to the canonical `otel.scope.*` keys.
-   Private Action Runner: the `api_key_only_enrollment` setting now defaults to `true`. Runners now enroll using only an API key by default, without requiring an application key. Your API key need to have "Private Action Runner" scoped enabled.
-   Add opt-in PAR split mode support to the containerized Agent.
-   Enables PAR action execution on-demand to significantly reduce idle memory usage. This is opt-in behind `private_action_runner.split_enabled` and is currently available on Linux host deployments.
-   Private Action Runner: add the `private_action_runner.restricted_shell.disable_detailed_telemetry` setting. When set to `true`, it suppresses the raw command text and effective sandbox configuration that rshell attaches to its "run" telemetry span for each `runCommand`/`runRemediationCommand` invocation. The setting defaults to `false`; other rshell telemetry (exit code, timing, command counts) is unaffected.
-   Adds support for on-demand PAR action execution to Windows host deployments.
-   Data Observability: Extended the `queryactions` component to support SQL Server in addition to PostgreSQL. The component now schedules `data_observability.queries` for `sqlserver` check instances and correctly resolves Azure SQL Database instances by requiring both host and database equality, preventing cross-database payload injection.
-   Autodiscovery now reports the check configuration keys it ignores when several annotation or label formats are set on the same entity. Only the format with the highest priority is applied (`checks`, then `check_names` with `init_configs` and `instances`, then the legacy `service-discovery.datadoghq.com` prefix), and the others used to be discarded silently. The ignored keys are now listed in the Autodiscovery section of the `agent status` output.
-   The SSI `injection-metadata` telemetry payload now supports an optional, free-form `metadata` field for carrying `result_class`-dependent data.
-   Bumped the Security Agent policies to v0.84.0
-   APM: Reduce allocations when decoding v0.4 traces by interning strings directly from the incoming payload instead of allocating a string for every value already present in the string table.
-   APM: Reserve room for the trailing end-of-body read when buffering an incoming trace payload, avoiding a reallocation and copy of the buffer.
-   A small sample of sketch metric flushes (0.1% by default) is now additionally sent to a v3beta metrics intake endpoint to validate the upcoming v3 metrics protocol. Shadow traffic is only sent for agents configured against the `datadoghq.com` (US1) site. To opt out, set `serializer_experimental_use_v3_api.sketches.shadow_sample_rate` to `0`.
-   The Agent now logs a warning instead of an informational message when it finds a Docker, CRI or PodResources socket that exists but cannot be reached. On Unix this condition is only reported when opening the socket fails with a permission error, so it means the Agent user lacks access to the socket. Because these messages are emitted while the configuration is still loading, they were previously discarded on Agents running with `log_level` set to `warn` or above.

### Deprecation Notes

-   Remove system-probe module-restart CLI command.

### Security Notes

-   Data security PostgreSQL scans now open connections as read-only, as a first layer of protection against accidental writes.
-   The Datadog flare extension (`ddflare`) in DDOT no longer serves its endpoint, by default `https://localhost:7777`, without authentication.
-   Update agent-payload to v5.0.209 to remove the legacy zstd\_0 dependency.

### Bug Fixes

-   Autodiscovery: a negative index in the `%%port_<index>%%` template variable (for example `%%port_-1%%`) no longer crashes the Agent. Negative indexes are now resolved Python-style, `-1` being the last port, `-2` the second to last, and so on, wrapping around the list of ports so that indexes beyond the number of ports still resolve.
-   The file-based secret backend now rejects directory paths passed as a secret name. On AIX, reading a directory with `os.ReadFile` returns raw directory bytes instead of an error, which could cause a directory's contents to be returned as a secret value.
-   APM: Raise the messagepack decoder allocation limit for the span `meta_struct` field to 10MiB, so large `meta_struct` entries (as written by LLM Observability) are no longer rejected by the package-wide 500,000 element limit. Other span fields keep the lower limit.
-   APM peer-tag aggregation now refreshes derived tag keys when Remote Configuration changes semantic registry mappings without changing the producer-declared content hash.
-   Add the origin product source to distribution metrics originating from checks.
-   Clarify the Process Component status when Service Discovery is enabled without Live Process collection. The enabled checks now report `service_discovery` instead of `process` and `rtprocess` in this configuration.
-   CSM Misconfigurations: a Kubernetes configuration file the Agent may not read is now reported with its content left out. An unreadable kubelet kubeconfig used to parse into an empty one, matching a kubeconfig that really names no cluster, and the managed environment detection concluded from it that an EKS, GKE or AKS node was unmanaged. The detection now sees an absent kubeconfig, and the payload says as much. The ownership and permissions of an unreadable file are reported too, so the benchmarks that check them see the real mode.
-   CSM Misconfigurations: the Kubernetes node configuration collected for the CIS Kubernetes benchmarks now assumes the `KubeletConfiguration` defaults for a kubelet started with `--config`, and folds in the drop-ins of `--config-dir`. Kubernetes keeps the historical command line defaults of `--read-only-port`, `--anonymous-auth` and `--authorization-mode` for a kubelet started on flags alone, and the Agent applied them in both cases. A node whose configuration file left those settings out was therefore reported with a read-only port on `10255`, anonymous authentication enabled and an `AlwaysAllow` authorization mode, while the kubelet was really running with the read-only port disabled, anonymous authentication disabled and `Webhook` authorization. OpenShift and kubeadm both leave `readOnlyPort` out of their rendered configuration, so their nodes failed the "kubelet read-only port should be disabled" rule with the port closed.
-   CWS: fix a regression where the process context updates carried by an event (`setuid`, `setgid`, `capset`, login UID and IMDS security credentials) were applied before the event was evaluated, preventing rules from matching on the process state that preceded the event.
-   Data Observability query actions now match PostgreSQL checks by their resolved database identifier. Queries now run when the identifier uses an agent hostname override or a `database_identifier` template.
-   Fixed DDOT being unable to reach the core Agent in containerized deployments that do not pass `--core-config` (Docker, ECS, and ECS Fargate). Without a core config path the connection failed with `x509: certificate signed by unknown authority`. The OTel Agent now falls back to the default `datadog.yaml` location when it exists and the collector is not running in standalone mode.
-   Fixed a bug in the trace-agent DogStatsD proxy endpoints (`/dogstatsd/v1/proxy` and `/dogstatsd/v2/proxy`) where a request body was split into all of its newline-separated payloads at once, so a body containing many newlines used several times its own size in memory. The body is now scanned one payload at a time, empty payloads are skipped, and the number of payloads relayed per request is capped.
-   The Agent no longer shows the content of check config files on its local expvar page. This includes the configuration of JMX checks. The content is still sent to Datadog and still included in the flare.
-   The metric filter list (`metric_filterlist`) now matches on the normalized metric name instead of the raw submitted name. Metric names are normalized by the Datadog intake on ingest, so a metric submitted as `my metric-name` is stored and displayed as `my_metric_name`. Previously a filter list entry using the normalized name that users see in Datadog would fail to match such a metric, and the metric was submitted anyway. Filter list entries themselves are matched as written, so they should be the metric name as it appears in Datadog.
-   Kubernetes orchestrator: Prevent unchanged clusters from being repeatedly reported as updated when the node listing order changes.
-   CWS: Network events are now attributed to the correct process when the same address and port are reused across different network namespaces.
-   Fix the default `query_timeout` for the Oracle check from 20,000 seconds to 20 seconds.
-   Fixed a bug in container image reporting where image references whose registry host included a port (for example `myregistry.local:5000/foo/bar:1.2.3`) were split on the first colon instead of the tag separator. This caused the `image_name` and `image_tag` tags, as well as the values shown in the Container Images view, to be incorrect for such images. The repo and tag are now parsed using the last colon following the last slash, matching standard image reference rules. Relatedly, the registry of images whose reference has a registry host followed by a single path component (for example `localhost:5000/service` or `registry.k8s.io/pause`) is now reported in the Container Images view instead of being left empty.
-   Fixes Agent installs and upgrades failing to create any systemd unit files on hosts whose kernel does not support ambient capabilities (kernel older than 4.3). On those hosts the installer selects the `-nocap` systemd unit templates, but those templates were never compiled into the installer binary, so unit generation failed with `failed to write stable units: open tmpl/gen/debrpm-nocap/datadog-agent.service: file does not exist` and the host was left with no Datadog units at all. Because the package manager scriptlet ignores this failure, the install appeared to succeed while `systemctl start datadog-agent` reported `Unit not found`. This affected both the classic DEB/RPM install path and Fleet Automation remote upgrades and configuration experiments, which use the equivalent `oci-nocap` templates.
-   Fix container log corruption when partial records from stdout and stderr are interleaved. The Agent now reconstructs partial CRI and Docker JSON-file records independently for each stream.
-   Cluster Agent (KSM check): fix a collision when collecting custom resource metrics for two custom resources that share the same `Kind` and plural name but belong to different API groups (for example `Project` in both `artifactory.example.com` and `sonarqube.example.com`). Previously the resources shared a single API client, causing repeated `Unexpected watch event object gvk` errors and mixed, incorrect metric counts. Custom resource clients are now keyed by their fully-qualified GroupVersionResource.
-   NCM check frequency will no longer default to a negative number; config values expressed as integers instead of durations (e.g. "5" instead of "30s" or "10m") will be parsed as seconds instead of nanoseconds.
-   DDOT: fix `DD_SITE` being ignored when `api.site` is absent from the Datadog exporter config. Telemetry was silently sent to `api.datadoghq.com` instead of the site configured via `DD_SITE` or `datadog.site`.
-   Fix an issue on Windows where uninstalling the Agent could fail if the configuration directory (`C:\ProgramData\Datadog` by default) had already been removed before running the uninstall.
-   On Windows, the `wlan` check no longer panics on hosts where `wlanapi.dll` is unavailable. The library ships with the `Wireless-Networking` feature, which is not installed by default on Windows Server. Such hosts are now reported as having no active Wi-Fi interface.
-   Fix an issue where the Agent logged a spurious error about Workloadmeta collectors not being ready on every startup in environments where no container runtime or orchestrator is detected, such as container sidecars.
-   Flare archive filenames now include a process ID and counter suffix (`datadog-agent-<timestamp>-<pid>-<counter>-<loglevel>.zip`), preventing two archives created by the same Agent process at the same second from overwriting each other.
-   Fleet Installer: Fix an issue where the installer's telemetry client repeatedly logged `failed to send telemetry payload` warnings with `404 Not Found` errors on GovCloud sites (`ddog-gov.com`), since no instrumentation telemetry intake exists there. Telemetry is now disabled outright for GovCloud sites, matching the existing behavior of the Agent's own resident telemetry.
-   HA Agent: a Remote Config document on the `HA_AGENT` product that belongs to the `comp/workloadbalancing` component (identified by an explicit `type` field) is now skipped instead of being processed as an invalid HA Agent document. If every document in an update batch turns out to belong to `comp/workloadbalancing`, HA Agent resets its state to `Unknown` rather than keeping a stale Active/Standby state.
-   Fix Private Action Runner startup delays when signing keys are already available from Remote Config.
-   Fix an Agent IPC client socket leak when a local service accepts a TLS connection but never completes the handshake.
-   KSM core check: Fixed a bug where the `kubernetes_state.configmap.count` metric could stop being reported. ConfigMaps are collected using a metadata-only Kubernetes client, and the conversion of watch events into ConfigMap objects was dropping annotations, including the `k8s.io/initial-events-end` bookmark annotation used by newer versions of client-go's watch-list feature to signal that the initial list has finished syncing. Without that signal, the underlying reflector would wait indefinitely and the metric would never be reported.
-   Windows: Fix the logon duration event reporting `boot_timeline` and `group_policy_details` entries out of chronological order. A milestone that ran while the machine sat at the login screen, such as Computer Group Policy on a domain-joined host, could render after milestones that actually followed it.
-   Add new `bind_host` configuration option to TCP and UDP log listeners.
-   Normalize Windows device tags to use forward-slash paths, preventing duplicate disk and IO metric device tags.
-   Fixed a standalone DDOT (`DD_OTEL_STANDALONE=true`) failing to start when configured without a Datadog exporter, for example when the standalone DDOT is used to forward telemetry to a separate gateway layer via an OTLP exporter instead.
-   Agent OTLP Ingest no longer attaches the OpenTelemetry `debugexporter` by default. The `debugexporter` is now attached only when the `otlp_config.debug` section is explicitly configured. Declaring the section without a verbosity uses the default verbosity (`basic`), while setting `otlp_config.debug.verbosity` (or `DD_OTLP_CONFIG_DEBUG_VERBOSITY`) to `basic`, `normal`, or `detailed` selects the verbosity. Setting it to `none` leaves the exporter detached.
-   The deprecated `process_config.enabled` setting no longer overrides `process_config.container_collection.enabled` and `process_config.process_collection.enabled` when those are configured directly, whether through the configuration file, an environment variable, or any higher-precedence source. Previously, setting `DD_PROCESS_CONFIG_ENABLED=false` together with `DD_PROCESS_CONFIG_CONTAINER_COLLECTION_ENABLED=false` left container collection enabled. `process_config.enabled` still applies to whichever of the two settings is left unset, so configurations that only use the deprecated setting are unaffected.
-   Autodiscovery now retries check configurations that failed secret resolution when `secret_refresh_interval` is enabled, allowing checks to recover after a transient secret backend outage without restarting the Agent.
-   Fix a crash of the Agent process when a Python check raises an exception whose message contains a Unicode lone surrogate.
-   Runtime usage enrichment keeps container image SBOM components that share a name and a version. Trivy reports one component per install location, so a library version pinned by two lockfiles appears twice, and the merge kept only the first, leaving the dependency graph referencing a component the payload had lost. Components are deduplicated by their CycloneDX bom-ref, which is what identifies one.
-   Fixed container image SBOMs being reported as in use after their last container had stopped. On Kubernetes nodes, an image that had run a container once kept that flag until the Agent was restarted.
-   The runtime usage properties merged onto container image SBOMs (`LastSeenRunning`, `HasSetSuidBit` and `RunningAsRoot`) now reach OS packages alone. The system-probe report is matched to components by name and version, so a language package sharing an OS package's name and version took the timestamp and flags of the OS package that had run. Components whose purl places them outside the dpkg, rpm and apk databases are left out of the match.
-   The runtime usage properties merged onto container image SBOMs (`LastSeenRunning`, `HasSetSuidBit` and `RunningAsRoot`) are now set on OS packages alone. The runtime scanner reads the dpkg, rpm and apk databases, so language packages (npm, pypi, golang and others), the image's operating-system component and the per-lockfile application components stay out of its scope. The absence of a property marks a component out of scope, where a `LastSeenRunning` of `0` states that the package was watched and found idle.
-   Fixed `sbom.container_image.use_spread_refresher` never refreshing container image SBOMs on hosts with fewer than ten images, and refreshing them more slowly than `periodic_refresh_seconds` on other hosts.
-   Fixed the Agent crashing when the stored SBOM of a container image could not be uncompressed. The image is now skipped instead.
-   ECS Fargate: skip EC2 IMDS instance-type lookups. The Agent no longer periodically queries `169.254.169.254/latest/meta-data/instance-type` on Fargate (where IMDS is unavailable), which removes recurring INFO log noise in CloudWatch.
-   SNMP: Fix the detection of profiles using the legacy Python metric syntax. The detection result is now cached along with the profiles, so every check instance is consistently handed over to the Python loader instead of only the first instance to be configured. Previously the remaining instances silently kept using the Core loader.
-   SNMP: The error reported when a legacy profile forces the fallback to the Python loader now names the profiles using the legacy syntax.
-   Stop the Agent from attempting to reach a Kubelet when running as a Cluster Checks Runner. Cluster Checks Runners are Deployment replicas, not DaemonSets, so they never have a locally-reachable Kubelet, since a CCR doesn't correspond to any one node. This removes the recurring `Impossible to reach Kubelet through HTTPS` warning logged by Cluster Checks Runner pods. There is no change in behavior on the node Agent or Cluster Agent.
-   Auto multi-line detection no longer concatenates consecutive IIS W3C extended-format access log entries. A client IPv4 address next to the leading timestamp was being scored as part of that timestamp, which dropped the match to the detection threshold so each single-line record was treated as a continuation of the previous one. IPv4 addresses are now recognized as their own token and no longer interfere with timestamp detection.
-   Fixed a crash in the trace-agent when processing a v0.7 payload whose trace chunk omits the `tags` field and whose spans carry a `_dd.p.dm` tag. Promoting the decision maker to the chunk level no longer writes to an uninitialized map.
-   APM: Fix several trace-agent debug/error log messages that printed incorrect info due to format-string bugs.

### Other Notes

-   Agent Data Plane has been bumped to version 1.6.0. See the [Agent Data Plane 1.6.0 release notes](https://github.com/DataDog/saluki/releases/tag/1.6.0).
-   Agent Data Plane has been bumped to version 1.6.1. See the [Agent Data Plane 1.6.1 release notes](https://github.com/DataDog/saluki/releases/tag/1.6.1).
-   Added origin mapping for the Cisco Catalyst Center integration.
-   Extended the delegated authentication (Workload Identity Federation) component so it can write a resolved API key into a map- or list-shaped `additional_endpoints` configuration value, replacing a placeholder `DELA(...)` directive. This is internal foundation work; no Agent subsystem reads `DELA(...)` directives yet, so this does not enable any user-facing behavior on its own.
-   Added origin mapping for the Kueue and External Secrets integrations.
-   The `battery` and `wlan` check configurations are no longer shipped in the Linux and AIX packages. Both checks can only collect on macOS and Windows, so on other platforms their configuration only mattered when `infrastructure_mode` was set to `end_user_device`, where it scheduled a check that could never report data. macOS and Windows packages are unchanged.
-   Add metric origin mapping for the thermal integration.

# Datadog Cluster Agent

### Prelude

Released on: 2026-09-30 Pinned to datadog-agent v7.84.0: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7840).

### Upgrade Notes

-   `DD_INSTRUMENTATION_INSTALL_TYPE=k8s_single_step` and SSI defaults (for example `DD_TRACE_ENABLED`) apply only when a target or policy matches the pod, not merely because the namespace could match some rule. Pods that only have library annotations and do not match a target or policy use `k8s_lib_injection`.

### New Features

-   `DatadogInstrumentation` checks and logs configurations can now target Strimzi `StrimziPodSet` workloads.
-   Add AppSec injection support for GKE managed Gateways (EXTERNAL mode only). The Cluster Agent now detects GKE managed Gateways whose `spec.gatewayClassName` is in an allowlist of external-managed GatewayClass names (`gke-l7-global-external-managed`, `gke-l7-regional-external-managed`) and creates one `GCPTrafficExtension` (`networking.gke.io/v1`) per Gateway to route edge traffic through a user-deployed Datadog AppSec callout service. SIDECAR mode is not supported because managed GKE has no in-cluster Envoy data plane. The callout Deployment, Service, and HealthCheckPolicy must be deployed by the user following the public GKE service-extensions documentation. Cluster-agent RBAC for `gcptrafficextensions.networking.gke.io` (get/list/watch/create/delete) is required. The GatewayClass allowlist is configurable via `appsec.proxy.gke.gateway_classes`. Multi-cluster GatewayClasses (names ending in `-mc`) are always skipped, including when added to that allowlist, because they require a `net.gke.io` `ServiceImport` callout backend that the Cluster Agent does not create. Each `GCPTrafficExtension` is owned by its Gateway via an owner reference, so Kubernetes garbage-collects it even if the Cluster Agent misses the Gateway deletion event.

### Enhancement Notes

-   Adds Cluster Agent telemetry for the `DatadogInstrumentation` controller, including the number of resources it tracks and reconciliation outcomes for checks and logs.
-   Added the `datadog.cluster_agent.autoscaling.workload.objective.target` gauge, which exposes the target value configured in a `DatadogPodAutoscaler` `spec.objectives`. The metric is tagged with `objective_type` (`pod_resource`, `container_resource` or `custom_query`), `value_type` (`utilization` or `absolute_value`), `objective_index` (the 0-based position in `spec.objectives` that keeps each objective a distinct timeseries), and, for resource objectives, `resource_name` and `kube_container_name`.
-   Collect NVIDIA Dynamo custom resources by default.
-   Collect KubeRay `RayCluster`, `RayCronJob`, `RayJob`, and `RayService` custom resources by default.
-   Single Step Instrumentation now evaluates local targeting (Helm, Operator, or `datadog.yaml`) before Remote Config policies. Local targets keep first-match-wins order. Remote Config policies use last-match-wins: the last matching policy applies, so a catch-all can be listed first and exceptions after. A workload that matches a local target is not overridden by a remote deny.
-   Single Step Instrumentation now decides SSI versus local library injection from whether a configuration target or remote-config policy matched the pod, instead of a namespace-level eligibility approximation. Library annotations still short-circuit target selection for library versions and tracer configs (existing GA precedence).

### Security Notes

-   The Cluster Agent's admission controller webhook now enforces a size limit on incoming request bodies and validates the request content type before reading the body.
-   The Cluster Agent no longer exposes the Go `pprof` profiling and `expvar` debug endpoints on its metrics port (`metrics_port`, default `5000`) to remote callers. These `/debug/` endpoints were previously served on all network interfaces without authentication; they are now restricted to loopback callers, and requests originating from any other address receive a `404`. The `/metrics` endpoint is unchanged and remains reachable off-host so the node Agent can continue to scrape Cluster Agent telemetry. Local tooling such as the Cluster Agent flare, which connects over loopback, is unaffected.

### Bug Fixes

-   Fixed an issue where the `DatadogPodAutoscaler` controller could silently drop a status update after an HTTP 409 (Conflict) caused by a stale `resourceVersion` read from the informer cache. The reconcile now requeues on such a conflict so a subsequent pass retries the update with a refreshed object.
-   Ensure Kubernetes endpoint check annotations take precedence over `DatadogInstrumentation` configurations that target the same Service and integration. This prevents duplicate endpoint checks and restores the CR-backed check when the overriding annotation is removed.
-   Fix an issue on AKS clusters where the Cluster Agent and the AKS admission enforcer would repeatedly overwrite each other's changes to the `datadog-webhook` `MutatingWebhookConfiguration`/`ValidatingWebhookConfiguration` objects, causing `the object has been modified` errors to be logged in a loop. This affected admission controller features whose webhook rule did not otherwise restrict which namespaces it applies to (for example the `DatadogInstrumentation` CRD validating webhook), even when `admission_controller.add_aks_selectors` (`DD_ADMISSION_CONTROLLER_ADD_AKS_SELECTORS`) was enabled.
-   Fix a crash in the `kubernetes_state_core` check (Cluster Agent or Cluster Check Runner) that occurred when using a wildcard entry (`"*"`) in `kubernetes_namespace_annotations_as_tags` or the equivalent `kubernetes_resources_annotations_as_tags` namespace configuration, on any namespace without annotations.
-   Fixed an issue where the Cluster Agent could schedule Prometheus Scrape OpenMetrics checks against Kubernetes services even when the autodiscovery configuration included `kubernetes_container_names`. The Cluster Agent now skips service and endpoint Prometheus Scrape scheduling for configurations that set `kubernetes_container_names`; scraping discovered by node Agents remains as is.
-   Fix a Cluster Agent API bug that could log spurious `superfluous response.WriteHeader call` warnings. The internal telemetry wrapper now correctly tracks the response status when a handler writes the response body before explicitly setting the status code.
-   Restore continuous `kubernetes_state.endpoint.address_available` and `kubernetes_state.endpoint.address_not_ready` reporting (including `0` for the opposite ready state). After the kube-state-metrics v2.18 bump, each metric was only emitted for addresses in that state, so healthy endpoints no longer reported `address_not_ready=0` and fully unready endpoints no longer reported `address_available=0`.

</Release>

<Release version="7.83.3" date="September 24, 2026" published="2026-09-24T09:46:35.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.83.3">
# Agent

### Prelude

Released on: 2026-09-24

- Please refer to the [7.83.3 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7833) for the list of changes on the Core Checks

### Bug Fixes

- Fix missing log source configuration fields in public inventory metadata, including Windows Event Log queries, processing options, auto-multiline settings, and maximum message size.

# Datadog Cluster Agent

### Prelude

Released on: 2026-09-24 Pinned to datadog-agent v7.83.3: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7833).

</Release>

<Release version="7.83.2" date="September 16, 2026" published="2026-09-16T15:31:13.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.83.2">
# Agent

### Prelude

Released on: 2026-09-16

- Please refer to the [7.83.2 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7832) for the list of changes on the Core Checks

### Enhancement Notes

- gpu: Constant metrics (`gpu.device.total`, `gpu.memory.limit`, `gpu.core.limit`, and `gpu.memory.bar1.total`) are reported on the cadence set by the new `gpu.static_metrics_reporting_interval` (15 seconds by default) to ensure accuracy when using weighted sums.
- Single Step Instrumentation's tracer config mechanisms (the `ddTraceConfigs` `Targets` field, remote-config policies, and the `admission.datadoghq.com/apm-inject.tracer-configs` pod annotation) now also accept `OTEL_` prefixed environment variable names, in addition to the existing `DD_` prefix. This allows configuring a tracer's native OpenTelemetry mode (e.g. `OTEL_TRACES_EXPORTER`, `OTEL_EXPORTER_OTLP_ENDPOINT`) through SSI.

### Bug Fixes

- Fixed the Agent Data Plane pre-flight sending its metrics and API key to `datadoghq.com` instead of the configured `site`. The generated pre-flight configuration was built from the fully resolved Agent configuration, so `dd_url`'s default value appeared in it as though it had been set explicitly, and an explicit `dd_url` takes precedence over `site`. The pre-flight configuration is now built from the settings the operator actually supplied, matching what a normally-supervised Agent Data Plane reads from `datadog.yaml`.
- Fix `DD_NETWORK_PATH_COLLECTOR_FILTERS` so JSON-encoded Network Path collector filters are parsed and applied correctly.

# Datadog Cluster Agent

### Prelude

Released on: 2026-09-16 Pinned to datadog-agent v7.83.2: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7832).

</Release>

<Release version="7.83.1" date="September 9, 2026" published="2026-09-09T12:28:46.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.83.1">
# Agent

### Prelude

Released on: 2026-09-09

- Please refer to the [7.83.1 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7831) for the list of changes on the Core Checks

### Bug Fixes

- Fix bug which made fast network path test billed to customer
- Release the containerd view snapshot and lease taken for a container image SBOM scan even when the scan is cancelled or times out. The release ran on the scan's own context, so a scan that hit its deadline left the snapshot behind, and on a lazy snapshotter that snapshot holds the layer it materialised.

### Other Notes

- The fleet installer daemon now reports the DDOT (OpenTelemetry Collector) process state as part of the agent state sent to Datadog, so DDOT version and configuration updates can be monitored. The state is read from the process manager when it supervises DDOT, and from systemd or the Windows service manager otherwise. It is also visible in the output of `datadog-installer status`.

# Datadog Cluster Agent

### Prelude

Released on: 2026-09-09 Pinned to datadog-agent v7.83.1: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7831).

### Bug Fixes

- Fix Cluster Agent graceful shutdown to release the Kubernetes leader-election lock before exiting, allowing another replica to take over without waiting for the lease to expire.

</Release>

<Release version="7.83.0" date="September 3, 2026" published="2026-09-03T13:18:20.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.83.0">
# Agent

### Prelude

Released on: 2026-09-03

- Please refer to the [7.83.0 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7830) for the list of changes on the Core Checks

### New Features

- Add a Data Security provider that schedules one-off database scan checks triggered through Remote Configuration. It is enabled when both `data_security.enabled` and `shared_library_check.enabled` are set, and currently targets PostgreSQL databases already monitored by the Agent (via the `postgres` check).

- Add `k8s cluster receiver`, `k8s leader elector extension`, and `count connector` to the DDOT (Datadog Distribution of OpenTelemetry Collector) default manifest, enabling collection of Kubernetes cluster-level metrics, leader election coordination for Kubernetes receivers, and count-based metric generation via the OpenTelemetry Collector pipeline.

- Add Helm rollback action

- Adds an Agent Data Plane (ADP) preflight mode, controlled by the new `data_plane.preflight_mode` setting (enabled by default).

  When `data_plane.enabled` has not been set at all, the Agent starts ADP once at startup for 90 seconds in an isolated configuration, sends a single throwaway metric through it, then stops it and reports any startup errors to Datadog as agent telemetry. This surfaces environment-specific ADP problems before ADP is enabled for real.

  The preflight process handles no customer data: it runs in standalone mode, listens only on a temporary DogStatsD endpoint under the Agent's run directory, does not register with the Agent's remote agent registry, and never takes over the Agent's own DogStatsD port. The temporary configuration it is given contains the Agent's resolved configuration, so it is written user-only and removed once the run finishes. To keep values that were only ever held in memory from being written out this way, the pre-flight does not run at all when secrets are in use: when `secret_backend_command`, `secret_backend_type` or `multi_secret_backends` is set, or when any setting has already been resolved from a secret. Setting `data_plane.enabled` explicitly to either `true` or `false`, or setting `data_plane.preflight_mode` to `false`, also disables the pre-flight, as does running an Agent package that does not ship ADP.

- APM : Add support for span-derived primary tags on span metrics produced by the Datadog Distribution of OpenTelemetry Collector (DDOT). Set `span_derived_primary_tags` on the `datadog` connector's `traces` section to a list of attribute keys, and the value of each key found on a span (or, failing that, on its resource) is attached to the APM stats the connector emits, letting you break down span metrics by those tags:

      connectors:
        datadog/connector:
          traces:
            span_derived_primary_tags: [team, region]

  Keys absent from both the span and its resource attributes are omitted. Each key must also be configured as a primary tag in your Datadog organization; keys that are not registered as primary tags are dropped by the intake and do not appear on the resulting span metrics.

- CWS `set` actions accept a new `capture` field, a regular expression with a single capture group that is applied to the value of `field` to extract part of it. This makes it possible to lift an identifier embedded in an event field, such as a command id inside a file path or an IAM role inside an IMDS url, and store it as a scoped variable rather than storing the whole field value. `capture` can only be used together with `field`, and only on fields holding a single string. A value that does not match the expression leaves the variable untouched.

- When `data_security.enabled` is set, the Agent now forwards sensitive-data-scanner findings to Datadog as structured `sds-result` payloads on the event platform.

- Scaffold PostgreSQL support to the data security check.

- DDOT (Datadog Distribution of OpenTelemetry Collector): the embedded `datadog` exporter now honors the `orchestrator_explorer` setting when the OpenTelemetry Agent runs in standalone mode (`DD_OTEL_STANDALONE=true`). When enabled, Kubernetes resource manifests collected by a `k8sobjects` receiver in the exporter's logs pipeline are forwarded to the Orchestrator Explorer (Kubernetes Resources) intake. In connected mode this setting is ignored, as the Datadog Cluster Agent already collects and ships orchestrator data.

- Add support for additional log collection options for Kubernetes workloads through DatadogInstrumentation resources.

- Adds support for configuring Network Path Dynamic Test filters through Remote Config.

- Adds support for scheduling Network Path tests through Remote Config.

- Added ConfigMap collection to the Kubernetes Orchestrator. ConfigMap manifests are sent with their `data` and `binaryData` fields stripped. The collector is disabled by default (`IsStable: false`) and must be activated explicitly by listing `configmaps` in the `collectors` field of the orchestrator check instance configuration.

- The DDOT (OpenTelemetry) config converter now automatically injects the `cumulativetodelta` processor into metrics pipelines that export to the `datadog` exporter, converting all cumulative metric types (sum, histogram and exponential histogram) to delta. The processor is added only to metrics pipelines, and is skipped for any pipeline where a `cumulativetodelta` processor is already defined. This behavior is controlled by the new `cumulativetodelta` entry in `otelcollector.converter.features`, which is enabled by default; remove it from that list to disable the auto-injection.

- OTLP ingestion: Adds a new `otlp_config.logs.infra_attributes.tags_as_ddtags` option. When enabled, custom tagger-derived tags (for example, tags configured via `kubernetesResourcesLabelsAsTags`/`kubernetesResourcesAnnotationsAsTags`) are written as real Datadog log tags instead of log attributes for OTLP logs ingested directly by the Agent. Default behavior is unchanged. (commit 522ee5fe634)

- DDOT: The `infraattributes` processor now supports a new `logs_tags_as_ddtags` option. When enabled, custom tagger-derived tags (for example, tags configured via `kubernetesResourcesLabelsAsTags`/`kubernetesResourcesAnnotationsAsTags`) are written as real Datadog log tags instead of log attributes. Default behavior is unchanged. (commit 522ee5fe634)

- Added a `com.datadoghq.remoteaction.agent` Private Action Runner bundle exposing read-only datadog-agent operations (status, diagnose, and configuration) as remote actions executed against the local Agent's authenticated IPC API. (Preview)

- Added a `generateFlare` action to the `com.datadoghq.remoteaction.agent` Private Action Runner bundle that builds a flare archive on the Agent host. (Preview)

- The Private Action Runner now supports a split deployment model in which a dedicated on-demand executor runs actions in a separate process, reachable over a local gRPC socket secured with mutual TLS.

- Added the `data_security.enabled` configuration flag (disabled by default) which enables the `sds-result` event platform forwarder used to send sensitive-data-scanner results to Datadog.

- Add the Data Security feature to scan monitored PostgreSQL databases for sensitive data, driven remotely through Remote Configuration (`DATA_SECURITY_DB_SCAN_TASKS`). Enabled with `data_security.enabled` and `shared_library_check.enabled`.

- Agent Cloud Auth (delegated authentication / Workload Identity Federation) on AWS now resolves credentials from EKS IRSA, ECS task roles, EKS Pod Identity and EC2 IMDS in the trace-agent, standalone DogStatsD, private action runner, IoT Agent and Heroku Agent. Previously only flavors built with the `ec2` build tag (main Agent, Cluster Agent, process-agent, security-agent, system-probe, installer) supported those credential sources; the others silently disabled the feature. Most notably the trace-agent is now covered, so APM no longer requires a statically configured `api_key` when Cloud Auth is in use. OpenTelemetry Collector (DDOT / `otel-agent`) is not covered: it does not load the delegated authentication component, and still requires a statically configured `api_key`.

- Setting `delegated_auth.aws.region` without `delegated_auth.provider` no longer skips cloud provider auto-detection. The configured region is now applied to the auto-detected provider, as intended, instead of being treated as an explicit provider configuration.

### Enhancement Notes

- Notable Events on Windows now reports critical temperature events: system shutdown or hibernation triggered by a critical thermal condition.

- The Linux Agent packages now ship a built-in `datasecurity` Rust-based check as a shared library under `/etc/datadog-agent/checks.d`. This is an initial scaffold and is not enabled by default.

- Rust artifacts are now built with `codegen-units = 1` to minimize the size of the produced binaries and shared libraries.

- The kubelet check now collects `kubelet.containers_per_pod` (`.count` and `.sum`), a histogram of the number of containers running per pod on a node. This provides visibility into the distribution of container counts across pods, which can help identify pods with unusually high sidecar/container density.

- Adds kubernetes-actions functionality to the private action runner in the DCA.

- Adds ability to patch daemonsets and statefulsets through the kubernetes-actions pipeline.

- Add `exporter.datadogexporter.AddUnits` feature gate that maps OTLP (UCUM) metric units to their Datadog equivalents.

- APM : Supported new `db.system.name` attribute replacing `db.system` according to changes in OpenTelemetry Semantic Conventions (v1.30.0+).

- APM: `agent status` now shows the live trace-semantics registry in the APM Agent section, reporting whether it comes from Remote Configuration or the embedded default along with its content hash and version.

- APM: Reduced memory allocations when decoding incoming v0.4 and v0.5 trace payloads with the `convert-traces` feature enabled. Span attributes are now batch-allocated while converting to the internal trace format, lowering allocation counts and garbage-collection pressure in the trace-agent receiver. This has no effect when `convert-traces` is disabled.

- The Rust shared-library checks are now built with Bazel, which enables running their unit tests and clippy lint checks in CI as part of the build.

- Agents are now built with Go `1.26.7`.

- Check instances scheduled via configuration discovery now carry the `dd_config_discovery:true` tag. This can be used to identify, and if needed exclude, metrics submitted by an autodiscovered check that duplicates a check configured manually elsewhere for the same service.

- Data Security scan results now report the total number of sensitive-data matches found in each column (`count_matches`), in addition to the number of distinct rows that contain a match. This gives more accurate visibility when a single row contains several matches.

- The Data Security check now resolves its PostgreSQL source from fully templated autodiscovery configurations, so scan targets are matched reliably when the PostgreSQL integration is configured with template variables such as `%%host%%`.

- The data security check now reports the number of scanned rows and the list of scanned columns (name and data type) for each scanned table in its scan results.

- Augment the `datasecurity` component with the sensitive data scanner library.

- Add a new `dogstatsd_require_listener` configuration option (disabled by default). When enabled, DogStatsD exits with a non-zero status if it cannot create any listener (UDP port, Unix socket, or named pipe), which would otherwise leave DogStatsD running with no way to receive metrics. Enable it so a process supervisor can detect and restart a non-functional DogStatsD.

- The `dogstatsd_stream_socket` configuration option, which lets DogStatsD listen for metrics on a Unix domain socket using stream mode (`SOCK_STREAM`), is now considered stable.

- gpu: add volatile ECC error metrics as a counterpart to the existing aggregate (lifetime) ones, including `errors.ecc.corrected.volatile` and `errors.ecc.sram.uncorrected_by_subtype.volatile`.

- The Agent flare now includes `ulimit.log` (the running Agent process' resource limits, on non-Windows platforms) and, on AIX, `svmon.log` (a per-segment virtual memory breakdown from `svmon -P`). These help diagnose resource-exhaustion issues without requiring a separate manual collection step from the host.

- gpu: Add NVLink fabric cluster UUID and clique ID to GPU tags (`gpu_fabric_cluster_uuid` and `gpu_fabric_clique_id`).

- GPU: emit `gpu.errors.xid`, a count of NVIDIA XID errors in each collection interval. `gpu.errors.xid.total` remains the lifetime total since the Agent started collecting events.

- In Kubernetes environments, the `apm_config.apm_non_local_traffic` and `jmx_use_container_support` defaults are now applied directly by the Agent binary instead of relying on the `datadog-kubernetes.yaml` file shipped in the container image. This preserves these Kubernetes defaults even when external tooling (such as the Datadog Operator or Helm chart) replaces `datadog.yaml`. Values explicitly set via a config file or environment variable continue to take precedence.

- The orchestrator check now collects `DatadogInstrumentation` (`datadoghq.com/v1alpha1`) custom resources as part of the out-of-the-box set indexed by the Kubernetes Explorer, alongside the other `datadoghq.com` custom resources. Collection requires `orchestrator_explorer.custom_resources.ootb.enabled` (enabled by default) and is skipped when the custom resource definition is absent from the cluster.

- OTLP: Mapped the `service.namespace` resource attribute to a `service.namespace` tag by default. OpenTelemetry semantic conventions only guarantee `service.name` / `service.instance.id` uniqueness within a `service.namespace`, so it is now preserved to keep service identity.

- On Windows and Linux, the fleet installer now also registers the Private Action Runner's on-demand executor with `dd-procmgr`.

- Private Action Runner in Datadog Agent now emits healthcheck metrics like its standalone counterpart.

- Private Action Runner: add the `private_action_runner.restricted_shell.allowed_system_services` setting to further restrict system-service action grants resolved by Datadog execution policies. Leaving the setting unset preserves the backend grants; configuring an empty map blocks all system-service operations.

- The Agent can now refresh secrets-managed API keys when a Remote Agent reports an Invalid API Key event.

- End User Device Monitoring no longer applies a preconfigured set of SaaS domain filters to `network_path.collector.filters`. Network Path filters are now empty by default in `end_user_device` mode, and user-configured filters are preserved unchanged.

- Bumped the Security Agent policies to [v0.83.0](https://github.com/DataDog/security-agent-policies/compare/v0.82.0...v0.83.0)

- Shared-library checks now honor an explicit `min_collection_interval: 0` as one-shot scheduling (the check runs a single time), matching the behavior of Python checks.

- Logs emitted by shared library checks are now routed through the Datadog Agent logger, so their output is formatted and level-filtered consistently with other checks instead of being written directly to standard output.

- SNMP device scans now walk devices using GetBulk by default, requesting only OIDs the device actually returns and adapting the max-repetitions on failures. This avoids the infinite loops and device crashes that could occur with the previous GetNext-based walk. SNMPv1 devices, which do not support GetBulk, continue to use the GetNext walk.

- SNMP device scans now report results incrementally while the scan is running instead of only after it completes, so large devices surface OIDs sooner.

- Upgrade OpenTelemetry Collector dependencies from v0.156.0 to v0.158.0 (core v1.62.0 to v1.64.0).

  See the full upstream changelogs: [collector-contrib v0.157.0](https://github.com/open-telemetry/opentelemetry-collector-contrib/releases/tag/v0.157.0), [collector core v0.157.0](https://github.com/open-telemetry/opentelemetry-collector/releases/tag/v0.157.0). [collector-contrib v0.158.0](https://github.com/open-telemetry/opentelemetry-collector-contrib/releases/tag/v0.158.0), [collector core v0.158.0](https://github.com/open-telemetry/opentelemetry-collector/releases/tag/v0.158.0).

- Agent Cloud Auth (delegated authentication) now reports why it did not start. When `org_uuid` is configured but no supported cloud provider is detected, the Agent logs a warning naming every credential source it checked instead of a debug-level message, and `agent status` shows the reason rather than only "not enabled".

- The `agent status` Delegated Authentication section now reports the AWS credential source in use for each managed API key (static environment variables, IRSA web identity, ECS/EKS container credentials, or EC2 IMDS), along with the last and next scheduled key refresh and the last error.

- Agent Cloud Auth (delegated authentication) failures now name the AWS credential mechanism that was attempted and what to check for it, instead of reporting a generic `missing AWS credentials`. Credential resolution that returns blank credentials, for example when the EC2 metadata service answers with an error document, is now treated as a failure rather than reported as a successful resolution.

- The `windows_certificate` check now supports `filters.include` and `filters.exclude` configuration to scope certificate collection by any emitted tag key (thumbprint, SAN, CN, friendly name, template, etc.) using Go regex patterns.

### Deprecation Notes

- The eBPF probes for GPU Monitoring are deprecated and are now disabled by default. Set `gpu_monitoring.enable_ebpf_probes` to `true` in `system-probe.yaml` to keep using them.

### Security Notes

- In FIPS mode, the default SNMPv3 authentication and privacy protocols used when `authKey`/`privKey` are set without an explicit `authProtocol`/`privProtocol` are now `SHA-256`/`AES` instead of `MD5`/`DES`, since `MD5` and `DES` are not FIPS 140-3 compatible. Outside FIPS mode, the defaults remain unchanged.
- The `otel-agent flare` command now writes its diagnostic archive into a private, unpredictably-named directory (restricted to the current user) instead of a predictable, world-readable path in the shared system temporary directory. Previously, on a multi-user host, another local user could read the flare contents (collected configuration, environment variables, and debug data), or redirect the archive by pre-creating a symlink at the predictable path.

### Bug Fixes

- APM: The trace-agent refreshes the API key and retries on a 403 only when `secret_refresh_on_api_key_failure_interval` is set (&gt; 0), now consistent with metrics and logs
- APM : Container tags resolution debug information is now stored using the tracer payload's deduplicated string table instead of inline strings, reducing the size of payloads that include this debug information.
- APM: Fix trace-agent crashes on malformed trace payloads that contain nil entries: nil spans, span links, or span events, and nil span-event attribute values. Nil entries are now dropped at decoding and conversion boundaries before payloads reach trace processing.
- APM: Fix a trace-agent crash when a `/v1.0/traces` trace chunk carries a trace ID that is not 16 bytes long. Chunk trace IDs are now normalized to 16 bytes during normalization instead of panicking on an out-of-bounds slice in the score and probabilistic samplers.
- APM: ProbabilisticSampler now properly uses the lower order bits of the trace ID on v1 traces.
- APM: Fix a trace-agent crash from unbounded recursion when decoding deeply nested attribute values; nesting depth is now bounded.
- APM: Fix a trace-agent crash when a span event attribute declares an array type but carries a missing array or array element.
- APM V1 trace endpoint now safely skips unknown fields, harvesting any inline strings they carry into the string table so that streaming-string references in later known fields continue to resolve correctly.
- Keep daemon specific default log filepaths now that the core agent's `log_file` configuration is populated.
- Windows: Fix a crash in the Agent when a statsd client configured to use a named pipe is closed without ever having successfully written to that pipe. The named pipe connection is established on first write, so closing such a client dereferenced a nil connection and terminated the process. Fixed by updating `datadog-go` to v5.9.1.
- \[DBM\] Bump `go-sqllexer` to v0.2.4 to fix a SQL normalization bug:
  - Stop treating backslash as a string escape character in SQL Server and Oracle string literals, which previously caused the obfuscator to swallow SQL past a literal like `ESCAPE '\'` and truncate the obfuscated query.
- Bump the embedded GoSNMP library to fix SNMPv3 engine-ID discovery and improve robustness of SNMP OID and varbind parsing.
- APM: Make the automatic library injection mode use the CSI driver only when the injector and library images come from configured Datadog registries. Images from other registries now fall back to init containers so Kubernetes can use the workload's image pull credentials.
- APM: Fixed Dynamic Instrumentation snapshot and log-probe uploads failing with a connection reset when the Logs product is disabled (`logs_enabled: false`). The debugger proxy now drains the request body before responding, so these uploads are dropped cleanly instead of resetting the tracer's connection.
- On ECS Managed Instances in daemon mode, the ECS workloadmeta collector no longer fails to start when the ECS Metadata v1 introspection endpoint is unreachable. The collector now falls back to the metadata v4 `/tasks` endpoint, which provides all the task data this deployment needs. Previously a v1 failure produced no `ECSTask` entities for the whole host, so `container.*` metrics were missing the `task_arn`, `task_family`, `task_version`, `ecs_cluster_name` and `ecs_container_name` tags. ECS EC2 daemon mode still requires metadata v1, which provides its task list.
- ECS cluster metadata (cluster name, cluster ID, region and AWS account ID) now falls back to the Agent's own task metadata when the ECS Metadata v1 introspection endpoint is unavailable. This fixes the orchestrator ECS check being skipped, the container lifecycle check reporting an empty cluster ID, and the flare missing its ECS section on ECS Managed Instances. The fallback applies to every non-Fargate launch type, so ECS EC2 deployments whose v1 introspection endpoint is unreachable now resolve cluster metadata instead of failing.
- The Cluster Agent no longer schedules endpoints checks against EndpointSlice endpoints that are not ready or are terminating. This restores the behavior of the `v1.Endpoints` code path, which only ever targeted ready addresses, and stops checks from erroring out against pods that are shutting down or failing their readiness probes.
- Fixed an issue where the Docker container collector could not inspect a container whose image declares a port *range* in its exposed ports (for example `EXPOSE 1061-1070`). Such containers were skipped entirely with an `invalid port '1061-1070': invalid syntax` error. Older Docker daemons return these ranges verbatim, which the container inspect decoder rejected. The Agent now expands port-range entries into individual ports so the container is collected normally.
- Fix `agent status` (and the `JSON`/`HTML` status renderers) showing a stale HA Agent `state` (active/standby). The status page previously read a snapshot cached by the periodic inventory metadata collector, which could lag up to `inventories_max_interval` (10 minutes by default) behind the agent's actual HA state. The status page now reflects the live state on every call, matching the behavior already used by the flare payload.
- Stop the Agent from repeatedly attempting to connect to a Kubelet on hosts that are not running on Kubernetes. This removes the recurring `Impossible to reach Kubelet through HTTPS, fallback to HTTP` warning logged by host-based and non-Kubernetes containerized installations. There is no change in behavior on Kubernetes nodes.
- Logs Agent: fixed a file source that could stop collecting permanently (`Bytes Read: 0`) after its log file was rotated or truncated below the offset already read. When a tailer starts, the Agent now checks that the offset stored for the file is still within it, and restarts from the beginning of the file when it is not.
- Fixed the kubernetes\_state.deployment.rollout\_duration metric occasionally reporting erroneous large values after a cluster agent restart
- Fix chassis type detection on Windows for convertible and detachable 2-in-1 devices (SMBIOS codes 31 and 32), which were previously reported as `Other`.
- Agent GUI now follows symbolic links when processing conf.d items.
- Fix the NVIDIA Jetson check so missing or frequency-only GPU fields in `tegrastats` output do not prevent other Jetson metrics from being collected.
- OTel Agent: The `datadog` extension no longer probes cloud metadata source providers when a hostname is already set in the configuration, avoiding spurious GCP metadata server requests on non-GCP hosts. See [open-telemetry/opentelemetry-collector-contrib#49241](https://github.com/open-telemetry/opentelemetry-collector-contrib/pull/49241).
- Fixed a crash loop in standalone `otel-agent` (`DD_OTEL_STANDALONE=true`) when deployed alongside a core Datadog Agent that injects `DD_REMOTE_CONFIGURATION_ENABLED=true` or `DD_AGENT_IPC_CONFIG_REFRESH_INTERVAL` into its environment, such as when using the Datadog Operator. Standalone mode now reliably disables Remote Configuration and config sync regardless of these environment variables, since it has no core agent IPC endpoint to use them with.
- The private action runner now stops heartbeating a task once the backend reports it no longer exists, instead of retrying indefinitely.
- Prevent double counting container memory limits when no limit is set on the container level and the limit is only set on the pod level. Use the metric `kubernetes.memory.limits` to only expose memory limits set explicitly at the container level.
- Fixed missing network IP metadata for the Agent when running containerized without host networking (the common case for most Kubernetes deployments). Previously, network metadata collection was skipped entirely in this case. The Agent now falls back to reporting the node's IP address, using the existing `kubernetes_kubelet_host` configuration value.
- A panic in a Rust-based shared-library check no longer aborts the whole Agent. The panic is now caught at the check boundary and surfaced as a check error, keeping the rest of the Agent running.
- Fixed the host SBOM scan (`sbom.host.enabled`) reporting no packages when the Agent runs as `dd-agent`. The scan walks the paths the enabled analyzers declare, one of which is `/root/buildinfo/content_manifests`, and `/root` is only traversable by root. Such a path is now skipped and the scan reports the packages it found.

### Other Notes

- Agent Data Plane has been bumped to version 1.4.0. See the [Agent Data Plane 1.4.0 release notes](https://github.com/DataDog/saluki/releases/tag/1.4.0).
- The `datasecurity` shared-library check now embeds the `datadog.sds` protobuf definitions used to report scan results. This is internal wiring; no scan results are emitted yet.
- The `ddot-collector` image now starts in standalone mode by default. Using it in bundled mode now requires to explicitly set `DD_OTEL_STANDALONE` to false.
- In ECS daemon mode, the ECS workloadmeta collector now returns a startup error when the ECS Metadata v1 instance data cannot be retrieved and no v1-independent task parser is available, instead of starting without a task parser. Startup is retried, so the collector recovers once the endpoint becomes reachable.
- The Agent no longer reports a `Check Execution Failure` health platform issue when a check run fails.
- Shared-library (Rust) checks can now submit event platform events as raw bytes, in addition to strings, allowing binary payloads such as protobuf.

# Datadog Cluster Agent

### Prelude

Released on: 2026-09-03 Pinned to datadog-agent v7.83.0: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7830).

### Upgrade Notes

- The `<namespace>` part of a `datadogmetric@<namespace>:<name>` external metric reference is now ignored. The referenced `DatadogMetric` is always looked up in the namespace of the `HorizontalPodAutoscaler` or `WatermarkPodAutoscaler` that holds the reference, so referencing a `DatadogMetric` owned by another namespace is no longer supported. Such a reference now resolves to a `DatadogMetric` that does not exist, which leaves the autoscaler without a metric value and unable to scale.

  To find out whether you are affected, list every external metric reference that carries an explicit namespace:

      kubectl get hpa --all-namespaces -o yaml | grep -E 'datadogmetric@[a-z0-9-]+:'
      kubectl get wpa --all-namespaces -o yaml | grep -E 'datadogmetric@[a-z0-9-]+:'

  References whose namespace is the namespace of the autoscaler holding them keep working unchanged. For every reference pointing at another namespace, create a `DatadogMetric` with the same query in the autoscaler's own namespace and point the autoscaler at it. The namespace can now be left out entirely, `datadogmetric@<name>` is a valid reference that resolves in the autoscaler's namespace.

### New Features

- Add the `kubernetes_state.pod.terminating` gauge to the Kubernetes State Core check. The metric reports a value of 1 for each pod from the moment its deletion timestamp is set until the pod leaves the informer.
- `DatadogInstrumentation` checks and logs configurations can now target Argo `Rollout` workloads.

### Enhancement Notes

- Add the `kube_argo_rollout` tag to pod metrics emitted by the Kubernetes State Core check for pods managed by Argo Rollouts.
- Add the `-l` and `--list` options to `datadog-cluster-agent status` to list available status sections. A section name can now be passed to the command to display only that section.
- The Cluster Agent's Prometheus HTTP Service Discovery provider now applies the OpenMetrics check template's `rename_labels` mapping to the tags derived from the SD target labels, in addition to the labels scraped from each target. Previously `rename_labels` only affected scraped metric labels, so a label supplied by the SD endpoint could not be renamed. No configuration change is required: the existing `rename_labels` in the `check_template` now covers both sources.
- The cluster-agent KSM auto-sharding dispatcher (`cluster_checks.ksm_sharding_enabled`) now supports a single `kubernetes_state_core` config that combines a shardable `cluster_unassigned` instance with a `cluster_aggregates_only` instance. The `cluster_unassigned` instance is sharded by resource type (pods/nodes/others) as before, and the `cluster_aggregates_only` instance (which does a full-pod watch and cannot be sharded) is dispatched alongside the pods shard. Previously such a multi-instance config disabled sharding. This lets KSM auto-sharding and the cluster-aggregate `.total` fix be enabled together from one config.
- The Cluster Agent now uses a single `ListWatch` call to track Kubernetes Node metadata, instead of one call per node.
- Karpenter `NodePool` autoscaling now tries to automatically resolve which `EC2NodeClass`/`NodeClass` to use when more than one exists for a given provider, based on the `NodePool`'s `kubernetes.io/os` and `kubernetes.io/arch` requirements. Each NodeClass's own `kubernetes.io/os`/`kubernetes.io/arch` labels are preferred when present, falling back to matching tokens in the NodeClass name (e.g. `linux-amd64`) otherwise; if neither signal uniquely identifies a NodeClass, or the label- and name-based signals disagree, the ambiguity is left unresolved. Previously, having more than one NodeClass of the same provider type always caused NodePool creation/update to fail with a "too many NodeClasses found" error; this is still the outcome when the disambiguation above can't resolve to a single NodeClass. Additionally, when both a manual Karpenter `EC2NodeClass` and an EKS Auto Mode `NodeClass` exist in the cluster, the EKS Auto Mode `NodeClass` is now preferred; previously the `EC2NodeClass` was always preferred.

### Security Notes

- The Cluster Agent external metrics provider now resolves `datadogmetric@` references in the namespace of the requesting object rather than the namespace embedded in the metric name. Previously, a workload could read the value of a `DatadogMetric` owned by another namespace, and keep that `DatadogMetric` active so that its Datadog queries kept running.

### Bug Fixes

- Fix `kubernetes_state.container.cpu_requested` and `kubernetes_state.container.memory_requested` to use the effective requests reported by Kubernetes after an in-place vertical resize. Pod spec requests remain the fallback when status resources are unavailable.
- The Cluster Agent no longer opens a second, redundant cluster-wide `ListWatch` for Kubernetes Nodes. Setting `kubernetes_node_labels_as_tags` or `kubernetes_node_annotations_as_tags` used to start an extra Node metadata watch in addition to the one already used to populate the Node cache; that extra watch has been removed and label/annotation-as-tags extraction now relies solely on the existing Node cache.

</Release>

<Release version="7.82.3" date="August 26, 2026" published="2026-08-26T11:41:13.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.82.3">
# Agent

### Prelude

Released on: 2026-08-26

- Please refer to the [7.82.3 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7823) for the list of changes on the Core Checks

### Enhancement Notes

- Agents are now built with Go `1.26.7`.

### Bug Fixes

- The Dynamic Instrumentation proxy in the trace-agent no longer drops debugger data when `logs_enabled` is left unset. Data is only dropped when `logs_enabled` (or the deprecated `log_enabled`) is explicitly set to `false`.
- On Windows, fixed an issue where upgrading the Agent without providing `DDAGENTUSER_PASSWORD` could lock out a domain Agent user account. The installer now leaves `dd-procmgr-service`, and the components it supervises, disabled until the Agent user password is provided again.

# Datadog Cluster Agent

### Prelude

Released on: 2026-08-26 Pinned to datadog-agent v7.82.3: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7823).

</Release>

<Release version="7.82.2" date="August 20, 2026" published="2026-08-20T07:20:26.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.82.2">
# Agent

### Prelude

Released on: 2026-08-19

- Please refer to the [7.82.2 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7822) for the list of changes on the Core Checks

### New Features

- When `infrastructure_mode: end_user_device` is set, the `logon_duration` feature is now enabled automatically, so operators no longer need to set `logon_duration.enabled: true` separately. This setting can still be overridden explicitly in the configuration file if needed.

### Enhancement Notes

- The Agent's embedded Python has been upgraded from 3.13.14 to 3.13.15
- Agents are now built with Go `1.26.6`.

### Bug Fixes

- APM: Raise the messagepack decoder allocation limit for the span `meta_struct` field to 10MiB, so large `meta_struct` entries (as written by LLM Observability) are no longer rejected by the package-wide 500,000 element limit. Other span fields keep the lower limit.
- Disable GPU parallel collection by default to avoid triggering NVML concurrency bugs.
- gpu: fix concurrent calls to NVML GetFieldValues API, that could cause missing NVLink metrics.

### Other Notes

- gpu: GPU inventory payload and GPU tags are only emitted when GPU monitoring is enabled

# Datadog Cluster Agent

### Prelude

Released on: 2026-08-19 Pinned to datadog-agent v7.82.2: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7822).

</Release>

<Release version="7.82.1" date="August 10, 2026" published="2026-08-10T15:37:55.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.82.1">
# Agent

### Prelude

Released on: 2026-08-11

- Please refer to the [7.82.1 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7821) for the list of changes on the Core Checks

### Bug Fixes

- Windows: Fixed an issue where an explicit `DDAGENTUSER_KEEP_RIGHTS` or `DDAGENTUSER_NAME` value passed as an install argument to a Fleet Automation-triggered Windows Agent install/upgrade could be silently overridden by a stale fallback value (respectively from the registry and from the running service account).
- Fix an issue where GPU monitoring could trigger a kernel panic on multi-GPU nodes with Hopper/Blackwell GPUs.

# Datadog Cluster Agent

### Prelude

Released on: 2026-08-11 Pinned to datadog-agent v7.82.1: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7821).

</Release>

<Release version="7.82.0" date="August 5, 2026" published="2026-08-05T13:58:53.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.82.0">
# Agent

### Prelude

Released on: 2026-08-05

- Please refer to the [7.82.0 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7820) for the list of changes on the Core Checks

### Upgrade Notes

- Automatic multi-line log detection (`logs_config.auto_multi_line_detection`) is now enabled by default. Multi-line log messages such as stack traces and JSON blobs are aggregated into a single log entry out of the box, instead of being split into separate entries. To restore the previous behavior, set `logs_config.auto_multi_line_detection` to `false` (or the environment variable `DD_LOGS_CONFIG_AUTO_MULTI_LINE_DETECTION=false`).
- Change default EVP track for AI usage to `eudm-intake` and disable AI usage agent desktop monitoring by default (the monitor stays in idle mode). AI usage agent's `ai_usage_native_host.yaml` configuration file is now regenerated from the packaged template on every Agent install and upgrade to ensure the default changes take effect. This is a deliberate short-term measure: it forces already-installed machines onto the updated defaults. Once the defaults are settled (expected within a few Agent releases), the installer will go back to preserving an existing `ai_usage_native_host.yaml` and only creating it when missing. Note that any customization made to the config file is discarded on every Agent install and upgrade, and must be re-applied afterwards.
- APM: Bump the default version of the `datadog-apm-library-js` package installed by the Datadog installer from major version 5 to major version 6, following the release of `dd-trace-js` v6.
- APM: Updates the default JS (Node.js) library used for Kubernetes auto-instrumentation (Cluster Agent admission controller) from major version 5 to major version 6.
- The legacy `viper` based configuration backend has been removed. The `DD_CONF_NODETREEMODEL` environment variable and the `conf_nodetreemodel` configuration setting no longer have any effect and can be removed from your configuration. The Agent now always uses the improved configuration implementation.
- serverless-init no longer forces `DD_TRACE_PROPAGATION_STYLE=datadog` during tracer auto-instrumentation. The tracer's own default (which includes W3C `tracecontext` and `baggage` in addition to `datadog`) now applies, and a customer-provided `DD_TRACE_PROPAGATION_STYLE` is respected. Applications that relied on serverless-init restricting propagation to `datadog` only should set `DD_TRACE_PROPAGATION_STYLE=datadog` explicitly.

### New Features

- Private Action Runner: add the `com.datadoghq.remoteaction.rshell.runRemediationCommand` action. It behaves like `runCommand` but runs the restricted shell in remediation mode, which additionally permits file-target output redirections (`>`, `>>`, `2>`, `&>`, `&>>`) and write-oriented builtins such as `truncate`, all confined to the configured allowed paths. The action is not enabled by default and must be explicitly added to the runner's actions allowlist.

- `agent flare` now includes diagnostic artifacts from the Agent Data Plane (ADP) process when `data_plane.enabled` is set to `true`. If ADP is unreachable at flare time, an `UNREACHABLE.txt` file containing the connection error is written to ADP's subdirectory and the rest of the flare completes normally.

- When `infrastructure_mode` is set to `cloud_cost_only`, the Agent adds an `infra_mode:cloud_cost_only` tag to metrics from selected integrations. Use `integration.cloud_cost_only.tagged` to list which checks receive the tag; when the list is empty (the default), all checks are tagged.

- Add a new `dogstatsd_no_aggregation_pipeline_workers_count` configuration option to control the number of parallel workers processing messages in the no-aggregation pipeline. Defaults to `1` to preserve existing behavior.

- Adds Datadog CSI driver telemetry to COAT, including volume publish and unpublish attempts as well as library resolution, download count and duration, cleanup, cache size, cached library count, and library volume link metrics.

- The Datadog OTLP connector now scales APM stats (hits, errors, duration) by the probabilistic head-based sampling weight carried in the W3C `tracestate` (`th` threshold and `p` power-of-two encodings). When upstream head-based sampling has dropped a fraction of traces, the computed trace metrics are scaled up to reflect the true traffic volume instead of only the sampled subset.

- Envoy Gateway AppSec protection can now run in `sidecar` mode over a Unix domain socket. Datadog injects the `serviceextensions` ext\_proc container into Envoy Gateway data-plane pods, and Envoy Gateway communicates with it through an Envoy Gateway `Backend`.

  This behavior is selected by `cluster_agent.appsec.injector.mode`, which now defaults to `sidecar`. Envoy Gateway must have the Backend extension API enabled with `extensionApis.enableBackend: true`; if it is disabled, the cluster agent warns and does not change Envoy Gateway configuration.

  This is a behavior change for AppSec-enabled Envoy Gateway deployments: they now default to sidecar injection instead of external Service mode. To keep the previous behavior, set `cluster_agent.appsec.injector.mode` to `external`.

- Added an experimental telemetry error log forwarder. Disabled by default, the Agent forwards records logged at `ERROR` level or higher to the COAT intake so Datadog Engineers can aggregate Agent errors across customer organizations. The forwarder shares the agent telemetry transport, inheriting endpoint and compression settings from the agent telemetry configuration.

- On Windows, `datadog-installer` now honors the `DD_AGENT_MAJOR_VERSION` and `DD_AGENT_MINOR_VERSION` environment variables, matching the Linux and macOS install scripts.

- GPU: add the `gpu.device.needs_recovery` metric, which reports whether a GPU requires a recovery action (such as a reset or node reboot) as exposed by NVML's GPU recovery action field. The value is `0` when no action is needed and `1` otherwise, and the metric is tagged with `recovery_action` (`none`, `reset`, `reboot`, `drain` or `drain_and_reset`).

- The `agent status` command now includes a "Logs Agent Backpressure" section reporting per-component utilization of the logs pipeline and an overall `HEALTHY`/`WARNING`/`SATURATED` state, making it easier to see which pipeline stage is the bottleneck when logs are delayed.

- On macOS, the Agent can now be restarted directly from the web-based GUI (Agent Manager).

- New Agent Secret Backend: "windows.regkey"

- APM Single Step Instrumentation now supports setting tracer configuration options via the `admission.datadoghq.com/apm-inject.tracer-configs` pod annotation, the annotation-based equivalent of the `apm_config.instrumentation.targets[].ddTraceConfigs` option. The value is a JSON array of objects (for example `[{"name":"DD_PROFILING_ENABLED","value":"true"}]`) and each entry's name must start with the `DD_` prefix.

- NetFlow: automatically detect and split Cisco FirePower/ASA bidirectional (NSEL) flow records into two unidirectional flow events. NSEL records carry initiator→responder and responder→initiator byte/packet counts in NFv9 fields 231/232/298/299; the agent now captures these fields via built-in mappings and emits a separate flow for each direction with correctly swapped src/dst addresses, ports, and interfaces. No user configuration is required.

### Enhancement Notes

- Malformed `ad.datadoghq.com/service.*` and `ad.datadoghq.com/endpoints.*` annotations on Kubernetes services are now reported as Autodiscovery misconfiguration health events when the health platform is enabled. The issue is resolved automatically once the annotation is fixed.
- Adds Agent Data Plane packaging and launchd service support to macOS Agent packages.
- Adds Agent Data Plane packaging to Windows Agent MSI installs. ADP is supervised by dd-procmgr via `processes.d/datadog-agent-data-plane.yaml`, written by the fleet installer during `postinst` (same pattern as DDOT on Windows).
- Emit datadog.cluster\_agent.kubernetes\_actions.running when kuberenetes actions product is enabled and running.
- Podman receiver metrics collected via the Datadog Distribution of OpenTelemetry (DDOT) Collector are now correctly classified with origin `opentelemetry_collector_podmanreceiver` instead of falling back to `opentelemetry_collector_unknown`.
- Added the `data_plane.stop_timeout` configuration setting, which controls the graceful shutdown budget for the Agent Data Plane (ADP). When unset, it derives its value from `aggregator_stop_timeout + forwarder_stop_timeout`, so customizing either of those component timeouts now extends ADP's shutdown window in lockstep with the core Agent.
- On startup the Datadog Agent now validates the system-probe configuration against its schema and reports any violations through the Agent Health pipeline. Only the values the customer set in the configuration are validated. This can be disabled with `health_platform.invalidsysprobeconfig_check.enabled`.
- On Windows, the AI Usage Chrome Native Messaging host is now delivered as a fleet-managed Agent extension that is only installed when End User Device Monitoring is enabled (`infrastructure_mode: end_user_device`). It is no longer unconditionally installed by the MSI, and is skipped on Agent upgrades when End User Device Monitoring is disabled.
- APM: Added cardinality limits to client-side stats computation in the stats concentrator. These limits are no-op in the agent and are intended for use by the Go tracer.
- Agents are now built with Go `1.26.5`.
- dd-procmgrd: Add write RPCs (Create, Start, Stop, ReloadConfig, GetConfig) for runtime control of managed processes.
- On Windows, the `ddinjector` system-probe telemetry now negotiates the counter contract version with the installed driver, reporting new crash and boot-recovery counters when the driver supports them and degrading gracefully against older drivers.
- Agent Cloud Authentication (delegated authentication) now discovers AWS credentials from additional sources on Agent builds that include EC2 metadata support. In addition to the previously supported static credentials (`AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY`) and EC2 instance metadata (IMDS), it now supports IAM Roles for Service Accounts (IRSA, via `AWS_WEB_IDENTITY_TOKEN_FILE` and `AWS_ROLE_ARN`) and ECS task role / EKS Pod Identity container credentials. The region used for the STS exchange follows the `delegated_auth.aws.region` setting, then `AWS_REGION` / `AWS_DEFAULT_REGION`, falling back to `us-east-1`. Shared config and profile loading (`AWS_PROFILE`, `~/.aws`) is not used.
- When `dogstatsd_flush_incomplete_buckets` is enabled, the DogStatsD server now also finishes processing in-flight packets before sending metrics in the current (incomplete) aggregation window.
- Add a `retry_on_failure` configuration block to the Datadog serializer exporter (DDOT and OSS Datadog exporter). The defaults (initial interval `2s`, multiplier `2`, maximum interval `64s`, maximum elapsed time `15m`) and the default `sending_queue` settings (`queue_size: 300`, `num_consumers: 1`) and `timeout: 20s` mirror the historical forwarder defaults so behavior remains consistent when upgrading. Users can tune these to increase throughput; for example, raising `num_consumers` lets the exporter dispatch multiple HTTP requests to the Datadog intake in parallel.
- Add the `datadog.serializerexporter.UseSyncForwarder` feature gate (Alpha, disabled by default) to the Datadog serializer exporter. When enabled via `--feature-gates=+datadog.serializerexporter.UseSyncForwarder`, metric submission becomes synchronous: HTTP errors propagate back through `ConsumeMetrics`, are counted in the standard OpenTelemetry exporter metrics, and are retried by the exporter helper — fixing the silent-drop behavior where `otelcol_exporter_send_failed_metric_points` would not increase even on intake failures. Applies to both the embedded Datadog OpenTelemetry Collector (DDOT) and the OSS Datadog exporter (`opentelemetry-collector-contrib`). The Agent OTLP Ingestion pipeline is unaffected.
- GPU: improve GPU metric collection latency by collecting NVIDIA GPU collector data in parallel.
- Add a new `trace_container_tag_promotion` configuration option to the `infraattributes` processor in the Datadog Agent's embedded OpenTelemetry Collector (DDOT). When set to `duplicate` or `rename`, custom tags emitted by the processor (for example tags produced by `podLabelsAsTags`) are written with a `datadog.container.tag.` prefix so they are promoted into Datadog container tags and become visible in the Infrastructure tab of a span. The default value `off` preserves the previous behavior. Known Datadog and OpenTelemetry container semantic conventions, as well as `service` / `env` / `version` tags, are exempt and always written under their canonical key.
- OTLP ingest: The `container_tag_promotion` mode used by the `infraattributes` processor on the OTLP ingest traces pipeline is now configurable through `otlp_config.traces.infra_attributes.container_tag_promotion` (environment variable `DD_OTLP_CONFIG_TRACES_INFRA_ATTRIBUTES_CONTAINER_TAG_PROMOTION`). Accepted values are `off`, `duplicate` and `rename`. The default `off` matches the processor's own default, so promotion is opt-in and the previous behavior is preserved.
- Installer: The default install script now honors the `DD_PROCESS_CONFIG_PROCESS_COLLECTION_ENABLED`, `DD_PROCESS_CONFIG_CONTAINER_COLLECTION_ENABLED`, and `DD_PROCESS_CONFIG_PROCESS_DISCOVERY_ENABLED` environment variables (and their `DD_PROCESS_AGENT_*` aliases) to set the corresponding `process_config` collection toggles in `datadog.yaml` at install time
- Migrate to a local version of the Datadog connector rather than the implementation from opentelemetry-collector-contrib. The connector's configuration and behavior are currently unchanged.
- \[netflow\] Add the `tos`, `dscp`, and `dscp_name` fields to NetFlow flow payloads, exposing the IP Type of Service byte along with its decoded DSCP value and standard RFC name (for example `EF` or `AF41`). The `dscp_name` mapping now includes all IANA-registered codepoints.
- Reduce idle agent memory usage when Network Device Monitoring NetFlow is disabled (the default). The NetFlow flow aggregator is no longer allocated when the feature is off.
- Reduced per-sample allocations in the metrics aggregator, lowering CPU and garbage-collection overhead on Agents processing a high volume of metrics.
- Reduced the CPU and memory overhead of sensitive-data scrubbing when collecting Kubernetes resources for the Orchestrator Explorer. The command-line tokenizer regular expression is now compiled once instead of on every scrubbing call, and redundant per-token work was removed.
- Reduced memory allocations when converting Python strings to C strings across the RTLoader boundary.
- OTLP trace ingestion now reports `otel.scope.name` and `otel.scope.version` in addition to the deprecated `otel.library.name` and `otel.library.version`. Use the `disable_otel_scope_convention` feature gate to stop reporting the new keys.
- The OTel Agent `--sync-delay` flag (env `DD_SYNC_DELAY`) now defaults to `30s` instead of `0`. The OTel Agent will retry synchronizing its configuration from the core Agent for up to 30 seconds at startup before failing, which makes startup more resilient when the core Agent is not yet ready. This default can still be overridden via the flag or environment variable.
- The Private Action Runner now submits execution metrics through DogStatsD.
- Private Action Runner: restricted shell actions now combine Datadog execution policy allowlists with operator-configured command and path allowlists. Read-only and read-write allowed paths are handled separately, so remediation commands can write only to paths allowed by both policies.
- Persist low-priority transactions (such as retried metrics) to disk during a graceful Agent shutdown instead of dropping them.
- Added GPU attribution support for Kubernetes Dynamic Resource Allocation allocations reported by the kubelet PodResources API. The Agent maps NVIDIA DRA `gpu-N` device names to local GPU indexes.
- Reduced the CPU usage of log file scanning when tailing container logs.
- The image-size portion of the available-disk check performed before a container image SBOM scan now runs only for scans that export the image to a tarball on disk (the containerd default mode and the Docker collector), and is skipped for scans that read the image layers in place (CRI-O, and containerd with `sbom.container_image.overlayfs_direct_scan` or `sbom.container_image.use_mount`). The flat `sbom.container_image.min_available_disk` floor still applies to every scan, and its default is lowered from 1GB to 10MB.
- Bumped the Security Agent policies to [v0.82.0](https://github.com/DataDog/security-agent-policies/compare/v0.81.0...v0.82.0)
- serverless-init no longer overrides customer-provided `CORECLR_ENABLE_PROFILING`, `CORECLR_PROFILER`, `CORECLR_PROFILER_PATH`, or `DD_DOTNET_TRACER_HOME` values when auto-instrumenting a bundled .NET tracer. Any value already present in the environment is now respected.
- The OTel Datadog exporter will now emit `otel.datadog_exporter.metrics.running.fargate{task_arn}` for AWS ECS Fargate workloads so that each workload has its own metric. Host-based workloads continue to use `otel.datadog_exporter.metrics.running{host}` unchanged.
- APM: The minimal OTel-to-Datadog span conversion used for APM stats now preserves the raw W3C `tracestate` in the `w3c.tracestate` span tag, allowing downstream consumers to recover head-sampling probability.
- Update OpenTelemetry Collector dependencies to version 0.155.0. Notable upstream changes included in this update:
  - **Datadog extension**: Fixed `tls.insecure_skip_verify` being ignored.
- The discovery service map (`discovery.service_map.enabled`) now captures HTTP/2 traffic in addition to HTTP and TLS, so gRPC and other HTTP/2 service-to-service calls appear in the service map.
- Universal Service Monitoring (USM) now uses the direct consumer for HTTP monitoring by default on kernels `>= 5.8.0`, reducing the latency of HTTP event collection. On older kernels the batch consumer continues to be used. The previous behavior can be restored by setting `service_monitoring_config.http.use_direct_consumer` to `false`.
- On Windows, Universal Service Monitoring now derives the `service`, `env`, and `version` tags for IIS-hosted applications from the environment variables configured in `applicationHost.config` (application pool and `aspNetCore` `environmentVariables`) and in each application's `web.config`, in addition to `appSettings` and `datadog.json`. The resolution follows the .NET tracer's precedence so USM tags match what APM reports for the same application.
- On Windows, when the Agent is installed with the fleet OCI layout and the **DDOT** extension, the OpenTelemetry Collector is now supervised by **dd-procmgr** using a process definition under the Agent install layout, consistent with Linux. The **datadog-otel-agent** Windows service may still be registered for rollback, but the core Agent does not start it when that process definition file is present under `processes.d` and `process_manager.enabled` is true, so DDOT is not run twice while procmgr is on. The core Agent starts **dd-procmgr-service** when `process_manager.enabled` is true (set it to false to avoid starting the process manager if needed). The DDOT extension install hook writes `processes.d` and restarts **dd-procmgr-service** only under the same `process_manager.enabled` setting. Without the fleet `processes.d` DDOT definition, behavior is unchanged: the Agent still starts **datadog-otel-agent** when the collector is enabled.
- On Windows, when `process_manager.enabled` is true and the fleet installer writes `processes.d/datadog-agent-action.yaml`, the Private Action Runner is supervised by dd-procmgr instead of the legacy `datadog-agent-action` Windows service. When the processes.d definition is absent or process manager is disabled, the Agent continues to start PAR via SCM when `private_action_runner.enabled` is true.

### Security Notes

- The trace-agent now scrubs sensitive values (passwords, tokens, API keys) from the `command_line` field of SSI `injection-metadata` telemetry payloads before forwarding them.

### Bug Fixes

- Logs file tailer: directory entries returned by a configured `path` glob are now skipped instead of being opened as files. Previously the tailer attempted to open the directory and surfaced a misleading `Access is denied` error from the OS (on Windows in particular), which sent troubleshooting toward permissions. A non-wildcard `path` that resolves to a directory now returns a descriptive error explaining that a file or glob is required.
- Removed `user_id` field from AI usage event to avoid inconsistency between Chrome extension mode and desktop-monitor mode. `user_id` is taken care of by Datadog backend.
- \[DBM\] Bump `go-sqllexer` to v0.2.3 to fix a SQL normalization bug:
  - Preserve bracket-quoted T-SQL identifiers containing spaces (e.g. `[Column With Spaces]`) so they are no longer corrupted during normalization.
- Fixed a bug in the Go-native disk check (diskv2) where disk IO metrics (such as `system.disk.read_time`, `system.disk.write_time`, `system.disk.read_time_pct`, and `system.disk.write_time_pct`) were reported for every device regardless of the `device_include` and `device_exclude` settings. IO metrics are emitted only for devices whose partitions pass the configured device filters, matching the behavior of partition metrics and of the Python disk check.
- ECS daemon-scheduled tasks now emit the `daemon_task_definition_arn` tag instead of `task_definition_arn`, enabling correct entity resolution in the ECS Explorer.
- Fix a bug where the `CELSelector` field was not included in the Autodiscovery config digest. Configs with different CEL selectors were incorrectly treated as identical, which could cause the wrong workload filter rules to be applied.
- Fixed automatic multi-line Go stack trace aggregation for container-based log formats.
- CWS: Network packets of NAT-translated connections (for example source-port masquerading) are now correctly attributed to the owning process on the ingress path.
- Fix a bug where a Data Observability query action targeting a bare host shared by multiple database instances on different ports (e.g. two postgres instances on the same host) could drop every instance on that host from the remaining file-provider configuration, silently stopping normal DBM collection on the untargeted ports. Only the instance actually selected as the Data Observability check is now excluded.
- Fix Docker log parsing for TTY-mode containers when a single log line exceeds the 16KB Docker buffer size. Previously, the Agent retained the per-chunk timestamp prefix that Docker inserts at every 16KB boundary, causing those timestamps to appear inside the collected log content. The parser now strips the intermediate timestamps so the log line is reassembled correctly.
- Fix `panic: runtime error: invalid memory address or nil pointer dereference` when a DogStatsD distribution metric is submitted with a non-finite sample rate such as `NaN` (for example `s.dist:1|d|@nan`). The invalid sample rate is now treated as unsampled instead of corrupting the sketch's sample count.
- Fix stale metrics for locally-owned `DatadogPodAutoscaler` objects. The cluster agent cached the CRD object only on the first reconcile; status subresource writes (which do not bump `.metadata.generation`) were never reflected in the cached copy, causing the following metrics to report their initial values until the next spec or annotation change: `datadog.cluster_agent.autoscaling.workload.status.desired.replicas`, `datadog.cluster_agent.autoscaling.workload.status.vertical.desired.container.cpu.request`, `datadog.cluster_agent.autoscaling.workload.status.vertical.desired.container.cpu.limit`, `datadog.cluster_agent.autoscaling.workload.status.vertical.desired.container.memory.request`, `datadog.cluster_agent.autoscaling.workload.status.vertical.desired.container.memory.limit`, `datadog.cluster_agent.autoscaling.workload.autoscaler_conditions`. The controller now unconditionally refreshes the cached object on every reconcile.
- Fix error raising `could not set '*.use_http' unknown key` raised by the Fips proxy.
- Fixed the Cluster Agent `kubeapiserver` workloadmeta collector so that resources are now discovered on a non-preferred API group version when they are not served on the group's preferred version.
- Fix a potential goroutine hang in the OTLP logs exporter during graceful shutdown. The blocking channel send now respects context cancellation, allowing the exporter to stop cleanly instead of waiting indefinitely for the downstream logs pipeline.
- Fix the OTLP logs exporter sending corrupt data downstream when JSON marshaling of a log record fails. The malformed record is now dropped and the error is logged.
- Fix a bug in the OTel Agent where OTLP explicit-bucket histograms with bounds beyond the internal sketch's representable range could exhaust memory and crash the process, or — once the loop was bounded — record non-finite `min`, `max`, `sum`, and `avg` in the resulting sketch. Saturating bounds are now clamped to the largest representable finite bin.
- Fixed a rare process-agent panic when collecting metrics for a container with an empty or short (&lt;12 character) ID.
- Fixed an issue where the Cluster Agent could fail to schedule Prometheus checks for Kubernetes Services newly annotated with `prometheus.io/scrape` until the Cluster Agent was restarted or the backing workload was scaled.
- Fixed a slow memory leak in the remote workloadmeta collector where a new context was created on every stream reconnection attempt without cancelling the previous one. This could cause unbounded growth in the number of orphaned contexts when the remote workloadmeta endpoint was repeatedly unreachable.
- Fix a bug where a SAP HANA instance targeted by a Data Observability query action could run as multiple duplicate check instances at once, causing duplicate database monitoring collection. The targeted instance is now correctly excluded from the remaining file-provider configuration.
- Fix a bug where editing an active Data Observability query action's monitor (e.g. changing its query count) could cause the previously excluded database instance to run as a duplicate check instance again, alongside its Data Observability check, causing duplicate database monitoring collection.
- Fixed several issues in the SBOM runtime usage enrichment ("package in use"). The `HasSetSuidBit` and `RunningAsRoot` properties are now reported as `false` for packages that are not in use (previously absent), so consumers can distinguish "not in use" from "unknown". A package's setuid observation is no longer cleared by a later access to one of its non-setuid files. An idle workload whose container image scan is slow is now enriched once the scan completes, instead of being skipped after a fixed number of retries.
- On Windows, fix Cross-Org Agent Telemetry (COAT) reporting for DDOT under **dd-procmgr**: `runtime.agent_service_procmgr_configured{service:ddot}` now checks `InstallPath\processes.d` where the installer writes the YAML, instead of `ProgramData\Datadog\dd-procmgr\processes.d`.
- On Windows, align the DDOT `processes.d` definition with the legacy **datadog-otel-agent** SCM service: run `otel-agent.exe` without CLI arguments and set only `DD_OTELCOLLECTOR_INSTALLATION_METHOD=bare-metal`. The collector resolves `otel-config.yaml` and `datadog.yaml` from ProgramData when paths are omitted, and reads fleet policies from the registry (with a stable managed-path fallback) when `DD_FLEET_POLICIES_DIR` is not set in the process environment. This fixes DDOT failing to start under **dd-procmgr** when the baked `processes.d` invocation or fleet policy handling did not match that SCM path (including an unsubstituted `DD_FLEET_POLICIES_DIR` placeholder).
- On Windows, stop baking `DD_FLEET_POLICIES_DIR` into **Private Action Runner** and **Agent Data Plane** `processes.d` definitions at install time. Those processes now resolve fleet policy location from the Windows registry via `FleetConfigOverride`, so fleet policy updates are not blocked by a stale install-time path.
- Exclude the workloadmeta process collector from the IoT Agent binary using the `systemprobechecks` build tag, reducing binary size.
- OTel Agent: Disable v3 series API shadow sampling, which is incompatible with the zlib compression the OTel Agent forces for the metrics intake.
- Fixed a bug where CPU core count fields (`cpu_cores`, `cpu_logical_processors`) were missing from OTel host metadata payloads in gateway topologies.
- Preserve OTLP attributes on Kubernetes manifests sent by the OTel Datadog exporter in a dedicated manifest attributes field.
- \[oracle\] Removed the `oracle_client_lib_dir` instance configuration option, which did not work as expected. The Oracle client library path must be set before runtime for the dynamic linker to properly link the library, so pointing to it from the check configuration is too late to take effect. Configure the client library location through the runtime linker instead, as described in the Oracle integration setup documentation: <https://docs.datadoghq.com/integrations/oracle/?tab=linux#prerequisite>
- Cluster Agent: fixes a bug where the Cluster Agent did not reset the dangling config and unscheduled check metrics when its internal state was reset.
- The container image SBOM collector no longer attempts to export an image to a tarball when the image's layers are not present in the containerd content store, as happens with remote snapshotters such as nydus. The export could never succeed for these images, and the scan was retried indefinitely, keeping the scan worker busy and increasing Agent memory usage. Such scans are now skipped instead of retried.
- Skip cloud provider network ID probes when `cloud_provider_metadata` is empty. Previously, `GetNetworkID` would probe GCE and EC2 metadata endpoints on every host metadata collection cycle (every 5-30 minutes) even when all cloud providers were disabled by configuration. The probes now short-circuit immediately, avoiding unnecessary work in on-premises environments.
- Reduce log noise in on-premises environments by changing the `could not get network metadata` message from INFO to DEBUG level. In environments where cloud provider metadata is disabled by configuration, this message was emitted at every metadata collection cycle (every 5-30 minutes), causing unnecessary log volume. The message is still available at DEBUG level for diagnostic purposes.
- Fix a potential panic and loss of latency data when discovery service map HTTP or HTTP/2 statistics are merged across collection cycles (for example when a client misses a cycle or multiple clients are registered).
- Disable v3beta metrics intake shadow payloads when zlib compression is used.
- On Windows, fixed an issue where a failed Agent upgrade could leave some files without their permissions restored after the MSI rolled back, which could prevent the Agent from starting.

### Other Notes

- Agent Data Plane has been bumped to version 1.3.1. See the [Agent Data Plane 1.3.1 release notes](https://github.com/DataDog/saluki/releases/tag/1.3.1).
- The anomaly detection observer component no longer allocates memory or starts background goroutines when anomaly detection is disabled (the default). This reduces idle Agent PSS.
- Update `libgcrypt` to 1.12.2.

# Datadog Cluster Agent

### Prelude

Released on: 2026-08-05 Pinned to datadog-agent v7.82.0: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7820).

### New Features

- The Cluster Agent's Prometheus HTTP Service Discovery provider supports an optional `exclude_filter` field per endpoint entry. The field accepts a CEL expression evaluated against each discovered target's `host`, `port`, and `labels` fields. Targets for which the expression returns `true` are skipped at collection time.

### Enhancement Notes

- The Cluster Agent's leader election now uses a dedicated Kubernetes API server client with independently managed client-side rate limiting to be more resilient.
- Reduce scale-up stabilization windows for built-in DPA presets: Optimize Cost from 300s to 190s, Optimize Balance from 600s to 130s, and Optimize Performance from 900s to 70s. This allows faster scale-up response for workloads using autoscaling profiles.
- Cluster check stickiness is now enabled by default. The dispatcher biases check placement toward the runner where a check previously ran, reducing unnecessary check migrations. The behavior can be tuned or disabled via the following configuration options:
  - `cluster_checks.stickiness_enabled` — enable or disable stickiness (default: `true`)
  - `cluster_checks.stickiness_factor` — multiplier applied to check cost when computing the bias (default: `4.0`)
  - `cluster_checks.stickiness_upper_limit` — maximum bias applied regardless of check cost (default: `1.0`)
  - `cluster_checks.stickiness_lower_limit` — minimum bias applied when stickiness is enabled (default: `0.05`)

### Bug Fixes

- Fixed APM Single Step Instrumentation injecting the library twice when the admission webhook is reinvoked (for example on GKE Autopilot, where another mutating webhook triggers reinvocation). In CSI injection mode the pod has no init container, so the re-admission guard failed to detect that the pod was already instrumented and appended the injector to `LD_PRELOAD` a second time. The guard now also checks for the instrumentation volume, which is present in every injection mode.
- Fixed an issue where the Cluster Agent could associate a pod's detected languages with the wrong Deployment. The Cluster Agent now only attributes a pod's detected languages to a Deployment when the pod is owned by a ReplicaSet and the ReplicaSet derived from the pod name matches the owner ReplicaSet. Pods that are not owned by a ReplicaSet are no longer considered for Deployment-level language detection.
- Fix `cluster_checks.nodes_reporting` gauge drifting upward across leader elections. The metric is now correctly decremented when the cluster agent loses leadership and the node store is reset.
- Fix <span class="title-ref">agent</span> commands in DCA (listener should always be started)
- Fix permission in docker image when executing "/readsecret.sh" script with dd-agent user
- Fixed an issue in the KSM check where cluster-aggregate metrics (`kubernetes_state.container.<cpu|memory>_requested.total`, `kubernetes_state.container.<cpu|memory|gpu|mig>_limit.total`, and the `initcontainer` equivalents) reported incorrect cluster totals when the check ran with `pod_collection_mode: node_kubelet`. The aggregate metrics are now computed from a dedicated instance (running on the cluster-agent or a cluster-checks runner) in the new `pod_collection_mode: cluster_aggregates_only` mode, which watches all pods directly from the API server. To enable the fix, set the `cluster_aggregates_enabled: true` instance option on the `node_kubelet` and `cluster_unassigned` instances; those instances then suppress the affected accumulators, eliminating multi-source gauge collision at ingestion. Without that option the previous (colliding) behavior is unchanged, so it must be set alongside deploying the `cluster_aggregates_only` instance. The fix preserves the per-pod metric scaling benefit of `node_kubelet` mode while restoring correct cluster aggregates.
- Fix a bug in the orchestrator explorer check that led to trying to collect Kubernetes subresources under certain custom resource API groups.

</Release>

<Release version="7.81.3" date="July 30, 2026" published="2026-07-30T15:19:04.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.81.3">
# Agent

### Prelude

Released on: 2026-07-30

- Please refer to the [7.81.3 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7813) for the list of changes on the Core Checks

### Bug Fixes

- Windows: Fixed an issue where the `DDAGENTUSER_KEEP_RIGHTS` opt-out was not preserved when the Agent was upgraded through Fleet Automation. Fleet-triggered upgrades uninstall and reinstall the Agent MSI as two separate steps, which cleared the stored opt-out before the reinstall could read it back, causing the `SeDeny*LogonRight` assignments on the Agent service account to be reapplied even when the customer had previously opted out with `DDAGENTUSER_KEEP_RIGHTS=1`. In-place MSI upgrades were not affected.
- Fixed an issue where Remote Configuration would sometimes attempt to process client requests that had already timed out.

# Datadog Cluster Agent

### Prelude

Released on: 2026-07-30 Pinned to datadog-agent v7.81.3: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7813).

</Release>

<Release version="7.81.2" date="July 22, 2026" published="2026-07-22T11:40:38.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.81.2">
# Agent

### Prelude

Released on: 2026-07-22

- Please refer to the [7.81.2 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7812) for the list of changes on the Core Checks

### Upgrade Notes

- Metrics now use the Datadog v3 intake by default for Datadog destinations.

  Metric destinations configured with non-Datadog-looking URLs, such as custom `additional_endpoints` and reverse proxies, continue to use the v2 intake by default. To enable v3 for every destination, set `use_v3_api.series.enabled: "true"`. To keep using v2 intake, set `use_v3_api.series.enabled: "false"` (global) or `use_v3_api.series.endpoints: { "<url>": "false" }` (per-endpoint).

### Bug Fixes

- On macOS, opening the Datadog Agent GUI via the fallback launch method no longer steals focus from the application the user is currently working in.

# Datadog Cluster Agent

### Prelude

Released on: 2026-07-22 Pinned to datadog-agent v7.81.2: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7812).

</Release>

<Release version="7.81.1" date="July 15, 2026" published="2026-07-15T12:42:25.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.81.1">
# Agent

### Prelude

Released on: 2026-07-15

- Please refer to the [7.81.1 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7811) for the list of changes on the Core Checks

### New Features

- Add a new reflector-based Kubernetes event collection path, enabled via `event_collection_mode: watch`.

### Enhancement Notes

- Agents are now built with Go `1.26.5`.

# Datadog Cluster Agent

### Prelude

Released on: 2026-07-15 Pinned to datadog-agent v7.81.1: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7811).

</Release>

<Release version="7.81.0" date="July 8, 2026" published="2026-07-08T12:09:40.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.81.0">
# Agent

### Prelude

Released on: 2026-07-08

- Please refer to the [7.81.0 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7810) for the list of changes on the Core Checks

Metrics now use the Datadog v3 intake by default. The v3 payload format is more compact, reducing outbound bandwidth from Agents to Datadog.

Metrics sent to Observability Pipelines Worker continue to use the v2 intake by default.

To keep using v2 intake, set `use_v3_api.series.enabled: "false"` (global) or `use_v3_api.series.endpoints: { "<url>": "false" }` (per-endpoint).

### Upgrade Notes

- The DDOT feature gate `exporter.datadogexporter.metricremappingdisabled` has been removed and replaced with `exporter.datadogexporter.DisableAllMetricRemapping`.

- Removed the `agent status py` subcommand (which wasn't officially supported)

- On Linux, the agent process manager systemd units were renamed from `datadog-agent-procmgrd.service` / `datadog-agent-procmgrd-exp.service` to `datadog-agent-procmgr.service` / `datadog-agent-procmgr-exp.service`. The `dd-procmgrd` binary and its paths are unchanged.

  On upgrade, the installer stops and removes the legacy `procmgrd`-suffixed unit files so only one process manager daemon binds the socket. Update any custom automation that referenced the old unit names.

- Upgrade OpenTelemetry Collector dependencies from v0.152.0 to v0.153.0 (core v1.58.0 to v1.59.0).

  See the full upstream changelogs: [collector-contrib v0.153.0](https://github.com/open-telemetry/opentelemetry-collector-contrib/releases/tag/v0.153.0), [collector core v0.153.0](https://github.com/open-telemetry/opentelemetry-collector/releases/tag/v0.153.0).

- Upgrade OpenTelemetry Collector dependencies from v0.153.0 to v0.154.0 (core v1.59.0 to v1.60.0).

  See the full upstream changelogs: [collector-contrib v0.154.0](https://github.com/open-telemetry/opentelemetry-collector-contrib/releases/tag/v0.154.0), [collector core v0.154.0](https://github.com/open-telemetry/opentelemetry-collector/releases/tag/v0.154.0).

### New Features

- In-place vertical scaling is enabled as the default strategy for workload autoscaling.

- New metrics for GPU memory have been added to the GPU Monitoring product:
  - `gpu.memory.utilization`: Ratio of used memory compared to total memory.

- Add passthrough entry for genresources EVP intake track.

- This change adds two new metric points for the GPU Monitoring product:  
  - `gpu.pci.link.speed.current`: Current usable bandwidth for the PCI link in bytes per second
  - `gpu.pci.link.speed.max`: Max usable bandwidth for the PCI link in bytes per second

- Add a new ReportIssue method to the Python bridge to report issues to Agent Health Platform

- APM: The trace-agent can now receive span tag equivalence and peer tag mapping updates over Remote Configuration and apply them at runtime, without an agent restart. The feature is opt-in via the new `remote_configuration.apm_semantics.enabled` setting (default `false`). Stats aggregation picks up the updated peer-tag keys on the next span processed. If the backend removes or untargets a previously-applied payload, the trace-agent reverts to the mappings it ships with. Existing deployments see no behavior change with default settings.

- APM: `remote_configuration.agent_config.enabled` is now a settable configuration entry that controls the trace-agent's Remote Configuration subscription for agent-config updates (such as runtime log-level overrides) independently from `remote_configuration.apm_sampling.enabled`. When the user has explicitly set `apm_sampling.enabled` but not `agent_config.enabled`, the trace-agent mirrors the former into the latter so existing configurations continue to behave exactly as before.

- On Linux, when the DDOT extension is installed with the Datadog Agent, DDOT is now managed by `dd-procmgrd` through `processes.d/datadog-agent-ddot.yaml` instead of relying on the legacy `datadog-agent-ddot` systemd unit. Uninstalling the extension removes that config file. To roll back to the legacy behavior manually, remove `processes.d/datadog-agent-ddot.yaml` and restart `datadog-agent`.

- Add Go stack trace aggregation to the auto multi-line log pipeline. When auto multi-line aggregation is enabled (`logs_config.auto_multi_line_detection`), multi-line Go crash dumps (`panic:`, `fatal error:`, `runtime:` errors, signal crashes, and unexpected faults) are automatically detected and combined into a single log entry using a streaming state-machine parser.

- gpu: all gpu.nvlink.\* metrics now have a nvlink\_port tag and are emitted per-port. We provide GPU-level alternatives for certain metrics such as gpu.nvlink.throughput.data.rx/tx.total

- Enable `instrumentation_crd_controller.enabled` and a new autodiscovery provider will schedule checks derived from `DatadogInstrumentation` custom resources deployed in the Kubernetes cluster.

- Parses and collects `kubernetes.pod.cpu.requests`, `kubernetes.pod.memory.requests`, `kubernetes.pod.cpu.limits`, and `kubernetes.pod.memory.limits`.

- Process Autodiscovery is now enabled by default on Linux through the `process` autoconfig feature. It can be disabled with `DD_AUTOCONFIG_EXCLUDE_FEATURES=process`.

- Register `process_manager.enabled` in the Agent configuration schema (`pkg/config/schema/core_schema.yaml`), set its default in `pkg/config/setup`, and document it in `config_template.yaml`. On Windows, this option controls whether the core Agent starts `dd-procmgr-service`. On Linux, `dd-procmgrd` is started by systemd; this setting is ignored there.

### Enhancement Notes

- Use compensated floating point summation to accurately calculate the sum and average aggregates of histograms for inputs where magnitudes significantly vary.

- Scale `.sum`, `.avg`, and `.count` aggregates by the exact `1/SampleRate` to avoid undercount of these aggregates for sample rates whose reciprocal is not an integer (e.g. `@0.21`).

- The macOS battery check now adds a `power_state:battery_critical` tag to the `system.battery.power_state` metric when the operating system reports a degraded battery.

- Update the SNMP traps database with new MIB additions, including `PANZURA-TRAP-MIB`.

- Updated the ntp check to support the default location of `systemd-timesyncd` (`/etc/systemd/timesyncd.conf`). The check now parses `NTP=` and `FallbackNTP=` keys in addition to the existing chrony/ntp.conf `server`/`pool`/`peer` directives.

- On startup the Datadog Agent will now validate its configuration against the schema and report any violations through the Agent Health pipeline.

- APM stats now mask additional metric tag values that exceed the value length or per-bucket cardinality limits.

- APM : The `enable_otlp_container_tags_v2` behavior is now enabled by default. Container tags on OTLP traces are now extracted using the infraattributes processor instead of calling the tagger directly, reducing redundant work and outgoing traffic. To opt out, set `disable_otlp_container_tags_v2` in `apm_config.features`.

- Agents are now built with Go `1.26.4`.

- CWS: Add support for monitoring the `socket` system call, enabling detection rules based on socket creation events (domain, type, protocol).

- The `comp/dataobs/queryactions` component now supports an optional `schedule` field on Data Observability monitor queries. The field accepts a standard 5-field cron expression (e.g. `"20 * * * *"` for 20 minutes past every hour) and enables wall-clock-aligned scheduling in place of the fixed `interval_seconds` cadence. When both `schedule` and `interval_seconds` are set on the same query, `schedule` takes precedence and `interval_seconds` is ignored. At least one of the two fields must be set; the agent rejects Remote Configuration payloads containing queries where neither field is provided or where the cron expression is syntactically invalid.

- Expanded the functionality of the experimental fentry-based network connection tracer. This tracer remains experimental and disabled by default.

- gpu: add new PCI link width metrics `gpu.pci.link.width.{current,max}` and add degraded PCI link metrics `gpu.pci.link.{width,speed}.degraded`.

- gpu: add `gpu.nvlink.errors.fec.{none,light,heavy}` metrics to easily group error thresholds

- The in-place vertical autoscaler throttles disruptive resizes to at most 15% of a workload's replicas per reconcile, configurable via `autoscaling.workload.in_place_vertical_scaling.disruption_tolerance_percent`.

- use the `/healthz` route to check and validate kubelet connection, instead of the deprecated `/spec` route.

- The agent automatically detects Kueue-related labels in pods and adds them as `kueue_local_queue` and `kueue_cluster_queue` tags.

- Network Config Management: Adds support for Cisco ASA firewalls by adding a new profile for these. Previously, Cisco ASA was not supported and would be unmonitored by the NCM integration.

- Extended the ntp check's `systemd-timesyncd` discovery to also read drop-in files under `/etc/systemd/timesyncd.conf.d/`, `/run/systemd/timesyncd.conf.d/`, `/usr/local/lib/systemd/timesyncd.conf.d/`, and `/usr/lib/systemd/timesyncd.conf.d/`. This covers hosts where `NTP=` is set by cloud-init or another tool that writes a drop-in instead of editing the main configuration file.

- Oracle: The Database Monitoring agent now derives blocking session and instance information from `v$lock`/`gv$lock` for sessions waiting on enqueue locks (`enq:` wait events) when the database does not auto-populate these fields. This requires granting `SELECT` on `v$lock` and `gv$lock` to the agent database user.

- Added `dockerstatsreceiver`, `kubeletstatsreceiver`, and `podmanreceiver` to the Datadog Distribution of OpenTelemetry (DDOT) Collector default component set.

- Add `datadog-private-action-runner rotate-identity` to force a new enrollment and rotate the runner's credentials. Restart the process to apply.

- Reduced memory allocation pressure in the orchestrator check by deferring deep copies and model extraction until after a cache miss is confirmed, eliminating the allocation cost for unchanged resources in steady-state clusters.

- Bumped the Security Agent policies to [v0.81.0](https://github.com/DataDog/security-agent-policies/compare/v0.80.0...v0.81.0)

- SNMP traps listener now supports a `network_devices.snmp_traps.tags` configuration option to attach a user-supplied list of tags to every forwarded trap and to every SNMP traps telemetry metric.

- Add a new `datadog-agent snmp walk --analyze` mode that converts SNMP walk output into a readable analysis report. The report summarizes matched and unmatched OIDs and includes profile context to help troubleshoot profile-to-device mismatches. Profile matching uses built-in and on-disk profiles only; profiles delivered via remote configuration (`use_remote_config_profiles`) are not loaded.

- The host Software Inventory metadata now reports the `install_paths` of each detected application, i.e. the filesystem location(s) where the software is installed, on macOS and Windows.

- Send CNM/USM data directly to Datadog from system-probe on Linux. This eliminates the need to run process-agent in certain configurations.

- Add `ebpf.core_load_success`, `ebpf.core_load_error`, `ebpf.core_remoteconfig_success`, and `ebpf.core_remoteconfig_error` to internal telemetry.

- System Probe will now download BTF (BPF Type Format) data, if needed, to support eBPF-based features. This behavior will only take effect in environments where the Linux kernel is newer than the Agent release and BTF is not available directly from the kernel.

- Windows: Added a new MSI property `DDAGENTUSER_KEEP_RIGHTS` that, when set to `1`, `true`, or `yes`, instructs the installer to skip re-applying the `SeDeny*LogonRight` assignments on the configured Agent service account (`ddagentuser` by default, or a custom account set via `DDAGENTUSER_NAME`). This lets customers preserve custom user-rights changes — for example, removing the service account from `SeDenyNetworkLogonRight` so the Agent can access network resources — across upgrades.

  `SeServiceLogonRight` is always granted regardless of this flag because the Agent service cannot start without it.

  Default behavior is unchanged: when the property is not set, the installer continues to enforce the hardened user-rights baseline.

### Deprecation Notes

- APM : The `evp_proxy_config.app_key` configuration option has been removed, along with support for the `X-Datadog-NeedsAppKey` request header in the trace-agent EVP proxy. The EVP proxy no longer attaches an Application key to forwarded requests.
- APM : The `enable_otlp_container_tags_v2` feature flag has been removed and no longer has any effect. Use `disable_otlp_container_tags_v2` to opt out of the new default behavior.

### Bug Fixes

- APM : On Windows, the .NET ETW tracer now forwards only the .NET runtime events that match its configured keywords, instead of all events delivered by the tracing session.
- APM : Prevent incoming msgpack payloads from over-allocating maps and lists.
- Fixed the Cluster Agent AppSec ingress-nginx injector so the injected init container starts on clusters that enforce `runAsNonRoot`. The init container set `runAsNonRoot: true` without an explicit `runAsUser`, and the injection image runs as root, so the kubelet rejected it with "container has runAsNonRoot and image will run as root", leaving the ingress-nginx controller pod stuck in `CreateContainerConfigError`. The init container now runs with an explicit non-root UID/GID, configurable via `admission_controller.appsec.nginx.init_run_as_user` and `admission_controller.appsec.nginx.init_run_as_group` (defaults 101/82; set a negative value to honor a custom init image's own user).
- Logs: Fix duplicate logs from containers tailed via the Docker socket when `logs_config.auto_multi_line_detection` is enabled. The auto-multiline aggregator emitted combined messages stamped with the first aggregated line's timestamp, which the Docker tailer committed as its resume offset; any reader restart then replayed lines 2..N of the group as duplicate, un-aggregated entries. The aggregator now carries the last aggregated line's timestamp through to the emitted message so the offset advances past the full group. Container logs tailed via files and the regex-driven `log_processing_rules` multi-line path were unaffected.
- Workload autoscaling: fixed a bug where custom tags added through the `ad.datadoghq.com/tags` annotation on a local-owner `DatadogPodAutoscaler` were not applied to the `datadog.cluster-agent.autoscaling.workload.*` metrics until the Cluster Agent was restarted (or the autoscaler spec was otherwise modified). The annotation is now part of the metadata fingerprint used to detect changes, so edits are picked up on the next reconcile.
- Enforce close socket after container stats read to prevent fd leaks.
- Fix the `health_platform.forwarder.interval` configuration field type from integer to string. Previously, setting an integer value (e.g. `900`) would be interpreted as nanoseconds by the agent instead of seconds, resulting in an unexpectedly short flush interval. The field now accepts duration strings such as `15m` or `5m30s`, consistent with other duration configuration fields in the agent.
- Fix Jetson check failing to parse `tegrastats` output on boards (e.g. Jetson AGX Thor) where `GR3D_FREQ` reports only per-GPC frequencies with no usage percentage (e.g. `GR3D_FREQ @[494,494,494]`). When no percentage is present, `nvidia.jetson.gpu.usage` is omitted instead of causing a parse error.
- Fixes a memory leak in the Kubernetes State Core check where Kubernetes watch connections and their cached data were not released when the check was unscheduled. In environments with frequent check rescheduling, this caused memory to grow unboundedly over time.
- Fix the kubelet check double-counting `kubernetes.cpu.usage.total`, `kubernetes.memory.usage`, `kubernetes.memory.working_set`, `kubernetes.filesystem.usage`, `kubernetes.filesystem.usage_pct`, `kubernetes.network.rx_bytes` and `kubernetes.network.tx_bytes` when `use_stats_summary_as_source` is enabled. The cAdvisor source no longer emits the metrics already produced by the kubelet `/stats/summary` endpoint, so `sum:` aggregations no longer roughly double when the option is turned on. The summary provider now also covers init and ephemeral containers, so enabling the option does not drop their CPU, memory, and filesystem metrics.
- Fix OTel Agent panic at startup when `DD_SYNC_DELAY` or `DD_SYNC_TO` is set to a bare number without a Go duration unit suffix (e.g. `30` instead of `30s`). The agent now prints a clear error message with a suggested fix instead of crashing with a stack trace.
- Fix a startup crash of the `simple-all-in-one` Docker image (used for ECS Fargate sidecar deployments) caused by a missing entrypoint wrapper for the Private Action Runner.
- Fixed an issue where the SBOM runtime usage enrichment (the "package in use" indicators such as `LastSeenRunning`) was never reported for containers managed by the kubelet. The enriched SBOM was keyed by the container image's manifest digest instead of its config digest, so it landed on a separate image entity that was never shipped to the backend.
- Fixed the SBOM runtime usage enrichment ("package in use") not being reported for some binaries on usr-merged Linux distributions (for example `mount` and `su` on Debian/Ubuntu). The package database records these under their pre-merge path such as `/bin/mount` while the kernel resolves the executed path to `/usr/bin/mount`. The resolver now normalizes both the `/bin` and `/usr/bin` layouts, so the affected packages' `HasSetSuidBit` and `LastSeenRunning` properties are populated correctly.
- Fixes a regression in serverless-init and the Lambda extension where the trace-agent's EVPProxy was disabled at startup, producing "EVPProxy is disabled: Has been disabled in config" 405 responses to clients sending LLM Observability spans (and other EVP-routed payloads). The `evp_proxy_config.*` and `ol_proxy_config.*` default registrations were only reachable under the non-serverless build, so the trace-agent's unconditional read of `evp_proxy_config.enabled` returned `false` and clobbered the package-level `true` default. The defaults have been moved into the shared APM config setup so they take effect under both the regular and `serverless` build tags.
- Fixes a segmentation fault that occurred when the Agent was shut down while `agent stream-logs` was running. The diagnostic message receiver's filter goroutine now detects a closed input channel during shutdown instead of dereferencing a nil message.
- Fix an issue on Windows where the `datadog-system-probe` service could hang during shutdown when Cloud Workload Security was enabled, resulting in delayed Agent restarts and a forced service termination. The same underlying issue could also prevent `datadog-system-probe` from initializing the DNS monitoring or Network Path features, leaving them unable to collect data.
- The `infra_mode` host tag emitted in `end_user_device` mode now uses the same key as the system CPU checks (`infra_mode:`), replacing the mismatched `infrastructure_mode:` key.
- Logs: Fix an issue where the TCP/Unix stream socket tailer would silently drop the final message of a connection when the peer closed the connection without sending a trailing newline. Forwarders that use the connect-send-close-per-message pattern (one TCP connection per log event) now have their final message emitted on end-of-stream rather than discarded.
- Fix F5 BIG-IP TMOS running-config validation for `#TMSH-VERSION` lines and configs that start with `ltm`.
- Fixed an OTLP explicit-bucket histogram bug where percentile aggregations (p50, p75, p90, p95, p99) for distribution metrics could collapse to 0 when a histogram's first non-empty bucket was `(0, B]`. This affected the default boundary set used by Micrometer's OTLP registry and many OTel SDKs (`[0, 5, 10, 25, 50, 75, 100, 250, …]`), most visibly when high-cardinality tagging and short delta intervals produced small per-bucket counts. `avg`, `sum`, `min`, `max`, and `count` aggregations were not affected.
- OTLP: The `_dd.stats_computed=false` resource attribute now overrides the `Datadog-Client-Computed-Stats: true` HTTP header. Collectors that cannot compute APM stats out-of-band can set this attribute to ensure APM metrics are computed by the Agent regardless of the header value. When the attribute is absent, the header still governs.
- Fixed an issue where the Private Action Runner (PAR) would fail to start up properly, causing it to not execute any tasks. This was caused by a race condition where the Remote Configuration notification could be fired before the PAR component had finished subscribing, causing it to miss the initial configuration.
- Fix an issue where the Private Action Runner binary was built without `zlib` and `zstd` compression support, causing invalid compressions when forwarding logs through the event-platform pipeline.
- Container image SBOMs no longer report a layer's diff\_id as its `LayerDigest`. A diff\_id is the uncompressed-content hash, while the `LayerDigest` is the compressed manifest blob digest; the two are distinct identifiers. The CRI-O `LayerDigest` is now read from the image manifest CRI-O stores on disk, and its `LayerDiffID` is taken from the image config rather than the containers-storage layer ID. Docker exposes no per-layer manifest digest and leaves `LayerDigest` empty instead of substituting the diff\_id.
- Fixed numeric list settings such as `network_config.dns_monitoring_ports` ignoring environment-variable overrides: `DD_NETWORK_CONFIG_DNS_MONITORING_PORTS` now accepts a JSON array (`"[53,5353]"`) or a space-separated list (`"53 5353"`) instead of only a single value.
- Fixed an issue on Windows where `C:\ProgramData\Datadog\application_monitoring.yaml` was not readable by IIS App Pool identities (and other non-administrator accounts), causing the .NET tracer to silently ignore fleet-managed stable configuration with an `Access is denied` error. The Fleet installer now grants `Everyone` read access on `application_monitoring.yaml` when the file is created or updated (at install time and via remote config experiments), matching the world-readable (`0644`) behavior on Linux.

### Other Notes

- The v3beta metrics series shadow sampling introduced in Agent 7.80.0 is now disabled by default.

- Internal refactoring of the health platform component: introduces an `egress` component that owns the periodic `store → intake` flush loop, and simplifies the `forwarder` component to a stateless HTTP client (`Send(ctx, *HealthReport) error`). The `SetProvider`/`IssueProvider` callback workaround for the circular dependency between the store and the forwarder has been removed.

- Internal refactoring of the health platform component: introduces a `runner` component that executes health check functions, decouples the scheduler from the store, and replaces the `SetReporter`/`SetProvider` callback pattern with direct fx dependencies.

- The health platform now uses a single shared issue registry component, eliminating duplicate registry construction and improving consistency of issue template lookups.

- The health platform store now accepts fully-built proto `Issue` objects directly via `ReportIssue`, removing its dependency on the issue template registry. Template resolution moves to the runner (for built-in health checks) and to direct callers (for AD misconfiguration and check-failure issues).

- Added Cross-Org Agent Telemetry (COAT) metrics to track whether agent services are supervised by `dd-procmgrd` or legacy supervisors (systemd on Linux, Windows Service Manager on Windows). DDOT is the first tracked service. Metrics are reported under the `procmgr` COAT profile:

  Metric Name | Type | Description |  
  --- | --- | --- |  
  `runtime.procmgr_daemon_reachable` | Gauge | Whether the agent can reach `dd-procmgrd` |  
  `runtime.procmgr_daemon_ready` | Gauge | Whether `dd-procmgrd` reports ready |  
  `runtime.procmgr_process_running` | Gauge | Whether a procmgr-managed process is running (`process` tag) |  
  `runtime.agent_service_installed` | Gauge | Whether a migratable service is installed (`service` tag) |  
  `runtime.agent_service_procmgr_configured` | Gauge | Whether a `processes.d` config exists (`service` tag) |  
  `runtime.agent_service_management_mode` | Gauge | Active supervisor for a service (`service`, `mode` tags) |

# Datadog Cluster Agent

### Prelude

Released on: 2026-07-08 Pinned to datadog-agent v7.81.0: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7810).

### New Features

- The admission controller can now automatically pick the Datadog CSI driver as the library injection mechanism for APM single step instrumentation when the `auto` injection mode is selected, the Datadog CSI driver is installed in the cluster and APM support is advertised on its annotations. Otherwise, the admission controller falls back to the init container injection mechanism. This auto-detection is disabled by default.

- The admission webhook now writes a set of APM Single Step Instrumentation (SSI) observability annotations directly on mutated pods, making the full injection outcome inspectable via `kubectl get pod -o yaml` without requiring cluster-level access.

  New annotations written by the webhook:

  - `internal.apm.datadoghq.com/injection-status`: overall outcome — `injected`, `partial`, `skipped`, or `error`.
  - `internal.apm.datadoghq.com/injected-libraries`: JSON array listing every component the webhook attempted to inject (injector + per-language libraries), each with its name, image, and individual status.
  - `internal.apm.datadoghq.com/effective-injection-mode`: the injection mode actually used (e.g. `csi`, `init_container`, `csi (auto)`), set immediately after provider selection so it is present even when injection is subsequently skipped.
  - `internal.apm.datadoghq.com/injection-error`: human-readable reason when injection was skipped or errored.
  - `internal.apm.datadoghq.com/csi-driver-status`: observed state of the Datadog CSI driver at injection time — `apm-enabled`, `apm-disabled` (driver present but APM SSI not advertised), or `not-installed`. Set independently of the configured injection mode.

  Per-library failures (unsupported language, library injection error) no longer prevent the webhook patch from being applied. The webhook now logs a warning and reflects the partial outcome in the annotations instead of discarding the entire mutation.

- Add support for `CPURequestsRemoveLimitsMemoryRequestsAndLimits` as a container `controlledValues` in `DatadogPodAutoscaler` and `DatadogPodAutoscalerClusterProfile`. When set, CPU requests are controlled and any existing CPU limits are removed, allowing containers to burst freely. Memory requests and limits are controlled as usual.

- A `DatadogInstrumentation` custom resource can now target a Kubernetes `Service` to run checks against each of its endpoints.

- Add `external_metrics_provider.autoscaler_autogen_label_selector` configuration option to the Cluster Agent. When set, only HPAs and WPAs matching the label selector trigger autogeneration of `DatadogMetric` objects. Autoscalers with explicit `datadogmetric@` references are always tracked regardless of the selector. This allows filtering out autoscalers managed by other controllers (e.g. KEDA) to avoid creating unwanted `DatadogMetric` objects.

### Enhancement Notes

- The Cluster Agent now reports its own pod name as `pod_name` in its inventory metadata payload (`datadog_cluster_agent_metadata`), providing a stable per-replica identifier for each Cluster Agent.
- Added `kubernetes_apiserver_client_qps` and `kubernetes_apiserver_client_burst` configuration options to control the rate limiter for the Cluster Agent's Kubernetes API server client. Default QPS and burst values are increased.
- Add `datadog-cluster-agent rotate-par-identity` to rotate the Private Action Runner credentials. The new identity is written to the shared Kubernetes secret. Run a Kubernetes rollout restart of the Cluster Agent deployment to apply the new identity.
- Cluster checks now keep the same check ID across Cluster Agent restarts when their configuration is unchanged.

### Bug Fixes

- Fixed the Cluster Agent's cluster check rebalancing algorithm to operate on configuration digests rather than individual instance IDs. Previously, multi-instance configurations could be incorrectly split across different runners, causing inaccurate workload estimates and suboptimal rebalancing decisions.
- Fix APM auto-injection being blocked when a container has no CPU or memory limit and requests below the minimum threshold. The Admission Controller now correctly distinguishes between "no limit set" (unlimited resources) and "low limit", preventing the request value from being incorrectly used as the effective limit.
- Fixed an issue in the algorithm used to rebalance cluster checks that could cause unnecessary check moves between runners.
- Fix nginx AppSec init container image having the controller version tag appended even when `admission_controller.appsec.nginx.init_image` is set to a fully-qualified image reference that already includes a tag.

</Release>

<Release version="7.80.4" date="July 1, 2026" published="2026-07-01T07:06:16.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.80.4">
# Agent

### Prelude

Released on: 2026-07-01

- Please refer to the [7.80.4 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7804) for the list of changes on the Core Checks

### Bug Fixes

- Add more traces during SSI installation on Linux host

# Datadog Cluster Agent

### Prelude

Released on: 2026-07-01 Pinned to datadog-agent v7.80.4: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7804).

</Release>

<Release version="7.80.3" date="June 24, 2026" published="2026-06-24T08:06:28.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.80.3">
# Agent

### Prelude

Released on: 2026-06-24

- Please refer to the [7.80.3 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7803) for the list of changes on the Core Checks

### Enhancement Notes

- Agents are now built with Go `1.25.11`.

### Bug Fixes

- Workload autoscaling: fixed a bug where, when running the Cluster Agent in high-availability mode (multiple replicas), the burstable mode of a `DatadogPodAutoscaler` could leave the CPU limit in place on a random subset of pods. The CPU-limit removal is now re-derived from the autoscaler spec in the admission controller, so every replica applies it consistently regardless of which one handles the admission request.
- Fix Private Action Runner self-enrollment failing silently on hosts with no direct internet access when a proxy is configured in `datadog.yaml`. Enrollment requests now respect the agent proxy settings (`proxy.https`, `proxy.http`, and `no_proxy`).
- Disable v3beta metrics intake shadow payloads when zlib compression is used.

# Datadog Cluster Agent

### Prelude

Released on: 2026-06-24 Pinned to datadog-agent v7.80.3: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7803).

</Release>

<Release version="7.80.2" date="June 17, 2026" published="2026-06-17T15:29:53.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.80.2">
# Agent

### Prelude

Released on: 2026-06-17

-   Please refer to the [7.80.2 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7802) for the list of changes on the Core Checks

### Enhancement Notes

-   Compliance: CIS Docker rules (`scope: docker`) are no longer evaluated on Kubernetes nodes where the kubelet's CRI runtime is not Docker (e.g. containerd, CRI-O), avoiding false positives on GKE Container-Optimized OS which ships dockerd alongside containerd. The runtime is read from the kubelet's `--container-runtime-endpoint` flag or the `containerRuntimeEndpoint` field of its `--config` YAML; if it cannot be determined the rules continue to evaluate.

### Security Notes

-   Fixed a confused-deputy vulnerability in the Cluster Agent's AppSec ingress-nginx admission mutator where the pod's `--configmap=<namespace>/<name>` argument was trusted verbatim, allowing a user with pod-create permission in one namespace to make the Cluster Agent service account create or update ConfigMaps and add labels and annotations in arbitrary namespaces. The mutator now requires the `<namespace>` portion to match the pod's own namespace (or use the `$(POD_NAMESPACE)` downward-API substitution) and skips mutation otherwise, emitting a warning event on the pod. The vulnerability affected Cluster Agent releases starting from 7.78.0.

### Bug Fixes

-   Fix an issue where container log collection could stop for an individual container without recovering and without any error in the Agent logs. When a container's log stream was idle longer than `logs_config.docker_client_read_timeout`, the read timeout could cause the underlying Docker connection to close in a way that the tailer treated as a permanent shutdown, silently stopping log collection for that container until it was recreated or the Agent was restarted. The tailer now reconnects in this case, and only stops when the Agent is intentionally shutting down. Low-volume containers (for example, services that log only periodically) were the most affected.
-   OTel Agent: Disable v3 series API shadow sampling, which is incompatible with the zlib compression the OTel Agent forces for the metrics intake.

# Datadog Cluster Agent

### Prelude

Released on: 2026-06-17 Pinned to datadog-agent v7.80.2: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7802).

### Bug Fixes

-   Fixed an issue where the admission controller connectivity probe webhook did not include the AKS selector requirements when `admission_controller.add_aks_selectors` was enabled, which could cause repeated webhook reconciliation conflicts on AKS.

</Release>

<Release version="7.80.1" date="June 12, 2026" published="2026-06-12T14:14:17.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.80.1">
# Agent

### Prelude

Released on: 2026-06-12

- Please refer to the [7.80.1 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7801) for the list of changes on the Core Checks

### Enhancement Notes

- The Agent's embedded Python has been upgraded from 3.13.13 to 3.13.14

# Datadog Cluster Agent

### Prelude

Released on: 2026-06-12 Pinned to datadog-agent v7.80.1: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7801).

</Release>

<Release version="7.80.0" date="June 11, 2026" published="2026-06-11T08:10:54.000Z" url="https://github.com/DataDog/datadog-agent/releases/tag/7.80.0">
# Agent

### Prelude

Released on: 2026-06-11

- Please refer to the [7.80.0 tag on integrations-core](https://github.com/DataDog/integrations-core/blob/master/AGENT_CHANGELOG.md#datadog-agent-version-7800) for the list of changes on the Core Checks

### Upgrade Notes

- Health Platform: the `ReportIssue` method now takes a single `IssueReport` argument instead of `(checkID, checkName string, report *IssueReport)`. The `IssueReport` struct carries three new fields — `IssueID` (unique instance id), `IssueType` (template id), and `Source` (reporting integration name) — replacing the separate `checkID` and `checkName` arguments.

  The health platform persistence file format has been bumped to version 2. Existing persistence files (`<run_path>/health-platform/issues.json`) written by a previous agent version will be detected, logged as incompatible, and discarded on startup; the agent starts with a fresh issue state. No data migration is performed.

  For integrations calling `ReportIssue`: construct an `IssueReport` with `IssueID` set to a unique instance key (e.g. `"check-execution-failure:<check-id>"`), `IssueType` set to the template identifier that was previously passed as the `IssueId` field of the proto `IssueReport`, and `Source` set to the integration name. To resolve an issue, call `ResolveIssue(issueID)` instead of passing `nil` to `ReportIssue`.

- Health Platform: the `health_platform.issues_detected` telemetry counter is now tagged with `issue_type` instead of `health_check_id`. Update any dashboards, monitors, or telemetry configuration that filtered or grouped by the `health_check_id` tag to use `issue_type` instead.

- **APM: On Linux, the trace agent process now only starts once data is sent to any of its configured listeners.**

  Previously, the trace agent started immediately on agent startup, it now starts lazily when needed, which reduces resource usage. To disable and restore the previous behavior, set `apm_config.socket_activation.enabled: false` in `datadog.yaml`, or set the environment variable `DD_APM_SOCKET_ACTIVATION_ENABLED=false`.

### New Features

- The Windows MSI installer now ships the AI usage Chrome native messaging host (`ai-prompt-logger-native-host.exe`) under `bin\agent`. The installer generates a Chrome Native Messaging Host manifest under `bin\agent\dist` and registers it machine-wide under `HKLM\SOFTWARE\Google\Chrome\NativeMessagingHosts` (including the `WOW6432Node` view for 32-bit Chrome). The host's runtime configuration is generated as `C:\ProgramData\Datadog\ai_usage_native_host.yaml` and uses the Agent's configured APM receiver port.

- Adds a new `discovery.service_map.enabled` system-probe configuration option that boots the universal service monitoring (USM) eBPF monitor in a restricted mode, capturing only the data needed to render a service dependency map (HTTP and HTTPS via TLS uprobes). Hosts running in this mode are not billed as USM customers, are not surfaced in USM dashboards, and do not produce `universal.http.*` metrics. Intended for non-APM customers as a free preview of application observability.

- Add `k8sobjectsreceiver` to the DDOT (Datadog Distribution of OpenTelemetry Collector) default manifest, enabling collection of Kubernetes object events and resource states via the OpenTelemetry Collector pipeline.

- Adds a new action `get-resource` in kubeactions.

- The fleet installer's `agent-package` OCI index now contains a FIPS-flavored sibling manifest for each platform, distinguished by the OCI `Platform.Variant` field. When `DD_FIPS_MODE=true` is set, the installer downloads the FIPS manifest; otherwise it downloads the base manifest. The package URL is unchanged in both cases.

- Add a new `nccl` core check that collects per-rank NCCL collective communication metrics from GPU training and inference workloads.

  The check listens on a Unix domain socket (default `/var/run/datadog/nccl.socket`) for JSON events emitted by the NCCL profiler plugin (`libnccl-profiler-dd.so`) running inside GPU pods. Each event is tagged with `rank`, `collective`, `n_ranks`, `kube_pod_name`, `kube_namespace`, and `kube_container_name`.

  Metrics emitted:

  - `nccl.collective.exec_time_us` — time a rank spends inside a collective operation. A rank with a significantly lower value than its peers is the straggler; ranks with higher values are waiting at the barrier.
  - `nccl.collective.algo_bandwidth_gbps` — algorithm bandwidth of the collective.
  - `nccl.collective.bus_bandwidth_gbps` — bus bandwidth normalised for the collective type.
  - `nccl.collective.msg_size_bytes` — tensor size being communicated.
  - `nccl.rank.seconds_since_last_event` — seconds since this rank last reported an event; non-zero values indicate a potential hang.

  Enable the check cluster-wide by setting `gpu.nccl.enabled: true` in the Agent configuration (or `DD_GPU_NCCL_ENABLED=true`). The socket path can be overridden via `gpu.nccl.socket_path`; the host directory mounted into training pods can be overridden via `gpu.nccl.host_socket_path`.

- Add support for the `datadog.metric.as_type` datapoint attribute on OTLP delta sum metrics. When this attribute is set to `"rate"`, the metric is sent to Datadog as a Rate (value divided by interval) instead of a Count. Accepted values are `"rate"`, `"count"`, and `"gauge"`; unknown values are logged and ignored. This allows users migrating from DogStatsD to OpenTelemetry to preserve rate-type metric behavior.

- Add `multi_secret_backends` in `datadog.yaml` so you can declare extra named secret backends (each with `type` and `config`). When no `secret_backend_type` is set, select the backend per handle using `ENC[backendID;secretKey]` (`backendID` matches a name under `multi_secret_backends`). Precedence is `secret_backend_command` (if set) over `secret_backend_type` over `multi_secret_backends`: a custom command wins over native type; when native `secret_backend_type` is set (and no custom command), every `ENC[...]` inner string is resolved only through that type and `multi_secret_backends` is not used for routing.

- Add `admission_controller.auto_instrumentation.container_registry_allow_list` configuration option (env var `DD_ADMISSION_CONTROLLER_AUTO_INSTRUMENTATION_CONTAINER_REGISTRY_ALLOW_LIST`) to restrict which container registries can be used as sources for APM library injection via Single Step Instrumentation. When set to a non-empty comma-separated list, the admission controller will skip injection for any pod whose injector image registry is not in the list, and will set the `internal.apm.datadoghq.com/injection-error` annotation with the reason. An empty list (the default) allows injection from any registry.

- Windows: `windows_certificate` check adds `certificate_store_regex`, a list of Go regular expressions matched against `HKLM` certificate store names. Patterns are matched case-insensitively. `certificate_store` and `certificate_store_regex` can be used together; at least one must be set.

### Enhancement Notes

- Process kubernetes actions asynchronously to avoid blocking the main thread.

- Add an example OpenMetrics check configuration for Agent Data Plane deployments to restore `datadog.agent.dogstatsd.*` and `datadog.agent.forwarder.transactions.*` metrics.

- Pre-register `datadog-apm-library-iis`, `datadog-apm-library-iis-rum`, and `datadog-apm-library-httpd` in the fleet installer. The packages are gated behind remote updates so they can be rolled out via remote configuration without a new installer release.

- Chunk remote workloadmeta messages in the Agent to avoid exceeding the gRPC max message size.

- APM : The Trace Agent `agent status` output now shows the UDS (Unix Domain Socket) receiver path when UDS is enabled, in addition to the existing TCP receiver address. Each per-client entry in the receiver stats section also displays the connection type (`tcp`, `uds`, or `pipe`), making it easier to distinguish traffic arriving via different transports.

- Autodiscovery template resolution failures are now logged at ERROR level instead of DEBUG, making them visible without enabling debug logging. Additionally, when the health platform is enabled, these failures are reported as AD misconfiguration health events with actionable remediation steps, providing proactive visibility when an autodiscovered check config is silently skipped due to unsupported template variables.

- When `infrastructure_mode` is `basic`, the Agent's default allowlist now includes the Directory, WMI Check, Windows Certificate, Windows Performance Counters, and Windows Registry integrations so they can run without extra `integration.additional` configuration on Windows-oriented deployments.

- Agents are now built with Go `1.25.10`.

- On Windows, network connections collected by Cloud Network Monitoring are now tagged with `interface_name` and `interface_type`.

- The Agent now streams Kubernetes metadata from the Cluster Agent by default, instead of polling for it periodically. This propagates tags derived from Kubernetes metadata (like `kube_service`) with less delay. This behavior is controlled by the `kubernetes_metadata_streaming` setting.

- `agent diagnose` now renders the check name as a prefix for all checks under the `check-datadog` suite. The JSON output gains a `check_name` field for the same purpose.

- The `--include` and `--exclude` flags of `agent diagnose` now match against the suite name, the owning check name, and the diagnosis category. For example, `agent diagnose --include postgres` now filters individual diagnoses across all suites instead of only matching suite names.

- DogStatsD timing metrics (`t` type) now include an explicit unit value (`millisecond`) in the metric payload sent to Datadog, allowing the Datadog UI to display the correct unit automatically.

- Dynamic Instrumentation now supports compound conditions using `&&`, `||`, and `!`.

- When `infrastructure_mode` is set to `none`, ECS task metadata collection is now disabled by default (see `ecs_task_collection_enabled`). Set `DD_ECS_TASK_COLLECTION_ENABLED` to `true` to override.

- When `infrastructure_mode` is set to `end_user_device`, the Agent now attaches additional host tags to identify the device: `infrastructure_mode:end_user_device`, `os_name`, `os_version`, `cpu_model`, `total_memory_gb`, and `device_model`. Hardware and OS tags are collected on macOS and Windows only.

- Add new `ad_tag_completeness_max_wait` configuration option. When set, autodiscovery waits up to that many seconds for an entity's tags to be complete before scheduling checks for it. This avoids checks running briefly with incomplete tags. It's disabled by default.

- Add `logs_config.use_container_timestamp` to optionally use the `time` field from container log files as the log timestamp instead of ingestion time, preserving container-provided per-line timestamps.

- Logs Agent: `logs_config.tag_multi_line_logs` and `logs_config.tag_truncated_logs` now default to `true` so file logs are tagged by default when they were aggregated as multiline logs or truncated by the Agent.

- Use native API requestWhenInUseAuthorization() to manage location permission prompt on MacOS.

- The OTel Agent standalone mode now automatically disables IPC with a core Datadog Agent. When `DD_OTEL_STANDALONE` is enabled, `DD_CMD_PORT` is forced to `-1`, so users no longer need to set it manually when running DDOT without a core Agent.

- The `service.instance.id` OpenTelemetry resource attribute is now mapped to the `service.instance.id` Datadog metric tag when converting OTLP metrics. This attribute is required for OTel traffic metrics in Datadog Fleet Automation.

- Private Action Runner: When `private_action_runner.api_key_only_enrollment` is enabled, the agent now enrolls via the new API-key-only OPMS endpoint (`/api/unstable/on_prem_runners/api_key_only`). This allows runners to self-enroll using only a scoped API key, without requiring an application key.

- The Private Action Runner now honors the `X-Retry-After-Ms` response header returned by the Datadog backend on workflow task dequeue and health check requests.

- The Private Action Runner now retries self-enrollment and auto-connection creation requests when the Datadog API returns a transient `5xx` response.

- Parse the ECS `/tasks` host metadata endpoint on Managed Instances and populate `DaemonName` for daemon-scheduled tasks.

- The default for `logs_config.file_scan_period` is now **1** second instead of 10, so the Agent discovers new and rotated log files on disk more quickly. Set `logs_config.file_scan_period` explicitly if you need a slower scan to reduce filesystem load (for example on network file systems).

- Bumped the Security Agent policies to [v0.80.0](https://github.com/DataDog/security-agent-policies/compare/v0.79.0...v0.80.0)

- Reduce the payload size of SNMP device metrics by letting the Datadog backend enrich device tags (such as `snmp_device`, `device_ip`, and `device_id`) from device metadata instead of attaching them to every metric. Existing queries and monitors continue to work, and no action is required. This only applies when `collect_device_metadata` is enabled (the default).

- SNMP network device metadata: when a profile lists multiple scalar `symbols` for the same metadata field (for example `serial_number`), the check now skips values that resolve to an empty string (after trimming whitespace) and continues to the next symbol, matching the intended fallback order when an OID exists but carries no usable serial.

- Upgrade OpenTelemetry Collector dependencies from v0.150.0 to v0.151.0 (core v1.56.0 to v1.57.0).

  Notable upstream changes:

  - Removed stable feature gates that are no longer needed: `connector.datadogconnector.NativeIngest`, `exporter.datadogexporter.UseLogsAgentExporter`, and `exporter.datadogexporter.metricexportnativeclient`.
  - Several collector-contrib components have been renamed with deprecated aliases (`spanmetrics` to `span_metrics`, `hostmetrics` to `host_metrics`, `fluentforward` to `fluent_forward`). The old names continue to work but will be removed in a future release.

  See the full upstream changelogs: [collector-contrib v0.151.0](https://github.com/open-telemetry/opentelemetry-collector-contrib/releases/tag/v0.151.0), [collector core v0.151.0](https://github.com/open-telemetry/opentelemetry-collector/releases/tag/v0.151.0).

- Upgrade OpenTelemetry Collector dependencies from v0.151.0 to v0.152.0 (core v1.57.0 to v1.58.0).

  See the full upstream changelogs: [collector-contrib v0.152.0](https://github.com/open-telemetry/opentelemetry-collector-contrib/releases/tag/v0.152.0), [collector core v0.152.0](https://github.com/open-telemetry/opentelemetry-collector/releases/tag/v0.152.0).

- A small sample of series metric flushes (0.1% by default) is now additionally sent to a v3beta metrics intake endpoint to validate the upcoming v3 metrics protocol. Shadow traffic is only sent for agents configured against the `datadoghq.com` (US1) site. To opt out, set `serializer_experimental_use_v3_api.series.shadow_sample_rate` to `0`.

- The OTel Agent now logs a warning and displays it in `agent status` when the `hostmetrics` receiver is configured while running in connected mode (`DD_OTEL_STANDALONE=false`). In connected mode the core Datadog Agent already collects host metrics, so enabling the `hostmetrics` receiver can lead to duplicate or conflicting metric names. To suppress the warning, either remove the `hostmetrics` receiver or switch to standalone mode (`DD_OTEL_STANDALONE=true`).

- Add six opt-in tag flags to the Windows Certificate Store integration: `certificate_template_tag`, `enhanced_key_usage_tag`, `friendly_name_tag`, `subject_alternative_names_tag`, `issuer_tag`, and `signature_algorithm_tag`. When enabled, each certificate's metrics and service checks are tagged with the corresponding X.509 or Windows certificate property. All flags default to `false`.

### Deprecation Notes

- APM: Restored the deprecated `DD_APM_SPAN_DERIVED_PRIMARY_TAGS` configuration option, but only in serverless contexts: the Datadog Azure App Services extension (`DD_AZURE_APP_SERVICES=1`) and `serverless-init` (Cloud Run, Container Apps, Cloud Run Functions). In all other deployments the option is silently ignored. Tracers should populate `additional_metric_tags` instead; do not use `DD_APM_SPAN_DERIVED_PRIMARY_TAGS` in new deployments.

### Security Notes

- Bumped pip to 26.1.1 in the embedded Python distribution to address CVE-2026-6357.
- Updated the Windows 1809 / LTSC 2019 Agent container base images from the deprecated `mcr.microsoft.com/powershell:*-1809` images to `mcr.microsoft.com/dotnet/sdk:9.0-nanoserver-1809` (nanoserver) and `mcr.microsoft.com/dotnet/sdk:9.0-windowsservercore-ltsc2019` (servercore). The previous PowerShell base images were unmaintained and still shipped PowerShell 7.1.0, which is affected by CVE-2022-26788.

### Bug Fixes

- The debugger proxy no longer forwards Exception Replay and Live Debugger logs when `logs_enabled` is `false`. This can be overridden using the new `apm_config.debugger_logs_enabled_override` setting (environment variable `DD_APM_DEBUGGER_LOGS_ENABLED_OVERRIDE`), which enables Exception Replay and Live Debugger when `logs_enabled` is `false`.
- APM : Fixed trace span obfuscation for OpenSearch request bodies when Elasticsearch JSON obfuscation is also enabled. Spans that only included the `opensearch.body` tag (and not `elasticsearch.body`) were previously left unobfuscated in that configuration.
- APM : Fix SQL obfuscation error when a query uses PostgreSQL array slice syntax with bind parameters (e.g. `arr[$1:]` or `arr[$1:$2]`). The tokenizer was incorrectly treating the `:` range separator as the start of a named bind variable, causing obfuscation to fail with a `LexError`.
- APM OTLP: Preserve gRPC status codes on trace metrics computed by DDOT and the OpenTelemetry Collector Datadog connector. This includes explicit gRPC status attributes such as `rpc.grpc.status_code` and the newer OpenTelemetry semantic convention `rpc.response.status_code` when `rpc.system.name` is `grpc`.
- APM : Enforce body-size limits on trace-agent proxy endpoints (DogStatsD, pipeline stats, OpenLineage, Debugger, SymDB). All endpoints are capped at `apm_config.max_request_bytes` (default 25 MB). The profiling proxy uses a separate limit configurable via `apm_config.profiling_max_request_bytes` (default 50 MB, env `DD_APM_PROFILING_MAX_REQUEST_BYTES`). The `Traces` and `ClientStatsPayload` msgpack decoders now reject payloads declaring more than 500,000 elements in a single array.
- APM : Fix an issue where converting traces to the v1 format did not prefer the root span's sampling priority when multiple `_sampling_priority_v1` values were present on spans in the same trace.
- Logs collected with automatic multiline detection now fall back to individual events when combining lines would exceed `logs_config.max_message_size_bytes`. Oversized single log lines continue to use the normal truncation path, and multiline logs that fit within the limit are still aggregated.
- \[DBM\] Bump `go-sqllexer` to v0.2.2 to fix the following bugs:
  - Obfuscate `EXTRACT` field keywords (e.g. `epoch`, `year`) so that queries from `pg_stat_activity` and `pg_stat_statements` converge on the same DBM signature.
  - Fix handling of PostgreSQL `VACUUM` commands so they are correctly extracted into statement metadata.
  - Fix lexer handling of multiline comments immediately following keywords.
- Fixed HTTP flows being incorrectly dropped when a request body arrives before the response. Fixed pending HTTP transactions not being finalized when a connection is closed by a bare FIN or RST. Fixed a bug check (BSOD) caused by mismatched `maxRequestFragment` values during HTTP initialization.
- Fixed the container ID to PID mapping for processes running in a sub-cgroup of their container's cgroup (for example, CrowdStrike Falcon's `sensor.falcon` scope nested under a container scope).
- Fix the `connection_reset_interval` setting not being applied to additional log endpoints and HTTP MRF endpoints. Previously, only the main log endpoint would periodically reset its connection, which could cause additional endpoints to send logs to stale destinations after a DNS failover. Additional endpoints now inherit the global `logs_config.connection_reset_interval` value by default, and can also be overridden per-endpoint in the `additional_endpoints` configuration.
- Fix a panic in the system-probe network tracer caused by concurrent access to the gateway lookup subnet cache. The cache now uses a thread-safe LRU implementation.
- Fixed the health platform forwarder using the wrong intake endpoint. Agent health reports are now sent to `agenthealth-intake.{site}` instead of `event-platform-intake.{site}`, which is only configured for the `logs` and `processesraw` tracks. This caused org\_id to be missing from all agent health recommendations.
- Fix a class of IPv6 `host:port` formatting bugs found in multiple call sites across the Agent. IPv6 literals in host configuration values were not bracketed when used to build URLs and dial addresses, producing strings like `http://fd38::1:5005` instead of `http://[fd38::1]:5005` and causing `too many colons in address` errors at runtime.
- Fixes the journald log tailer skipping the first journal entry when `start_position` is set to `beginning` or `forceBeginning`.
- Fix chassis type detection for Mac mini and Mac Pro hosts, which were previously reported as `Other`.
- Fix the macOS battery check detecting battery on Mac minis.
- Logs: Fixed a bug where the MultiLineParser did not mark truncation when reassembled log lines exceeded the 900KB size cap. Oversized lines are now properly flagged with `IsTruncated` so that downstream handlers can apply truncation markers and increment telemetry.
- Fix spurious "Unknown environment variable" warnings for `DD_SYNC_DELAY`, `DD_SYNC_TO`, and `DD_CORE_CONFIG` when running the OTel Agent.
- Fix a panic in the OTLP metrics pipeline when a sender submits a histogram with more `BucketCounts` entries than `ExplicitBounds` allows (violating the `counts == bounds + 1` OpenTelemetry specification invariant). Such data points are now rejected with an error instead of crashing the agent.
- Fix a C-memory leak in the logs batch sender where `resetBatch()` replaced the zstd `StreamCompressor` without closing the previous one.
- The `datadog-installer.exe` install script now adds `datadog.yaml.example` template comments to the config files on fresh installs.
- Fixed the gohai resource check silently dropping processes whose UID does not exist in the host's `/etc/passwd`. This commonly affects containerized processes running as UIDs created inside container images. The "Processes memory usage" widget on the host infrastructure page now correctly includes these processes by falling back to the numeric UID string when username lookup fails.
- gpu: fix an issue where some GPM metrics (gpu.gr\_engine\_active, gpu.sm\_utilization, gpu.sm\_occupancy, gpu.integer\_active, gpu.fp16\_active, gpu.fp32\_active, gpu.fp64\_active, gpu.tensor\_active) were only emitted correctly one out of eight times on average.
- NDM SNMP: fix IP metadata fields rendered as `<nil>` for OIDs declared with SYNTAX `IpAddress` (e.g. Cisco IPsec tunnel local/remote outside IPs, CDP remote addresses). gosnmp decodes these values as Go strings, but the metadata store previously only handled the raw-bytes path.
- Fixes an issue with the NetFlow collector where certain packets would be dropped, producing error logs and lost data. The issue was caused when a packet had trailing padding which was not properly handled.
- Fixed a bug where empty log sources would get orphaned due to empty `serviceID` due the agent attempting to collect logs from short lived containers that exit quickly.
- Fixed a regression introduced around Agent 7.40 where non-template check configurations containing unresolvable `ENC[...]` secrets were still scheduled with raw secret handles in their config. Checks are now correctly dropped when all instances fail secret decryption. When only some instances fail, the surviving instances are scheduled and the failing instances are dropped, preserving the pre-regression per-instance behavior.
- SNMP: Detect GetBulk response truncation (fewer varbinds than requested OIDs) and automatically reduce the batch size, preventing silent metric loss on devices that truncate large SNMP responses.

### Other Notes

- Add handling for dbm-column-statistics events in the event platform forwarder. These events are used by Database Monitoring integrations to report column statistics from database catalogs.
- Add metrics origins for Cisco SD-WAN and Versa integrations.
- Add metrics origins for HPE Aruba EdgeConnect and NiFi.
- Removed support for using the OpenTelemetry components contained in this repo from an external collector, using OCB. The equivalent components in the opentelemetry-collector-contrib repository are designed exactly for this use case and should be used instead: - [datadog exporter](https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/exporter/datadogexporter) - [datadog extension](https://github.com/open-telemetry/opentelemetry-collector-contrib/tree/main/extension/datadogextension)

# Datadog Cluster Agent

### Prelude

Released on: 2026-06-11 Pinned to datadog-agent v7.80.0: [CHANGELOG](https://github.com/DataDog/datadog-agent/blob/main/CHANGELOG.rst#7800).

### Upgrade Notes

- Updated the bundled kube-state-metrics library from v2.13 to v2.18. The kube-state-metrics metric allow/deny list now uses ECMAScript regular expression syntax instead of Go `regexp` syntax. Most patterns are compatible, but users relying on Go-specific regex features (e.g. `(?s)` flag) in `metric_allowlist` or `metric_denylist` should update their patterns.

### New Features

- The Cluster Agent admission controller now reports connectivity probe failures to the Datadog Health Platform. When the admission webhook becomes unreachable, an `admission-controller-connectivity-failure` health issue is raised with severity `high` and category `availability`, including remediation steps. The issue is automatically resolved when connectivity is restored.
- Add a Prometheus HTTP Service Discovery (HTTP SD) provider for the Cluster Agent. The provider polls Prometheus-compatible HTTP SD endpoints and generates check configurations for each discovered target. Configure endpoints under `prometheus_http_sd.configs`, each providing its own `url` and `check_template`.
- Autoscaling profiles (`DatadogPodAutoscalerClusterProfile`) now support Argo Rollouts as a target workload type. The Cluster Agent automatically detects whether the Argo Rollouts CRD is installed at startup and, if present, watches Rollout resources for profile labels alongside Deployments and StatefulSets.
- The `kubernetes_state` core check now collects both `endpoints` and `endpointslices` resources by default, and emits new `kubernetes_state.endpointslice.address_available` and `kubernetes_state.endpointslice.address_not_ready` metrics for Kubernetes EndpointSlice objects, mirroring the existing `kubernetes_state.endpoint.address_available` and `kubernetes_state.endpoint.address_not_ready` metrics.

### Enhancement Notes

- The orchestrator check now collects force-deleted pods by default. The `orchestrator_explorer.terminated_pods_improved.enabled` option will be removed in a future release.

</Release>

<Pagination cursor="2026-06-11T08:10:54.000Z|2026-06-11T10:05:33.063Z|rel_5jN5pbubTyyeSqoCnjO0s" next="https://releases.sh/datadog/datadog-agent-releases.md?cursor=2026-06-11T08%3A10%3A54.000Z%7C2026-06-11T10%3A05%3A33.063Z%7Crel_5jN5pbubTyyeSqoCnjO0s&limit=20" />
