Tips and tricks for managing Application Insights ingestion volume
Several Neos clusters can share the same Log Analytics workspace behind Application Insights, and that workspace often has a daily ingestion cap. When the cap is reached, no cluster sharing the workspace is observable until the next reset, whether or not it caused the spike. This article collects tips and tricks to keep a cluster's telemetry useful while staying within a shared or capped budget.
Measure before you cut
Application Insights bills ingested data by table (AppMetrics, AppDependencies, AppTraces, AppRequests, AppExceptions, ...) and exposes the billed size of every row through the _BilledSize column. Before changing any configuration, query the workspace to find out which table and which resource actually drive the volume.
// Billed volume by table over the last 7 days
AppDependencies
| where TimeGenerated > ago(7d)
| extend ResourceName = tostring(split(_ResourceId, "/")[-1])
| summarize TotalMB = sum(_BilledSize) / 1e6 by ResourceName
| order by TotalMB desc
Tip
Run the query above once per table (AppMetrics, AppTraces, AppRequests, ...) to see which resource and which signal type dominate. A single high-frequency, low-value span or a single chatty dependency can outweigh everything else combined.
Note
Prefer a direct query against the Log Analytics workspace API over the Azure CLI's az monitor app-insights query --app <name>: when several Application Insights components share the same workspace, that command can resolve to the wrong component and return misleading or empty results. Query the workspace itself and filter on _ResourceId instead.
Stop sending what another tool already collects
Application metrics exported through OpenTelemetry duplicate what a Prometheus scrape endpoint already collects on most Neos deployments. The chart-generated collector configuration already reflects this: its metrics pipeline never carries an Application Insights exporter, so application and Dapr sidecar metrics stay in Prometheus by default, at no cost against the ingestion cap.
If you supply a custom collector configuration (observability.collector.config.inline or existingConfigMap), apply the same principle yourself: omit an Application Insights exporter from your metrics pipeline, or omit the pipeline entirely. See Send telemetry to Application Insights through the collector for how to structure service.pipelines in that case.
Separate what must be visible from what should be on demand
Not every signal deserves the same retention. A practical split:
- Always visible: warnings, errors, and the structuring calls of a business flow (a service invocation, a publish/subscribe message, a state store read or write that is part of the request path). This is what a first responder needs to understand that something went wrong and where.
- On demand only: detailed informational logs, full request/response payloads, and anything only useful once an issue has already been reproduced and more detail is needed.
Neos backend containers accept a logLevel setting (Warning, Error, Information, ...) per component and per business cluster in the Helm chart's values file. Set it explicitly to Warning in each environment's values file and raise it to Information only on the cluster or component under investigation - do not rely on the chart's own built-in default, which can differ between Neos versions.
# helm-<environment>.yaml
gateway:
logLevel: Warning
clusters:
- name: <ClusterName>
logLevel: Warning
Important
This setting has no hot-reload path: raising or lowering it requires updating the Helm values and rolling out the affected pods. Plan for a short redeploy when you need more detail during an investigation, and revert it once you are done.
For the full list of logLevel properties and their defaults, see Helm configuration reference.
Filter pure noise at the collector, not in the application
Some spans carry no diagnostic value regardless of the environment: Kubernetes liveness and readiness probes, a sidecar's own health-check polling, or a background poll that only matters when it fails (and a failure already surfaces elsewhere, for example as an application exception). Dropping these at the OpenTelemetry Collector, with a filter processor, keeps the application code free of telemetry concerns and applies the same rule to every signal source hitting the collector.
processors:
filter/noise:
error_mode: ignore
traces:
span:
- 'IsMatch(name, "(?i)my-noisy-polling-pattern")'
Note
The chart-generated collector configuration already ships a filter/noise processor with a starter set of rules (SignalR internals and Dapr TaskHub sidecar polling). See Default trace filters for the exact list, and extend it there instead of duplicating the processor.
Caution
A span name pattern matched with strict equality (name == "...") only matches an exact string. If the emitting code prefixes the name with an icon, a category tag, or any other decoration, the rule silently never matches. Prefer IsMatch(name, "(?i)...") for a case-insensitive, prefix-tolerant match, and verify with a query against the workspace that the volume you expected to filter actually disappears.
Do not filter a low-volume signal just because it looks noisy
A signal being frequent does not automatically make it expensive, and a signal being rare does not make it safe to drop. Database dependency calls, for example, are usually a small fraction of the total ingested volume compared to metrics or informational traces, but they carry a lot of diagnostic value when investigating a slow request or a data issue. Decide by measured volume and by diagnostic value, not by intuition about what "must" be noisy.
Leave a margin, not just a break-even
Aim for a projected daily volume with headroom below the cap, not a projection that exactly matches it. Deployment events, batch jobs, and incident investigations all produce short bursts of extra telemetry; a budget with no margin turns any of them into a full-day outage of observability for every cluster sharing the workspace.