08/11/2026 | Press release | Distributed by Public on 08/11/2026 06:58
Dynatrace simplifies Databricks observability by unifying cost, performance, job health, and security insights in one place-enabling faster, data-driven decisions for FinOps, DevOps, and platform teams.
Databricks has become a critical part of modern cloud operations. It runs scheduled jobs, interactive workloads, SQL warehouses, and increasingly, AI services. But when telemetry is split across logs, system tables, dashboards, and product views, even basic questions can take too long to answer. Why did this job fail? What caused this cost spike? Is a serving endpoint under strain, or just busy?
The Databricks Workspace extension for Dynatrace can bring those signals together, helping teams troubleshoot faster, make better capacity and cost decisions, and understand how their Databricks environment is behaving as it grows more complex. It pulls telemetry from across the Databricks platform into a single, correlated view-enriched with context and queryable. This helps give teams the context to investigate issues more efficiently, make more confident resource decisions, and strengthen governance with real data.
In this post, we'll look at how the extension helps with job and Spark reliability, cost and capacity optimization, AI-serving observability, and audit telemetry for governance and security.
For operations teams, one of the most common questions is simple: which of my runs are failing, and why? Without correlated telemetry, the answer requires checking job run logs, pivoting to Spark UI for stage output, and cross-referencing cluster metrics separately - all before any real analysis begins.
The Databricks Workspace extension for Dynatrace helps answer that by combining:
Job duration and success rate trends over time, broken down by individual job - useful for spotting regressions against a baseline.
The duration breakdown is particularly useful: when a job slows down, you can see immediately whether the delay is in the queue before execution starts, or in a specific stage mid-run. The distributed traces then let you drill into the exact task that regressed, rather than re-running the whole job to reproduce.
Example use cases:
With this context, teams can move beyond broad retries and multi-tool investigation towards more evidence-based root cause analysis.
Databricks billing data alone tells you how much you spent but it doesn't tell you what you spent it on. The extension ingests billing system table data and enriches it with workload metadata, so when a cost spike appears, you can trace it to the specific job, cluster, or SQL warehouse responsible.
Spend can be broken down by:
On the capacity side, compute node timeline metrics (CPU, memory, swap, network) support right-sizing decisions with real usage data rather than assumptions.
Automated right-sizing signals based on p50/p90 CPU and memory utilization, with explicit recommendations per cluster role.The right-sizing table in the Cluster Resource Utilization dashboard makes this concrete: Clusters flagged as over-provisioned come with a specific recommendation (reduce CPU, downsize both CPU and memory) derived from actual utilization percentiles.
Example use cases:
This gives teams the workload context they can use to pursue more targeted cost optimizations instead of relying on blanket cuts.
As model-serving endpoints move into production, teams face questions that go beyond uptime: who is consuming capacity, how much, and at what cost? Without usage visibility, token consumption and serving costs accumulate without accountability.
The extension captures model serving endpoint usage and health signals, including:
The requester breakdown is especially useful for governance: you can see which users or services are driving the most traffic and token consumption, giving you the data you need to set quotas or support internal chargebacks before costs become a conversation.
Example use cases:
These signals can help teams make more informed decisions as they scale AI services, with visibility into reliability, performance, and cost.
Audit data answers the questions that come after something goes wrong: who changed what, when it happened, and what the platform did in response. The challenge is usually speed - getting from "something looks off" to a concrete answer before the incident grows.
The extension surfaces Databricks audit events with enough context to investigate quickly: the service and action involved, the user or identity behind it, and the resulting status or error. Events can be filtered and sliced by IP, kind (accounts, clusters, Unity Catalog, etc.), status code, and more - turning raw audit data into a navigable investigation surface.
Audit events broken down by IP, event kind, and more - with a filterable event table showing individual actions, their outcomes, and the identities behind them.Example use cases:
This can help turn audit logs into a more usable investigation surface, helping teams explore events in context instead of treating logs as a static archive.
The extension includes six bundled dashboards to accelerate time to value:
These dashboards provide immediate visibility so you can drill down into the questions that matter most for your environment.
Dynatrace helps teams operationalize Databricks telemetry with unified context across logs, metrics, and traces so they can move from insight to action faster. It helps transform Databricks from being yet another silo to monitor into part of a more proactive, resilient cloud operations model.
If your goals include higher reliability, better cost efficiency, stronger governance, and scalable AI operations, the Databricks Workspace extension can help you get there.