Skip to content

Connected Diagnostics

Connected Diagnostics starts from one or more resources, follows approved relations to build a bounded scope, and collects diagnostic findings from multiple resources in one run. It presents operational risk, coverage, evidence confidence, impact, and urgency as separate dimensions.

This is a read-oriented analysis that helps an operator review risk and recommended actions. It does not change resource configuration or alert state, and a suspected cause is not a confirmed root cause.

Before You Start

  • A valid Enterprise license is required. There is no Connected Diagnostics-specific entitlement.
  • The current account must have read access to every starting resource and every resource included in the resolved scope.
  • Workspace access never replaces resource-read access.
  • Resource plugins must expose standard diagnostics/summary or diagnostics/history findings for the run to collect useful results.
  • Traversal uses approved depends_on and hosted_on relations plus valid metric and log-source links. Suggested and rejected relations are not used.

Open Connected Diagnostics

Connected Diagnostics is available from:

  • the Resources index;
  • resource Monitor, Manage, and Edit pages;
  • a resource Relations page; and
  • an Analysis Workspace.

Opening it from a resource page preselects that resource. Opening it from an Analysis Workspace uses readable resources from ready panels on the active page and maps the current time range into the evidence lookback.

From the Resources index, or whenever the starting scope needs to change, select Add resources. Search by resource name, ID, or plugin instead of entering an ID manually. Only active resources readable by the current account appear.

Configure the Diagnostic Scope

The page automatically refreshes the scope preview whenever starting resources or analysis settings change. There is no separate preview button. The preview shows included resources, relation count, and estimated execution time.

Traversal Modes

ModeMeaning
DependenciesFollow resources that a starting resource depends on. Use it to inspect downstream services or data stores.
ImpactFollow resources that depend on a starting resource. Use it to inspect upstream consumers that could be affected.
NeighborhoodFollow nearby approved relations in either direction, including valid metric and log-source links.
Workspace onlyKeep the scope within starting resources selected from the workspace.

Depth controls how many relation hops are followed from a starting resource. The default is 1 and the maximum is 3. A larger depth can increase both scope and execution time.

Lookback controls the recent diagnostic history and evidence window. The UI offers 1 hour, 6 hours, 24 hours, and 7 days. The complete server-accepted range is 15 minutes through 7 days.

Resource limit is the maximum number of resources in one run. The default is 25 and the maximum is 100. The server uses at most 500 relations and always processes resource and relation identities deterministically without duplicates.

Inactive or unreadable resources are removed without disclosing their name, ID, or excluded count. When the preview is smaller than expected, review both relation status and resource-read access.

Run and Track a Diagnosis

  1. Review the starting resources and automatic scope preview.
  2. Set traversal mode, depth, lookback, and resource limit.
  3. Select Run connected diagnosis.
  4. Follow the progress bar and completed, failed, and unsupported counts.
  5. Select Cancel while the run is active when necessary.

Runs use these states:

StateMeaning
QueuedThe run is waiting for an execution slot.
RunningScope resolution or per-resource collection is in progress.
CompletedEvery supported target completed collection.
Partially completedSome targets succeeded while others failed or were unsupported.
FailedThe run could not produce a usable diagnostic result.
CancelledThe operator cancelled the run.

A failure on one target does not discard results already collected from other targets. Up to four runs execute globally, with two active runs per owner and four resource collectors per run. Collection is bounded to six seconds per resource and 45 seconds for the complete run. Retry later when the execution-capacity message appears.

Interpret the Scores

Every dimension ranges from 0 to 100. The direction differs by dimension, so interpret a value together with its card title.

DimensionMeaning of a higher value
Operational riskHigher-priority diagnostic risks need attention.
CoverageMore applicable findings were collected. Higher is better.
Evidence confidenceCollection is broader and evidence is fresher. Higher is better.
ImpactA broader scope may be affected through approved topology.
UrgencyCombined risk, impact, and persistence or recurrence justify higher response priority.

The shared numeric bands are None at 0, Low above 0 and below 30, Moderate from 30 to below 55, High from 55 to below 80, and Critical at 80 or higher. A High or Critical band on Coverage or Evidence confidence is not a bad state. For these two cards, a higher value means more complete evidence; read the numeric value and reason text first.

Operational Risk

  • Critical, warning, and informational findings have base contributions of 85, 55, and 20.
  • A plugin-provided persistence or recurrence count can add a bounded amount to one finding contribution.
  • The highest contribution is kept in full. Only eight percent of each of the next three contributions is added.
  • One critical finding therefore produces a risk of at least 85 and enters the Critical band.
  • Unknown and unavailable findings never lower an established risk; they lower Coverage and Evidence confidence instead.
  • Not applicable findings are excluded from both risk and coverage denominators.

For example, a critical contribution of 85 followed by a warning contribution of 55 produces 85 + 55 × 0.08 = 89.4. This prevents a critical finding from being hidden by averaging many lower-risk findings.

Coverage and Evidence Confidence

Coverage is the percentage of applicable findings for which a collected status exists. Not applicable findings are excluded, while unknown and unavailable findings are not counted as collected.

Evidence confidence combines coverage with freshness and can never exceed coverage. For example, if 6 of 8 applicable findings are collected and 5 of the 8 are fresh, Coverage is 75 and Evidence confidence is 63.75. When risk appears low alongside poor coverage, restore the collection path before treating the result as healthy.

Impact and Urgency

Impact uses approved relation paths, maximum path depth, and resources marked with high criticality. Suggested and rejected relations do not contribute.

Urgency combines 65 percent Operational risk, 25 percent Impact, and 10 percent bounded persistence or recurrence. Two runs with the same risk can therefore have different urgency when one affects more critical resources or includes stronger recurrence evidence.

Review Findings and Collection Gaps

Filter the finding list by category and status. Select a finding to review:

  • interpretation;
  • impact;
  • recommended action;
  • navigation to the source resource; and
  • when launched from a workspace, an action to pin a note to the active page.

The main finding states are:

StateInterpretation
CriticalA high operational risk needs immediate review.
WarningReview a potential degradation, configuration, capacity, or reliability risk.
InfoA low-risk fact may still help the operational decision.
OKThe check verified its healthy condition.
UnknownRequired evidence is insufficient for a decision.
UnavailableA collector or source-response problem prevented evaluation.
Not applicableThe check does not apply to this resource.

Collection gaps list every resource that did not complete collection.

  • Unsupported means that the plugin or resource does not expose the required standard diagnostic route.
  • Unavailable means that the resource, configuration, or diagnostic query could not be used.
  • Invalid response means that the plugin result did not satisfy the canonical diagnostic contract.

No collection gaps means that every collector responded. It does not mean there is no risk; review findings and scores together.

Pin a Finding to a Workspace

When Connected Diagnostics is opened from an Analysis Workspace, the finding detail includes Pin to workspace. The action saves the finding title and state as a note on the active page.

The note is not a stored copy of the complete run or raw evidence. Add the resource, time range, and intended response to the note when that context must remain available for later investigation.

Result Retention

Connected Diagnostics results are transient.

  • Once a Completed, Partially completed, or Failed result is displayed, the UI immediately consumes its server-side execution buffer.
  • The result remains visible while the current page stays open, but it cannot be restored after refresh or navigation.
  • There is no result history, report, run comparison, retention control, or daily or weekly schedule.
  • If a browser closes, a network disconnects, or an interrupted client leaves data unconsumed, the HA leader cleanup removes it after one hour.
  • Audit records retain create, run, cancel, and result-consumption events, but not the finding payload.

Pin decisions that must be revisited to an Analysis Workspace, or record them in an approved operational system. Do not use browser refresh as a retention mechanism.

Use Connected Diagnostics through MCP

When both the MCP gateway and Connected Diagnostics are enabled, the EE core catalog exposes:

  1. preview_connected_diagnostics to resolve the readable scope before execution;
  2. start_connected_diagnostics to start an asynchronous run and return a run_id;
  3. get_connected_diagnostics_result to read progress or a terminal result; and
  4. cancel_connected_diagnostics to cancel a queued or running run.

The tools are advertised only to a Konduo user-backed API key with the mcp:invoke scope. An OAuth or OIDC MCP subject is not guessed to be a local user. API-key resource IDs and wildcard restrictions apply to every resource discovered through traversal, not only to requested roots.

start_connected_diagnostics accepts the same starting resources, traversal mode, depth, lookback, and resource limit as the UI. Supply a stable idempotency_key for a call that may be retried.

get_connected_diagnostics_result consumes a successfully returned terminal result by default. Use consume=false when one more read is required, then consume the result on the final read. Deferring consumption never extends storage beyond the fixed one-hour safety cleanup.

Security and Audit

  • Connected Diagnostics is read-oriented and does not mutate managed-resource configuration or state.
  • Scope preview, execution, and result reads all reapply current resource permissions.
  • Credentials, unrestricted logs, secret-shaped fields, and arbitrary plugin payloads are not stored.
  • Run creation, cancellation, and result consumption are audited.
  • Suspected causes are high-risk evidence correlations, not confirmed root causes.
  • A diagnostic result never changes existing alert or anomaly lifecycle state.

Troubleshooting

Connected Diagnostics is not visible

Confirm that the Enterprise license is valid and the Connected Diagnostics feature is enabled. On resource pages, the current account needs resource-read access. In an Analysis Workspace, the active page must contain at least one ready resource panel.

The scope is empty or smaller than expected

Confirm that starting resources are active and readable. In Relations, verify that required relations are approved, then review traversal mode, depth, and resource limit. Suggested, rejected, and unreadable resources are intentionally excluded.

The run is partial or shows collection gaps

Review Unsupported, Unavailable, or Invalid response reasons in Collection gaps. Check plugin availability, resource connectivity, diagnostic-history routes, and metric or log-source configuration before rerunning. Findings from successful targets remain usable in a partial result.

The result disappeared after refresh

This is expected. Connected Diagnostics does not save result history. Pin required findings to the workspace or record them in an approved operational system before leaving the page.

Coverage and Operational risk are both low

Do not conclude that the system is healthy. Unknown and unavailable findings do not create artificial risk, but they lower Coverage and Evidence confidence. Resolve collection gaps and run the diagnosis again.

MCP tools are not listed

Confirm that both the MCP gateway and Connected Diagnostics are enabled. Use an API key mapped to a Konduo user, verify the mcp:invoke scope, and check that starting resources are within the API key's resource restrictions.

  • enterprise/docs/connected-diagnostics-contract.md
  • enterprise/docs/diagnostic-risk-v1.md
  • enterprise/backend/docs/openapi-connected-diagnostics.yaml
  • enterprise/docs/analysis-workspace-contract.md
  • enterprise/docs/instance-relations-contract.md