Skip to main content

Debugging, Audit, and Observability

TeaQL debugging is a chain of evidence, not a single “print SQL” switch. Start with the business intent, follow the typed execution path, inspect safe SQL evidence, and use operator-only output only inside a trusted diagnostic boundary. Audit and RuntimeTelemetry answer different questions and must remain separate.

Choose the evidence that answers the question​

EvidenceWhat it answersSafe application output?
Query comment and purposeWhat was requested, and why?Yes, after the application's own classification policy
Mutation audit reasonWhy was state changed?Audit-controlled; not ordinary telemetry
Typed trace pathWhich request, relation, provider, and SQL operation produced this work?Yes, when entity/relation names are approved
Safe SQL evidenceWhich parameterized SQL shape ran, how long it took, and how many rows it returned or changed?Yes; it excludes values
Diagnostic SQLWhat exact value-rendered SQL can an operator replay?No; trusted operator boundary only
RuntimeTelemetryWhere time and failures occur across query, mutation, relation, provider, cache, TFP, and audit operationsYes, using only the low-cardinality contract
Application audit eventWho changed what, why, and through which business action?Send to the governed audit sink, not a debug response
Request-policy/readiness evidenceWas the trusted tenant, policy, provider, and schema boundary installed?Report pass/fail, never secret context values

A repeatable diagnostic window​

Use the same order in every language:

  1. Start a fresh evidence window and select all, query-only, or mutation-only. Mode changes clear earlier evidence in runtimes with an evidence store; clear an application sink explicitly where the language exposes Clear.
  2. Run one bounded operation with non-empty query comment and purpose, or a non-empty mutation audit reason.
  3. Read only the safe projection first: operation, parameterized SQL, parameter count, elapsed microseconds, result/affected count, and bounded summary.
  4. Correlate the typed trace path with the Query, Facet, aggregate, or selected relation that initiated the SQL.
  5. Escalate to value-rendered diagnostic SQL only in an access-controlled environment. Never put it in HTTP errors, agent responses, or ordinary telemetry.
  6. Disable or remove the diagnostic sink after the investigation, and verify that sensitive output is not retained.

Diagnose Query, Facet, and Aggregation​

The business syntax is documented in Query, Mutation, Analytics, and Facets. Use its language examples, then interpret the captured evidence as follows.

Query and deep relations​

For the main Request, expect one trace beginning at query and ending at the provider SQL operation. Every explicitly loaded relation has its own typed lineage. A useful trace reads conceptually as:

query → School → students → enrollments → provider → select

This distinguishes an outer-query problem from a child relation problem. A surprising number of nearly identical child traces is a practical N+1 signal; an absent child trace usually means the relation was not selected, policy removed it, or construction stopped before execution. A loaded null, an empty loaded list, and a relation that was not selected are different outcomes.

For per-parent Top-N, verify that the provider evidence represents the runtime's database-side partition/window plan. Loading every child and slicing in application memory is not an equivalent result.

Facets​

For a Facet investigation, record all of these inputs before comparing counts:

  • the normalized outer filter;
  • the Facet name;
  • the typed child Request and allowed bucket predicate;
  • includeAllFacets/language equivalent;
  • every requested metric alias.

includeAll = true keeps allowed unmatched buckets with a zero count; includeAll = false returns matched buckets only. A child predicate restricts the allowed domain and is not interchangeable with that flag. The main rows and Facet counts must be diagnosed from the same trusted context and normalized outer population.

Aggregates​

For COUNT, SUM, AVG, MIN, and MAX, check the filtered population, group fields, metric aliases, SQL operation count, and returned aggregation carrier. Relation aggregates belong to each parent; grouped aggregates produce distribution rows. Do not compare either one with an in-memory total computed over a paged subset.

Diagnose Mutation and audit​

A mutation diagnostic should preserve four independent outcomes:

  1. Checker/Fix and field-specific validation result.
  2. Request-policy authorization result.
  3. Provider mutation, optimistic-version check, and affected-row count.
  4. Raw row audit plus the separately attributable, masked application audit event.

Missing audit intent, not-found, validation failure, optimistic conflict, and provider failure remain failures; a diagnostic adapter must not convert them to an empty successful result. RuntimeTelemetry is fail-open and may never decide whether the audit event is delivered. An application may choose a durable or fail-closed audit policy independently.

RuntimeTelemetry and OpenTelemetry​

All seven runtimes expose a provider-neutral, no-op-by-default telemetry contract for these operation families:

  • query, mutation, and relation_load;
  • provider, cache, and tfp;
  • audit.

The lifecycle is start/success/failure with duration, safe attributes, result cardinality where appropriate, and an error category derived from the native error type. The application owns the OpenTelemetry SDK, processors, exporter, Collector destination, sampling, flush, and shutdown. Exporter failure must not change a Query, save, audit, or readiness result.

Do not export entity IDs, user IDs, tenant IDs, query or database parameters, field values, audit reasons, request bodies, full URLs, or rendered SQL. W3C trace context may travel through TFP string-map metadata; it does not belong in the Query or Mutation payload.

Diagnose cache and lock behavior​

Cache operations use the cache telemetry family. Diagnose local.get, local.put, local.remove, and remote equivalents by operation, hit/miss or stored/removed result, duration, and error category. Cache keys and values are forbidden telemetry fields. A miss is not itself an error; an unavailable remote provider, serialization failure, and an expired entry must remain distinguishable in application tests even where the selected adapter chooses a fail-open miss.

For an inconsistent Query/Facet screen, verify that row, aggregate, Facet, and relation cache keys share the same tenant, policy, normalized filter, locale, and version inputs, and that an audited Mutation invalidates all affected namespaces.

For locks, record provider presence, local versus remote scope, key namespace shape without its sensitive value, wait timeout, lease, acquisition outcome, critical-section duration, and owner-safe release result. Never emit the full business lock key when it contains tenant or entity identifiers. A successful no-op result from a missing optional remote provider is not evidence of mutual exclusion; readiness must prove required provider installation.

The current provider maturity and owner-release qualifications are documented in Cache, Lock, and Cloud Runtime Infrastructure.

Diagnose cloud lifecycle​

Keep liveness, readiness, and dependency health separate:

  • liveness answers whether the process/event loop is alive;
  • readiness proves provider/schema plus required Redis, Nacos, Consul, policy, tenant, and audit dependencies for this deployment profile;
  • registration/discovery evidence proves the intended service name, namespace/group/datacenter, address, port, health filter, and deregistration;
  • graceful shutdown marks readiness out-of-service before deregistration and connection draining.

Do not log ACL tokens, Nacos credentials, returned configuration values, full service metadata, or discovery URLs containing secrets. Export dependency duration and safe error category through application telemetry. Java and Go Nacos/Consul evidence is currently protocol-level HTTP evidence; Rust's cloud stack additionally exposes Actuator-compatible endpoints, Prometheus metrics, and lifecycle composition. None of these claims implies a live production cluster has been tested for the reader's deployment.

Sensitive diagnostic output​

All runtimes require the exact opt-in below before sensitive plaintext logging is allowed:

TEAQL_ALLOW_SENSITIVE_PLAINTEXT_LOGS=I_UNDERSTAND_SENSITIVE_DATA_MAY_BE_WRITTEN_TO_DISK

The opt-in does not sanitize existing files, third-party driver logs, crash dumps, or application logging. Credentials remain redacted. Java and Rust also support TEAQL_SQL_LOG, TEAQL_SQL_LOG_TABLES, TEAQL_AUDIT_LOG, and TEAQL_AUDIT_LOG_ENTITIES, with _silent, _summary, _full, and _full_with_payload levels. Treat payload mode as privileged production access, not as the default troubleshooting recipe.

Trusted tenant and Request Policy​

The invariant is shared: tenant, provider/data service, request policy, audit sink, and safety limits come from the authenticated server composition root. They must be recursively rejected from untrusted JSON and TFP payloads. Readiness must reach the real provider/schema path and fail when trusted tenant or policy dependencies are missing.

The current Request Policy hook depth is not identical in every runtime:

RuntimeCurrent policy hook coverage to document and test
JavaSelect, insert, update, delete, and recover
RustSelect, insert, update, delete, and recover, plus entity data-service behavior hooks
GoSelect, insert, update, delete, and recover
SwiftQuery policy hook; mutation isolation must also be enforced at the provider/application/TFP boundary
PythonQuery policy hook; mutation isolation must also be enforced at the provider/application/TFP boundary
.NETQuery policy hook; mutation isolation must also be enforced at the provider/application/TFP boundary
TypeScriptComposition stores requestPolicy, but the current core prepareQuery() does not automatically apply that resource; enforce tenant scope in the server adapter/custom data service and retain negative tests

This table describes the current general runtime hooks, not a waiver for any language. A protected application must test list, direct ID, deep relation, Facet, aggregate, create, update, delete, background-job, and federation paths. Never use a caller-supplied tenant filter as the security boundary.

Language guides​

LanguageEnable and interpret evidence
JavaJava debugging and observability
RustRust debugging and observability
TypeScriptTypeScript debugging and observability
SwiftSwift debugging and observability
PythonPython debugging and observability
C#/.NET.NET debugging and observability
GoGo debugging and observability

Before closing an investigation, run the Debugging and governance checklist.