Debugging, Audit, and Observability
TeaQL debugging is a chain of evidence, not a single “print SQL” switch. Start with the business intent, follow the typed execution path, inspect safe SQL evidence, and use operator-only output only inside a trusted diagnostic boundary. Audit and RuntimeTelemetry answer different questions and must remain separate.
Choose the evidence that answers the question
| Evidence | What it answers | Safe application output? |
|---|---|---|
Query comment and purpose | What was requested, and why? | Yes, after the application's own classification policy |
| Mutation audit reason | Why was state changed? | Audit-controlled; not ordinary telemetry |
| Typed trace path | Which request, relation, provider, and SQL operation produced this work? | Yes, when entity/relation names are approved |
| Safe SQL evidence | Which parameterized SQL shape ran, how long it took, and how many rows it returned or changed? | Yes; it excludes values |
| Diagnostic SQL | What exact value-rendered SQL can an operator replay? | No; trusted operator boundary only |
| RuntimeTelemetry | Where time and failures occur across query, mutation, relation, provider, cache, TFP, and audit operations | Yes, using only the low-cardinality contract |
| Application audit event | Who changed what, why, and through which business action? | Send to the governed audit sink, not a debug response |
| Request-policy/readiness evidence | Was the trusted tenant, policy, provider, and schema boundary installed? | Report pass/fail, never secret context values |
A repeatable diagnostic window
Use the same order in every language:
- Start a fresh evidence window and select
all, query-only, or mutation-only. Mode changes clear earlier evidence in runtimes with an evidence store; clear an application sink explicitly where the language exposesClear. - Run one bounded operation with non-empty query comment and purpose, or a non-empty mutation audit reason.
- Read only the safe projection first: operation, parameterized SQL, parameter count, elapsed microseconds, result/affected count, and bounded summary.
- Correlate the typed trace path with the Query, Facet, aggregate, or selected relation that initiated the SQL.
- Escalate to value-rendered diagnostic SQL only in an access-controlled environment. Never put it in HTTP errors, agent responses, or ordinary telemetry.
- Disable or remove the diagnostic sink after the investigation, and verify that sensitive output is not retained.
Diagnose Query, Facet, and Aggregation
The business syntax is documented in Query, Mutation, Analytics, and Facets. Use its language examples, then interpret the captured evidence as follows.
Query and deep relations
For the main Request, expect one trace beginning at query and ending at the
provider SQL operation. Every explicitly loaded relation has its own typed
lineage. A useful trace reads conceptually as:
query → School → students → enrollments → provider → select
This distinguishes an outer-query problem from a child relation problem. A surprising number of nearly identical child traces is a practical N+1 signal; an absent child trace usually means the relation was not selected, policy removed it, or construction stopped before execution. A loaded null, an empty loaded list, and a relation that was not selected are different outcomes.
For per-parent Top-N, verify that the provider evidence represents the runtime's database-side partition/window plan. Loading every child and slicing in application memory is not an equivalent result.
Facets
For a Facet investigation, record all of these inputs before comparing counts:
- the normalized outer filter;
- the Facet name;
- the typed child Request and allowed bucket predicate;
includeAllFacets/language equivalent;- every requested metric alias.
includeAll = true keeps allowed unmatched buckets with a zero count;
includeAll = false returns matched buckets only. A child predicate restricts
the allowed domain and is not interchangeable with that flag. The main rows
and Facet counts must be diagnosed from the same trusted context and normalized
outer population.
Aggregates
For COUNT, SUM, AVG, MIN, and MAX, check the filtered population,
group fields, metric aliases, SQL operation count, and returned aggregation
carrier. Relation aggregates belong to each parent; grouped aggregates produce
distribution rows. Do not compare either one with an in-memory total computed
over a paged subset.
Diagnose Mutation and audit
A mutation diagnostic should preserve four independent outcomes:
- Checker/Fix and field-specific validation result.
- Request-policy authorization result.
- Provider mutation, optimistic-version check, and affected-row count.
- Raw row audit plus the separately attributable, masked application audit event.
Missing audit intent, not-found, validation failure, optimistic conflict, and provider failure remain failures; a diagnostic adapter must not convert them to an empty successful result. RuntimeTelemetry is fail-open and may never decide whether the audit event is delivered. An application may choose a durable or fail-closed audit policy independently.
RuntimeTelemetry and OpenTelemetry
All seven runtimes expose a provider-neutral, no-op-by-default telemetry contract for these operation families:
query,mutation, andrelation_load;provider,cache, andtfp;audit.
The lifecycle is start/success/failure with duration, safe attributes, result cardinality where appropriate, and an error category derived from the native error type. The application owns the OpenTelemetry SDK, processors, exporter, Collector destination, sampling, flush, and shutdown. Exporter failure must not change a Query, save, audit, or readiness result.
Do not export entity IDs, user IDs, tenant IDs, query or database parameters, field values, audit reasons, request bodies, full URLs, or rendered SQL. W3C trace context may travel through TFP string-map metadata; it does not belong in the Query or Mutation payload.
Diagnose cache and lock behavior
Cache operations use the cache telemetry family. Diagnose local.get,
local.put, local.remove, and remote equivalents by operation, hit/miss or
stored/removed result, duration, and error category. Cache keys and values are
forbidden telemetry fields. A miss is not itself an error; an unavailable
remote provider, serialization failure, and an expired entry must remain
distinguishable in application tests even where the selected adapter chooses a
fail-open miss.
For an inconsistent Query/Facet screen, verify that row, aggregate, Facet, and relation cache keys share the same tenant, policy, normalized filter, locale, and version inputs, and that an audited Mutation invalidates all affected namespaces.
For locks, record provider presence, local versus remote scope, key namespace shape without its sensitive value, wait timeout, lease, acquisition outcome, critical-section duration, and owner-safe release result. Never emit the full business lock key when it contains tenant or entity identifiers. A successful no-op result from a missing optional remote provider is not evidence of mutual exclusion; readiness must prove required provider installation.
The current provider maturity and owner-release qualifications are documented in Cache, Lock, and Cloud Runtime Infrastructure.
Diagnose cloud lifecycle
Keep liveness, readiness, and dependency health separate:
- liveness answers whether the process/event loop is alive;
- readiness proves provider/schema plus required Redis, Nacos, Consul, policy, tenant, and audit dependencies for this deployment profile;
- registration/discovery evidence proves the intended service name, namespace/group/datacenter, address, port, health filter, and deregistration;
- graceful shutdown marks readiness out-of-service before deregistration and connection draining.
Do not log ACL tokens, Nacos credentials, returned configuration values, full service metadata, or discovery URLs containing secrets. Export dependency duration and safe error category through application telemetry. Java and Go Nacos/Consul evidence is currently protocol-level HTTP evidence; Rust's cloud stack additionally exposes Actuator-compatible endpoints, Prometheus metrics, and lifecycle composition. None of these claims implies a live production cluster has been tested for the reader's deployment.
Sensitive diagnostic output
All runtimes require the exact opt-in below before sensitive plaintext logging is allowed:
TEAQL_ALLOW_SENSITIVE_PLAINTEXT_LOGS=I_UNDERSTAND_SENSITIVE_DATA_MAY_BE_WRITTEN_TO_DISK
The opt-in does not sanitize existing files, third-party driver logs, crash
dumps, or application logging. Credentials remain redacted. Java and Rust also
support TEAQL_SQL_LOG, TEAQL_SQL_LOG_TABLES, TEAQL_AUDIT_LOG, and
TEAQL_AUDIT_LOG_ENTITIES, with _silent, _summary, _full, and
_full_with_payload levels. Treat payload mode as privileged production
access, not as the default troubleshooting recipe.
Trusted tenant and Request Policy
The invariant is shared: tenant, provider/data service, request policy, audit sink, and safety limits come from the authenticated server composition root. They must be recursively rejected from untrusted JSON and TFP payloads. Readiness must reach the real provider/schema path and fail when trusted tenant or policy dependencies are missing.
The current Request Policy hook depth is not identical in every runtime:
| Runtime | Current policy hook coverage to document and test |
|---|---|
| Java | Select, insert, update, delete, and recover |
| Rust | Select, insert, update, delete, and recover, plus entity data-service behavior hooks |
| Go | Select, insert, update, delete, and recover |
| Swift | Query policy hook; mutation isolation must also be enforced at the provider/application/TFP boundary |
| Python | Query policy hook; mutation isolation must also be enforced at the provider/application/TFP boundary |
| .NET | Query policy hook; mutation isolation must also be enforced at the provider/application/TFP boundary |
| TypeScript | Composition stores requestPolicy, but the current core prepareQuery() does not automatically apply that resource; enforce tenant scope in the server adapter/custom data service and retain negative tests |
This table describes the current general runtime hooks, not a waiver for any language. A protected application must test list, direct ID, deep relation, Facet, aggregate, create, update, delete, background-job, and federation paths. Never use a caller-supplied tenant filter as the security boundary.
Language guides
| Language | Enable and interpret evidence |
|---|---|
| Java | Java debugging and observability |
| Rust | Rust debugging and observability |
| TypeScript | TypeScript debugging and observability |
| Swift | Swift debugging and observability |
| Python | Python debugging and observability |
| C#/.NET | .NET debugging and observability |
| Go | Go debugging and observability |
Before closing an investigation, run the Debugging and governance checklist.