Observability Guide

Cloudflare already observes a Worker. Workers Logs ingests every line console.* writes and indexes its JSON fields; Workers Traces spans the bindings a request touched; Workers Metrics counts invocations, errors, and CPU time. None of that is code you write, and none of it is reimplemented here.

What this package adds is what the platform's own view cannot travel with: a stable event name on every line, a request id that survives a Service Binding hop, enough context that a line means something on its own, and the discipline about what must not go on one.

The single rule everything below follows:

A log line must mean something to whoever reads it, wherever they read it.

Where they read it is the part that decides how much the line has to carry, and it is a deployment's answer rather than this package's. Inside the Workers Logs console, Cloudflare has already put the script name, the severity, the message and the trigger on the record; a line repeating them is a column paid for twice. Outside it — an export, a Logpush sink, an OTel endpoint, a file — none of that travels, and a line that assumed it did says nothing.

So the line is configurable, and the default assumes the console. See Choosing what a line carries.

What Cloudflare writes, and what we write

Worth knowing before deciding anything else, because half of it is free.

Every stored record has three parts. source is the JSON object we logged. $workers is the raw trace envelope — useful, ugly, and no place for a dashboard column. $metadata is what the dashboard actually renders from:

$metadata field Written by
service Cloudflare — the script name
trigger Cloudflare — the method and path, e.g. POST /
rayId, requestId, type, id, account Cloudflare
level us — lifted from the top-level level key
message us — lifted from the top-level message key

Only those two cross over. This is a fixed contract, not a convention: in a 330-row production export $metadata.message matched source.message on 171 of 171 rows and source.event on 0 of 136. No other key we write reaches $metadata, and there is no documented way to add one.

Two consequences:

  • A line must emit a top-level message or the dashboard's Message column is blank. logToAnalytics does; so does every logger.* call, falling back to the event name.
  • service, and the method and path, are already there. Which is why the default field list leaves service off and a Worker inside the console loses nothing.

event and message

Two keys, two audiences, and they are not interchangeable.

Key For Example
event machines — filtering, grouping, counting request.completed
message people — the column a log viewer shows Request completed: POST /api/v1/search (400)
logger.info('notification.sent', { notificationId }, `Notification sent to ${provider}`)
level  event                 message
info   notification.sent     Notification sent to telegram
warn   notification.failed   Telegram rejected the send (502)
warn   message.refused       Refused a message on conversation c-91
error  invoice.fetch.failed  Could not fetch invoices for August

The third argument is optional. Omit it and message reads the event name, so the column is never blank — good enough for a line nobody scans, and worth a sentence for one somebody will.

Write the event for a query and the message for a person. A dotted name you can count, and prose that says what happened:

// ✗ Prose in the event slot: a different string on every line, so `event` counts nothing.
logger.info('Sent a notification to telegram')

// ✗ Interpolation in the event slot, same problem with extra steps.
logger.info(`notification.sent.${provider}`)

// ✗ The Worker's name in the event costs `event = notification.sent` its meaning across services.
logger.info('example-worker.notification.sent', { provider })

// ✓
logger.info('notification.sent', { provider }, `Notification sent to ${provider}`)

The event should stay meaningful if the code moves to another Worker. $metadata.service will follow it there on its own.

Choosing what a line carries

LOG_FIELDS in wrangler.json, as a comma-separated list:

{
  "vars": {
    "SERVICE_NAME": "example-worker",
    "LOG_LEVEL": "info",
    "LOG_FIELDS": "requestId,method,path,ip,status"
  }
}

That is also the default, from cloudflareDefaults.logging.fields — the correlation id, the method, the path, the client address, and the status. Everything Cloudflare does not already record.

Add When
service lines leave Workers Logs — Logpush, OTel, a file
colo, country chasing something regional
durationMs you want timing on the line rather than from a trace
component you are reading a component's lines rather than a request's
time the consumer does not stamp its own ingestion timestamp

* keeps every field. "" keeps none, leaving the level, the event and the message — which is a complete line if the dashboard is your only reader.

The list governs ambient fields only. What a call site passes to a single log call is always emitted:

// `reason` survives whatever LOG_FIELDS says. Somebody chose to log it.
logger.warn('authentication.failed', { reason: 'api_key_absent' })

That split is the whole design. Ambient fields are stamped by middleware for the span of a request and are the ones a deployment might reasonably not want; a payload was typed out for one call, and filtering it would throw away the answer somebody logged the line to get.

Where service comes from

SERVICE_NAME in wrangler.json, declared beside the Worker's own name:

{
  "name": "example-worker",
  "vars": {
    "SERVICE_NAME": "example-worker",
    "LOG_LEVEL": "info"
  }
}

Both in one file, on purpose. A rename that misses one shows up in the same diff. The alternative — a config/LoggingConfig.ts holding the name — is a file a rename does not touch, which can leave production logs emitting a stale service name.

The fallback order is SERVICE_NAME → logging.service on the configuration surface → the literal string "unknown". That last step matters: a deployment that forgets the var still emits a service field, and service = unknown is a query that finds every misconfigured Worker in the account. An omitted field would make the same mistake invisible.

What is never restated

The invocation's outcome, CPU time, wall time, and script version — no matter what LOG_FIELDS says. None of them describe the request in a way this middleware could compute more cheaply than Cloudflare already does, and they live on the invocation record, which invocation_logs: false turns off anyway. Timing comes from a trace.

colo and country are read once per request from request.cf and offered as context, so LOG_FIELDS decides whether they reach the line. They reach Analytics Engine either way, through AnalyticsService — for the unrelated reason that an aggregate is joined to nothing and has to carry its own dimensions.

Using the logger

The logger is Logger from @bayudwiyansatria/core. There is no Cloudflare-specific logger: everything it needs from the environment — SERVICE_NAME, LOG_LEVEL, LOG_FIELDS — is a plain variable, so the kernel reads it without knowing which platform supplied it. What is Cloudflare-specific is the default field list, and that lives in cloudflareDefaults rather than in the kernel.

import { Logger } from '@bayudwiyansatria/core'

const log = Logger.fromEnv(env, { component: 'NotificationService' })

log.info('notification.sent', { notificationId, provider, durationMs })
log.warn('cache.miss', { key })
log.error('external_api.error', { provider, status }, `${provider} rejected the send (${status})`)

Inside a request, take the logger off the context instead of building one — it already carries the request id, method, and path:

api.post('/send', async ctx => {
  const log = ctx.get('logger')

  log.info('notification.sent', { notificationId })
})

Four things to decide, and only the first three are required:

Argument Is Example
the level how bad it is info, warn, error, debug
the event what happened, as a stable dotted name 'notification.sent'
the data what varied this time { notificationId, durationMs }
the message the same thing, said to a person `Sent to ${provider}`

Never build the JSON yourself, and never put the varying part in the event. console.log(`Sent ${id} in ${ms}ms`) produces a different string every time it runs, which makes it un-filterable, un-countable, and un-chartable — the three things a log stream is for. The sentence is where that string belongs, beside an event that stays the same.

component, not service

A line that needs to say which part of this Worker wrote it names a component:

Logger.fromEnv(env, { component: 'IngestionService' })

service answers "which Worker" — $metadata.service has it, from the script name. component answers "where inside it". Before 1.2.0 call sites passed a class name as service, so the column meant "which Worker" on some lines and "which class" on others, and neither reading was reliable.

component is context, so it is subject to LOG_FIELDS and is not in the default list. Add it when you are reading a component's lines rather than a request's; a well-named event usually already says which part of the Worker wrote it.

Event naming

subject.outcome — lowercase, dotted, snake_case inside a segment:

request.completed          authentication.failed      external_api.error
request.failed             authorization.denied       ai.error
request.unhandled          rate_limit.exceeded        kv.error

The cross-cutting names — the ones that mean the same thing in every Worker — are exported as events, so the spelling is not available to typos:

import { events } from '@bayudwiyansatria/cloudflare'

logger.error(events.EXTERNAL_API_ERROR, { provider, status })

A Worker's own vocabulary stays in that Worker as string literals. order.sent, message.received, and invoice.fetch.completed are each meaningful in exactly one place, and hoisting them into a shared constant would centralise a list nothing shares.

Rules of thumb:

  • An event marks a meaningful occurrence, not the entry to a function. Do not log every call.
  • One occurrence, one line. If a "starting" line is immediately followed by a "completed" line carrying the same identifiers, the first belongs at debug or nowhere.
  • Do not rename an existing event without a reason. A rename is a gap in every chart built on the old name.

request.failed and request.unhandled

Two names, deliberately, and the pair is worth understanding.

Inside a Hono app the request middleware never sees an exception — the framework has already turned it into a 500 by the time next() returns — so the only place the stack exists is app.onError. One name would make every unhandled failure count twice.

request.failed       one per failed request, from logToAnalytics — status, duration
request.unhandled    the exception and its stack, from app.onError

Both carry the same request id, so one query on that id returns the failure and its cause together.

Request correlation

Cloudflare gives every request from the internet a CF-Ray, and CloudflareRequestMetadata reads it as the request id — the id the dashboard's own logs are keyed by, so a line here and a line there can be joined without inventing anything.

A subrequest gets no new one. When a Worker calls another over a Service Binding, the callee sees whatever headers the caller built and nothing the edge added. Without a header of its own the trail stops at each hop, and a failure three Workers deep cannot be walked back to the request that caused it.

RequestCorrelation is that header. Stamp it on anything leaving the Worker:

import { RequestCorrelation } from '@bayudwiyansatria/cloudflare'

const requestId = ctx.get('requestId')

await env.NLP_WORKER.fetch('https://nlp/analyse', {
  method: 'POST',
  headers: { 'content-type': 'application/json', ...RequestCorrelation.of(requestId) },
  body
})

The callee picks it up automatically — CloudflareRequestMetadata prefers an inbound X-Request-Id over minting a fresh id — so one id spans the whole fan-out with no further work at either end.

It is correlation and only correlation. The value is attacker-controlled: anyone may send X-Request-Id. That is fine for joining log lines and unacceptable for anything else. It must never gate access, identify a caller, or stand in for a credential. The worst a forged value achieves is two unrelated requests appearing under one id in a query.

What must never be logged

Never, under any circumstances:

API keys              JWTs                  passwords
Authorization headers access tokens         cookies
secrets               refresh tokens        card numbers

Be especially careful with the things that are not obviously credentials but are just as damaging in a log stream that is indexed, exported, and retained on its own schedule:

  • chatbot messages, AI prompts, and AI responses
  • authentication payloads
  • request and response bodies
  • external API payloads
  • personal information

Do not log complete request or response payloads. Log the few fields that explain what happened. In practice that means shapes and sizes rather than contents:

// ✓ enough to diagnose, nothing to leak
log.debug('nlp.analysis.started', { requestId, messageLength: userMessage.length })
log.warn('authentication.failed', { reason: 'api_key_mismatch' })
log.error('external_api.error', { provider, status, destination: `${url.host}/…` })
// ✗
log.info('request received', { body: await ctx.req.json() })
log.warn('bad key', { presented: ctx.req.header('x-api-key') })

The safety net

Logger runs every payload through redact from @bayudwiyansatria/core before it is serialised. Fields whose names look like credentials — apiKey, API_KEY, x-api-key, authorization, set-cookie, refresh_token, clientSecret, and the rest — are replaced with [redacted], recursively, through arrays and nested objects.

log.error('external_api.error', { provider: 'telegram', apiKey: 'live-key' })
// data: { provider: 'telegram', apiKey: '[redacted]' }

Treat it as the last line of defence, not the first. It matches on field names, so a secret under an innocent name gets through — and a whole request body will be logged in full apart from the fields it happens to recognise. The rule at the call site is what actually protects you.

Wrangler configuration

Enable logs and traces per Worker:

{
  "name": "example-worker",
  "vars": {
    // Beside `name`, so a rename that misses it is visible in the same diff.
    "SERVICE_NAME": "example-worker",
    "LOG_LEVEL": "info",
    "LOG_FIELDS": "requestId,method,path,ip,status"
  },
  "observability": {
    "enabled": true,
    "logs": { "head_sampling_rate": 1, "invocation_logs": false },
    "traces": { "enabled": true, "head_sampling_rate": 0.1 }
  },
  "upload_source_maps": true
}

upload_source_maps is what makes a production stack trace readable. Without it every frame points into one minified bundle. Wrangler generates and uploads the map on wrangler deploy; the remapping happens out of band, so it costs nothing at runtime.

invocation_logs: false

Cloudflare writes one invocation record per request in addition to whatever you log. In a 330-row production export, 152 of them — 46% — were invocation records, and they carry none of your fields, so every application column renders blank on those rows. What they say about the request, request.completed and request.failed already say.

Turning them off is the single biggest reduction available, and the cost is specific and worth knowing: outcome, cpuTimeMs and wallTimeMs exist only on those records. Timing then comes from traces, and the outcome from the events. Leave them on for a Worker where you read CPU time out of Workers Logs regularly.

Choosing a sampling rate

Do not apply one number everywhere. Weigh volume, criticality, and what you would actually want to have when something breaks at 03:00.

Environment Logs Traces
Development 100% 100%
Staging 100% 100%
Production 100% 1–20%

Logs at 100% for anything below serious volume. Logs are the answer to "what happened to this request", and a sampled log stream cannot answer it — the one request you are looking for is the one that was dropped. Lower it only when volume makes the bill the bigger problem.

Traces well below 100%. A trace answers "where did the time go", which is a question about the shape of traffic, not about one request. A few percent of a busy Worker describes it as well as all of it, at a fraction of the cost. Raise it for a Worker that runs rarely and matters when it does — a nightly cron is a handful of invocations a day, so tracing every one costs nothing.

Where a Worker defines environments, each gets its own block:

{
  "env": {
    "staging": {
      "observability": {
        "enabled": true,
        "logs": { "head_sampling_rate": 1 },
        "traces": { "enabled": true, "head_sampling_rate": 1 }
      }
    }
  }
}

Tracing

Use Cloudflare's native tracing. It already spans D1, KV, R2, Workers AI, Queues, Durable Objects, Service Bindings, and outbound fetch — every binding this package wraps. Do not add a third-party tracing library, and do not hand-instrument an operation Cloudflare already traces.

Add a log line beside a traced operation only when it carries application context a span has no field for — which provider was chosen and why, which model, which cache key. The span says the call took 340ms; your line says it went to the fallback provider.

Metrics

Cloudflare's own metrics answer request count, error rate, 4xx and 5xx rates, CPU time, and request duration. Do not build custom metrics for any of them.

The debugging path is:

Metrics  →  something is wrong
Logs     →  what happened
Traces   →  where the time or the failure went

Investigating a production failure

  1. Metrics — error rate is up on example-worker. Note the window.
  2. Logs — filter service = example-worker and event = request.failed over that window. The status and duration are on the line.
  3. Cause — take a requestId from one of those lines and filter on it alone. Every line the request produced comes back: the request.unhandled line with the stack, the external_api.error line with the provider and status, and whatever the handler logged in between.
  4. Across Workers — the same id spans Service Binding hops, so drop the service filter and the downstream Worker's lines appear on the same trail.
  5. Traces — if the failure is latency rather than an exception, open a trace for the window and read which binding or subrequest consumed the time.

Steps 3 and 4 are why the request id matters, and why propagating it on outbound calls is not optional.

results matching ""

    No results matching ""