@outputai/core worker and in @outputai/llm cost reporting. Nothing in your code needs to change, but four runtime behaviors differ: the worker now bounds its own shutdown, every way it can end now has a defined outcome and exit code, a worker refuses to start unless it can publish a catalog matching its own source code, and LLM costs fall back to a bundled pricing snapshot instead of coming back empty.
Worker shutdown is now bounded
TEMPORAL_SHUTDOWN_GRACE_TIME and TEMPORAL_SHUTDOWN_FORCE_TIME have defaults, and the previously hardcoded hook flush allowance is now OUTPUT_HOOK_FLUSH_TIMEOUT_MS:
Previously the worker never force-terminated: it waited indefinitely for in-flight activities and was stopped only when the platform sent
SIGKILL, which meant the hook flush was never reached at all on a deploy that caught in-flight work. It now stops draining after 20s and exits under its own control with code 1 - closing its Temporal connection, flushing hooks, and logging first.
For most projects this is an improvement, because activities get 15 seconds of runway instead of being asked to cancel the instant a deploy starts, and buffered hook data actually gets written.
The three values are one budget
The drain runs toTEMPORAL_SHUTDOWN_FORCE_TIME and only then does the flush start, so the worst-case shutdown is force time plus flush timeout. That total has to fit inside the window your platform allows between SIGTERM and SIGKILL. The defaults total 25 seconds against the 30 seconds most platforms allow, leaving 5 seconds of margin.
The hook flush is the part to watch when retuning. It used to be 30 seconds, which on a default platform window could not complete; it is now 5 seconds, which fits. Five seconds is ample for handlers that write to stdout or an in-process buffer, which is what hook handlers should be doing - they all share this one budget, so a single handler awaiting remote I/O spends the allowance the others need. If you do have a handler that must await a network round trip, batch it and raise both this value and your platform’s window, in that order.
Set the values explicitly if your activities drain slowly
If your activities routinely run longer than 20 seconds and you raised your platform’s shutdown window to accommodate them, set the variables explicitly. Otherwise those activities are cancelled at 15 seconds and abandoned at 20 on every deploy, then retried on the new instance. Keep grace below force, and force plus flush below your platform’s shutdown window, so the worker always exits before it is killed:maxShutdownDelaySeconds, which defaults to 30 seconds and can be raised to 300. On Kubernetes it is terminationGracePeriodSeconds.
Raising these only extends shutdowns started by a signal. When an uncaught exception or unhandled rejection starts the shutdown instead, the worker force quits 60 seconds later regardless of TEMPORAL_SHUTDOWN_FORCE_TIME, exiting 1 after logging Uncaught exception handling timed out, force quitting.... That cap is fixed, so a force time above 60 seconds only fully applies to signal shutdowns.
An empty value falls back to the default rather than disabling the bound, so there is no longer a way to configure an unbounded drain.
Shutdown scenarios
Every path that ends the worker changed shape. The exit code is also explicit now: v0.12.0 only calledprocess.exit on the failure path and let the event loop empty on the success path.
If you alert on non-zero worker exits, the situations that produce one changed. A drain that outlives its window is now one of them, where v0.12.0 hung until the platform stepped in, so expect an exit
1 from any deploy that lands on work longer than TEMPORAL_SHUTDOWN_FORCE_TIME. Nothing in the logs separates it from another failure to stop: the Stopping Worker error warning reports IllegalStateError: Worker still in use, because the forced termination leaves in-flight polls behind, and Temporal’s own Worker failed line serializes its error as empty. Treat a non-zero exit on a deploy you triggered as routine. An uncaught error still exits 1, but now only after draining.
Two log lines moved with this. The failure log is now Worker error rather than Fatal error, and the last line is always Bye, where a failing worker used to end on Exiting....
Workers exit when the catalog does not match
The worker publishes its catalog during startup and treats it as a gate. If it cannot publish a catalog matching its own source code, it retries twice and then exits rather than serving requests against an outdated catalog. A worker that would previously have started with a stale catalog now fails its deploy instead. The failure is logged asWorker error with the underlying Temporal error.
Reconciliation also changed: a stale catalog workflow is now terminated rather than asked to complete. A catalog that stopped processing workflow tasks no longer blocks every deploy that follows it.
LLM cost survives a pricing outage
@outputai/llm now ships a snapshot of the pricing table. When the live catalog is unreachable and nothing is cached, costs are calculated from that snapshot instead of coming back as null, and are reported with status: "imprecise" and pricingFreshness: "snapshot".
If you branch on cost === null to skip cost accounting during an outage, that branch stops firing and you start recording snapshot-derived figures instead. Check pricingFreshness or status for that case. cost is still null when usage cannot be normalized, and a provider or model missing from the table still produces a non-null, incomplete cost.
Checklist
- Set
TEMPORAL_SHUTDOWN_GRACE_TIME,TEMPORAL_SHUTDOWN_FORCE_TIMEandOUTPUT_HOOK_FLUSH_TIMEOUT_MSexplicitly if your activities need more than 20 seconds to drain, keeping force time plus flush timeout under your platform’s shutdown window. - Raise
OUTPUT_HOOK_FLUSH_TIMEOUT_MSif your hook callbacks do real I/O, since the allowance dropped from a hardcoded 30 seconds to 5. - Update alerts that key off worker exits or the
Fatal errorlog line, which is nowWorker error. - Expect a worker whose catalog cannot be published to exit instead of starting with a stale catalog.
- Replace any
cost === nullcheck used to detect a pricing outage with apricingFreshnessorstatuscheck.