Documentation

Everything you need to run the engine, write durable workflows, and integrate with it from your own services.

Pre-alpha. Weeks 1–10 of a 17-week plan are complete. The core (durability, replay, retries, leases, observability) is implemented and tested, but the public API is still moving and there is no HA yet. Don't run payroll on it.

Introduction

workflowd is a durable workflow execution engine: you write ordinary Go functions, and the engine guarantees they run to completion even if the process is killed halfway through. It is a lightweight, self-hostable alternative to Temporal — a single binary with no external database.

A workflow is a good fit when your business process:

  • spans multiple steps that each might fail (charge, ship, notify),
  • needs to wait — minutes, days, or for a human approval,
  • must never half-execute (charged but never shipped),
  • needs an audit trail of exactly what happened.

Quickstart

Requires Go 1.26+. No database, no broker, no container runtime.

1. Build and run the engine

git clone https://github.com/olayanju-1234/durable-workflow-engine
cd durable-workflow-engine
make build
go build -o workflowd ./cmd/workflowd

./workflowd --serve

That single process now exposes three things:

AddressProtocolPurpose
:7233gRPCWorkflow API — start workflows, poll and complete tasks
:8233HTTPDashboard + JSON API
:9090HTTP/metrics for Prometheus

2. Open the dashboard

Visit http://localhost:8233. It is compiled into the binary — nothing to deploy separately.

3. Start a workflow

The bundled demo registers an OrderSaga workflow. Start one over HTTP:

curl -X POST http://localhost:8233/api/workflows \
  -H 'Content-Type: application/json' \
  -d '{"type":"OrderSaga","input":"order-42"}'

It immediately parks in WAITING — no worker is running yet, so nobody can execute the reserve activity. Watch it in the dashboard, then start a worker (see Writing a worker) and it will run to completion.

4. Prove it's durable

Kill the engine with kill -9 while a workflow is mid-flight, then start it again. It re-reads the log, replays, and resumes at the exact step it was on. That behaviour is enforced by a test in the repo (kill_test.go) that SIGKILLs a real engine process at step 3 of 5 and asserts the workflow finishes with no duplicated steps.

Mental model

Everything rests on one idea: a workflow is an append-only list of events. The engine never stores "current state" anywhere — it derives state by replaying events.

read history → replay to find current position
             → schedule next step → record result → repeat

Because state is a pure function of the log, crash recovery is not a special code path. Restarting is just running the loop again.

TermMeaning
WorkflowYour orchestration function. Deterministic, no I/O. Runs on the engine.
ActivityOne unit of real work (charge a card, send an email). Runs on a worker. May fail and be retried.
StepOne entry in a workflow's history: an activity, a timer, or a signal wait.
HistoryThe append-only event log for one workflow execution. The only source of truth.
WorkerA separate process that polls for activity tasks over gRPC and reports results.
LeaseA time-boxed claim a worker holds on a dispatched task.

History & durability

Each workflow execution gets its own file: <data-dir>/<workflow-id>.history. Records are appended one at a time and fsync'd before the write is acknowledged, so a "recorded" event has genuinely reached the disk.

┌───────────┬──────────────┬────────────────────────────┐
│ 4B length │ 4B CRC32C    │ N-byte protobuf HistoryEvent │
└───────────┴──────────────┴────────────────────────────┘
  • Length prefix — lets records be read back sequentially.
  • CRC32C — detects bit rot and torn writes.
  • Torn tail — a partial record from a crash mid-append is truncated on the next open, so the log always resumes on a clean boundary. A checksum failure in the middle of a log is a hard error, never silently skipped.
  • Directory fsync — a new workflow's file is fsync'd into its parent directory, so the file itself can't vanish in a power loss.

Event types recorded today:

  • WorkflowStarted, WorkflowCompleted, WorkflowFailed
  • StepScheduled, StepCompleted, StepFailed
  • TimerStarted, TimerFired, SignalReceived

Replay

To decide what a workflow should do next, the engine re-runs your workflow function from the top. Each durable call it makes is matched against history:

  • Step already completed → the call returns the recorded result instantly. The real work is not repeated.
  • Step still in flight → the function is suspended; the engine waits.
  • Step not in history → this is the next action. The engine records it and dispatches it.

So your function runs many times, but each side effect happens exactly once. This is why workflow code must be deterministic — replay must take the same path every time.

Determinism rules

Breaking these makes a workflow diverge from its recorded history. The engine detects step-level divergence and refuses to proceed rather than corrupting the run.
Don'tDoWhy
time.Now()ctx.Sleep(...), or pass time in as activity outputWall clock differs on every replay
rand, uuid.New()Generate inside an activityNew value each replay
HTTP calls, DB queries, file I/Octx.ExecuteActivity(...)Side effects would repeat on every replay
go func(){}, channelsSequential code (parallel steps land in Phase 5)Goroutine scheduling is non-deterministic
Reading global mutable statePass data through step inputs/outputsProcess memory is gone after a restart
Reordering or renaming steps in a running deploymentVersion your workflow (Phase 5)In-flight histories won't match the new code

Also avoid defer with side effects inside workflow functions — the deferred call runs on every replay pass.

Steps: activity, timer, signal

Three kinds of step, all recorded in the same history:

KindCreated byCompleted by
ACTIVITYctx.ExecuteActivityA worker reporting a result over gRPC
TIMERctx.SleepThe engine's timer service when the deadline passes
SIGNALctx.WaitForSignalAn external SignalWorkflow call

A timer's absolute deadline is written to disk, so a sleep of 24 hours survives any number of restarts — on recovery the timer is re-armed against the recorded deadline (and fires immediately if it already elapsed while the engine was down).

Writing workflows

A workflow is a Go function matching sdk.WorkflowFunc. It interacts with the outside world only through the sdk.Context it is given.

The Context API

type Context interface {
	// Run a named activity on a worker and wait for its result.
	// A failed activity returns an error you can handle or propagate.
	ExecuteActivity(name string, input []byte) ([]byte, error)

	// Pause durably for d. Survives restarts.
	Sleep(name string, d time.Duration)

	// Block until a matching external signal arrives; returns its payload.
	// A signal that arrived earlier is buffered and returned immediately.
	WaitForSignal(name, signal string) []byte

	// The workflow's start input.
	Input() []byte
}

A complete example

func OrderWorkflow(ctx sdk.Context, input []byte) ([]byte, error) {
	// 1. Reserve stock. If this fails, the workflow fails — nothing to undo.
	reserved, err := ctx.ExecuteActivity("reserve", input)
	if err != nil {
		return nil, err
	}

	// 2. Wait for a human to approve. Could be seconds or days.
	approval := ctx.WaitForSignal("await-approval", "approved")
	if string(approval) != "yes" {
		ctx.ExecuteActivity("release", reserved)
		return nil, errors.New("order rejected by reviewer")
	}

	// 3. Give the payment processor a moment to settle.
	ctx.Sleep("settle", 2*time.Second)

	// 4. Charge. Retries are automatic (see Retries below); we only see an
	//    error once the policy is exhausted — so compensate and fail.
	if _, err := ctx.ExecuteActivity("charge", reserved); err != nil {
		ctx.ExecuteActivity("release", reserved)
		return nil, fmt.Errorf("charge failed, reservation released: %w", err)
	}

	return ctx.ExecuteActivity("ship", reserved)
}
Step names matter. Each call's name is recorded in history and checked on replay. Names must be stable and unique within a workflow — they are how the engine lines your code up against the durable record.

Registering & running your workflow

Today the engine is embedded, not a library. Go's internal/ rule means packages under internal/ cannot be imported from outside this module, so you register workflows by building the daemon from this repo. A properly importable client/worker package is the Phase 5 deliverable.

Register your workflow in cmd/workflowd/main.go, in runServe, before the engine starts serving:

engine := server.NewEngine(store, logger)
defer engine.Close()

// Register a code-based workflow.
engine.RegisterWorkflow("OrderWorkflow", sdk.Workflow{
	Type: "OrderWorkflow",
	Fn:   OrderWorkflow,
})

// Optional: retry policy for a specific activity, by step name.
engine.SetRetryPolicy("charge", server.RetryPolicy{
	MaxAttempts: 3,
	Backoff:     200 * time.Millisecond,
	MaxBackoff:  5 * time.Second,
})

// Rebuild state from disk before accepting traffic.
if err := engine.Recover(); err != nil {
	return err
}

Then rebuild and run: go build -o workflowd ./cmd/workflowd && ./workflowd --serve.

Writing a worker

A worker is any process that speaks the gRPC service. It long-polls for a task, does the real work, and reports the result. Workers are stateless — all replay happens on the engine — so you can run as many as you like.

From outside this repo, generate your own stubs from proto/service.proto (the proto file is the stable contract). The loop looks like this:

for {
	// Long-poll. Blocks until the engine has work.
	resp, err := client.PollTask(ctx, &pb.PollTaskRequest{WorkerId: "worker-1"})
	if err != nil {
		return // context cancelled or engine gone
	}
	task := resp.GetTask()

	// Do the actual work, dispatching on the step name.
	output, execErr := run(task.GetStepName(), task.GetInput())

	req := &pb.CompleteTaskRequest{
		WorkflowId: task.GetWorkflowId(),
		StepName:   task.GetStepName(),
		Sequence:   task.GetSequence(),
		LeaseToken: task.GetLeaseToken(), // proof you still own the task
	}
	if execErr != nil {
		req.Result = &pb.CompleteTaskRequest_Error{Error: execErr.Error()}
	} else {
		req.Result = &pb.CompleteTaskRequest_Output{Output: output}
	}
	client.CompleteTask(ctx, req)
}

Long-running activities must heartbeat

Every dispatched task carries a lease_duration_ms. If your activity will run longer than that, call Heartbeat periodically (a third of the lease duration is a sane interval). A rejected heartbeat means you lost the lease — stop working and discard the result, because the task has been given to someone else.

Retries & Saga compensation

Retries are handled by the engine and are invisible to your workflow. Register a policy per activity step name:

engine.SetRetryPolicy("charge", server.RetryPolicy{
	MaxAttempts: 3,                      // counts the first try; 1 = no retry
	Backoff:     200 * time.Millisecond, // doubles each attempt
	MaxBackoff:  5 * time.Second,        // ceiling
})

The default is no retry (MaxAttempts: 1) — a failure surfaces to the workflow immediately. Retries reuse the same step sequence, so a retried activity stays one logical step with an attempts count (visible in the dashboard). Backoffs are durable: if the engine restarts during one, the retry is re-armed.

Saga compensation needs no special machinery. Once retries are exhausted, the activity error is returned to your Go code — so "undo" is just an ordinary error branch that calls more activities, as in the release call in the example above.

Polyglot workflows

Workflow code does not have to be Go. The engine owns durability — history, replay position, timers, retries, leases — while an SDK in your language owns the orchestration logic. They talk over a small protocol.

SDKLocationStatus
Gosdk/In-process (compiled into the engine)
TypeScript / Nodesdks/typescript/Verified end-to-end
Pythonsdks/python/Verified end-to-end
Javasdks/java/Written but not yet compiled — reference implementation

Every non-Go SDK is dependency-free: Node's fetch, Python's urllib and Java's HttpClient are all built in. No npm install, no pip install, no protobuf toolchain.

Activities were already language-agnostic. Any process that speaks the API can execute activities — that has been true since Week 3. What the SDKs add is running the workflow function itself outside the engine.

TypeScript

import { Worker, WorkflowFailure } from './workflowd.mjs';

const worker = new Worker({ address: 'http://localhost:8233' });

worker.workflow('OrderWorkflow', async (ctx) => {
  const reserved = await ctx.activity('reserve', ctx.input);

  await ctx.sleep('settle', 2000);                    // durable
  const ok = await ctx.waitForSignal('gate', 'approved');
  if (ok !== 'yes') throw new WorkflowFailure('rejected');

  try {
    await ctx.activity('charge', reserved);
  } catch (err) {
    await ctx.activity('release', reserved);          // compensate
    throw new WorkflowFailure(`charge failed: ${err.message}`);
  }
  return await ctx.activity('ship', reserved);
});

worker.activity('reserve', async (input) => `reserved:${input}`);
await worker.run();

Run it with node example.mjs — no install step.

Python

from workflowd import Worker, WorkflowFailure, ActivityFailure

worker = Worker(address="http://localhost:8233")

@worker.workflow("OrderWorkflow")
def order(ctx):
    reserved = ctx.activity("reserve", ctx.input)

    ctx.sleep("settle", 2.0)                          # durable
    if ctx.wait_for_signal("gate", "approved") != "yes":
        raise WorkflowFailure("rejected")

    try:
        ctx.activity("charge", reserved)
    except ActivityFailure as err:
        ctx.activity("release", reserved)             # compensate
        raise WorkflowFailure(f"charge failed: {err}")
    return ctx.activity("ship", reserved)

@worker.activity("reserve")
def reserve(payload):
    return f"reserved:{payload}"

worker.run()

The SDK protocol

Three HTTP endpoints. If your language isn't listed above, this is everything you need to add it — a complete SDK is about 250 lines.

POST /api/workflow-types            {"workflow_types":["OrderWorkflow"],"worker_id":"w1"}
GET  /api/workflow-tasks?types=OrderWorkflow    (long-poll; 204 = nothing yet)
POST /api/workflow-tasks/complete   {"workflow_id":...,"task_token":...,"command":{...}}

A workflow task carries the already-reconstructed state — input, steps in order with their results, and buffered signals — so your SDK never has to implement the history fold:

{
  "workflow_id": "order-42",
  "workflow_type": "OrderWorkflow",
  "task_token": "9f3c...",
  "history_length": 4,
  "input": "order-42",
  "steps": [
    {"step_name":"reserve","sequence":1,"kind":"ACTIVITY","status":"COMPLETED","output":"reserved:order-42"},
    {"step_name":"charge","sequence":2,"kind":"ACTIVITY","status":"SCHEDULED","attempts":2}
  ],
  "signals": [{"name":"approved","payload":"yes"}]
}

Replay your workflow function against steps in call order, then return one command:

{"type":"schedule_activity","step_name":"charge","sequence":2,"input":"..."}
{"type":"start_timer","step_name":"wait","sequence":3,"delay_ms":60000}
{"type":"await_signal","step_name":"gate","sequence":4,"signal_name":"approved"}
{"type":"complete_workflow","output":"..."}
{"type":"fail_workflow","error":"..."}
{"type":"noop"}

How replay suspends — the one hard part

When replay reaches a step with no recorded result, your workflow function must stop without workflow code being able to intercept it. This matters more than it sounds, because workflows legitimately write try { await activity() } catch { compensate() }. If your suspend signal is an ordinary exception, that catch swallows it and compensation runs on every replay pass.

LanguageMechanismWhy it's safe
Gopanic / recoverRecovered only in the SDK's own frame
Pythonsubclass BaseExceptionexcept Exception doesn't catch it
Javasubclass Errorcatch (Exception) doesn't catch it
TypeScripta promise that never settlescatch catches everything in JS, so suspension must not be a throw at all
The TypeScript case is not hypothetical: the first version of that SDK used a thrown sentinel, and a workflow with a try/catch around an activity silently stalled. It now suspends with a non-settling promise, which workflow code cannot observe.

Safety properties you get for free

  • SDK workers are stateless. Every task carries everything needed to replay, so a crashed SDK worker costs nothing — the task is re-dispatched.
  • Stale decisions are fenced. Each task records the history length it was based on. If history moved on while your worker was thinking, the command is rejected with 409 — poll again. This is what makes re-dispatch safe: two workers may both decide, but only one command can ever apply.
  • Malformed commands can't corrupt a workflow. The engine validates step sequencing before anything reaches the log, and a rejected command leaves the task pending so a fixed deployment heals the workflow automatically.
  • Replay errors fail loudly. If your SDK reports that it can't replay a history (the code changed incompatibly), the workflow fails with that message rather than spinning forever.

CLI flags

workflowd --serve [flags]
FlagDefaultDescription
--servefalseRun the engine (gRPC + HTTP). Without it the daemon just bootstraps and exits.
--data-dirdataDirectory holding *.history files. This is your database — back it up.
--addr:7233gRPC listen address.
--http-addr:8233Dashboard + JSON API. Empty string disables.
--metrics-addr:9090Prometheus /metrics. Empty string disables.
--demofalseDrive a sample workflow to completion against a real store, then exit.
--auth-tokenemptyRequire this bearer token on mutating /api routes (Authorization: Bearer <token>). Empty disables auth. /api/health and the dashboard's static assets stay open.
--version—Print version and exit.
Auth is optional and there is no TLS. Without --auth-token, anyone who can reach these ports can start, signal and terminate workflows. With it, the JSON API requires the token but gRPC and metrics remain unauthenticated. Bind everything to localhost or a private network either way.

gRPC API

Service workflow.v1.WorkflowService, defined in proto/service.proto.

RPCPurpose
StartWorkflowBegin a new execution. Takes workflow_type, input, optional workflow_id (generated when blank; a supplied id must not already exist).
PollTaskLong-poll for the next activity task. Returns a Task with a lease_token and lease_duration_ms.
CompleteTaskReport a result. Requires the current lease_token; sets exactly one of output or error.
HeartbeatExtend a lease on a long-running task. A rejection means the lease was lost.
SignalWorkflowDeliver an external event. Buffered if the workflow isn't waiting yet; rejected for unknown or terminal workflows.
ListWorkflowsSummaries of all workflows, with an optional status filter.
GetWorkflowOne workflow in full, including every step.
TerminateWorkflowOperator kill switch. Durably fails the workflow and fences any worker still holding its task.

Status codes

CodeMeaning
INVALID_ARGUMENTUnknown workflow type, or a duplicate workflow id.
FAILED_PRECONDITIONStale/expired lease, result doesn't match the in-flight step, or the workflow is already terminal.
NOT_FOUNDNo such workflow.

HTTP / JSON API

Served on --http-addr (default :8233) alongside the dashboard. Browsers can't speak raw gRPC, so this is the integration point for UIs, scripts and curl. Status values are the bare names RUNNING, WAITING, COMPLETED, FAILED, TERMINATED, TIMED_OUT, or UNKNOWN (history that could not be replayed). Timestamps are RFC 3339.

List workflows

Supports a status filter, a case-insensitive substring search over workflow id and type (q), and offset/limit pagination. The response carries total (matches after status+q filtering, before pagination) so clients can page through.

GET /api/workflows
GET /api/workflows?status=RUNNING
GET /api/workflows?q=order-42&limit=50&offset=100
{
  "total": 1,
  "offset": 0,
  "limit": 200,
  "workflows": [
    {
      "workflow_id": "OrderSaga-1752...-a1b2c3d4",
      "workflow_type": "OrderSaga",
      "status": "WAITING",
      "started_at": "2026-07-13T00:35:06.915Z",
      "updated_at": "2026-07-13T00:35:07.002Z",
      "history_length": 3
    }
  ]
}

Get one workflow

GET /api/workflows/{id}
{
  "workflow_id": "order-42",
  "workflow_type": "OrderSaga",
  "status": "WAITING",
  "history_length": 5,
  "input": "order-42",
  "steps": [
    {
      "step_name": "reserve", "sequence": 1, "kind": "ACTIVITY",
      "status": "COMPLETED", "attempts": 1, "output": "reserved:order-42"
    },
    {
      "step_name": "charge", "sequence": 2, "kind": "ACTIVITY",
      "status": "SCHEDULED", "attempts": 2
    }
  ]
}

Start a workflow

POST /api/workflows
{"type":"OrderSaga","input":"order-42","workflow_id":"optional"}
→ 201 {"workflow_id":"order-42"}

A duplicate workflow_id returns 409; an unregistered type returns 400.

Signal a workflow

POST /api/workflows/{id}/signal
{"name":"approved","payload":"yes"}
→ 200 {"status":"delivered"}

Terminate a workflow

POST /api/workflows/{id}/terminate
{"reason":"stuck on a bad payload"}
→ 200 {"status":"terminated"}

Health

GET /api/health → 200 {"status":"ok","version":"...","uptime_seconds":1234,
                        "workflows_active":2,"task_queue_depth":0,
                        "workflow_task_queue_depth":0}

/api/health never requires the auth token, so liveness probes work unconfigured.

Errors

Failures return {"error": "..."} with a matching status code:

CodeWhen
400Malformed JSON, unknown field, unknown workflow type, bad status filter, empty signal name
401Missing or wrong bearer token (when --auth-token is set)
404No such workflow
409Workflow already terminal (double terminate, signal after finish), or duplicate workflow id on start
500Unreadable or structurally invalid history

Metrics

Prometheus text format at --metrics-addr (default :9090/metrics). The registry is hand-rolled, so there is no client_golang dependency.

MetricTypeMeaning
workflowd_workflows_started_totalcounterExecutions started
workflowd_workflows_finished_total{status}counterTerminal outcomes: completed, failed, terminated
workflowd_workflows_activegaugeRunning or waiting right now
workflowd_tasks_dispatched_totalcounterTasks handed to workers (including retries)
workflowd_task_results_total{result}counterWorker-reported results: completed / failed
workflowd_leases_expired_totalcounterTasks reclaimed from dead workers
workflowd_retries_scheduled_totalcounterRetries armed after a failed attempt
workflowd_signals_received_totalcounterSignals durably recorded
workflowd_timers_fired_totalcounterDurable timers that reached their deadline
workflowd_task_queue_depthgaugeTasks buffered awaiting a worker
workflowd_step_duration_secondshistogramDispatch → worker result latency

Useful queries

# Workflow failure rate over 5 minutes
rate(workflowd_workflows_finished_total{status="failed"}[5m])

# Are workers keeping up? (a rising queue means add workers)
workflowd_task_queue_depth

# 95th percentile activity latency
histogram_quantile(0.95, rate(workflowd_step_duration_seconds_bucket[5m]))

# Dead workers
rate(workflowd_leases_expired_total[5m]) > 0

Dashboard

Served at --http-addr and compiled into the binary with go:embed. It gives you:

  • counts of total / active / completed / failed workflows,
  • a filterable workflow list that auto-refreshes every 2 seconds,
  • client-side search over workflow ids and types, plus per-status counts on the filter chips,
  • a dialog to start a workflow directly from the browser (type, optional id, input),
  • a detail view with the full step timeline rendered as woven weft picks — kinds, statuses, attempt counts, per-step errors and outputs, and which signal a step is parked on,
  • actions to send a signal to, or terminate, a live workflow.

It is plain HTML, CSS and vanilla JS — no build step, no npm, nothing to bundle.

Crash recovery

On startup (before serving traffic) the engine scans the data directory and, for each workflow:

  • Terminal or not started — nothing to do.
  • Waiting on an activity — the task is re-queued for a worker.
  • Waiting on a timer — re-armed against the recorded deadline (fires immediately if it already passed).
  • Waiting on a signal — any signal buffered before the crash is delivered.
  • Mid-decision — the decision loop advances one step.
  • Failed with retries remaining — the backoff timer is re-armed.
At-least-once, not exactly-once. If a worker finished an activity but died before reporting the result, that activity runs again after recovery. Make your activities idempotent — use the workflow id plus step name as an idempotency key.

Leases & dead workers

When a task is dispatched, the worker receives a random lease token and a deadline (30s by default). CompleteTask and Heartbeat must present that token.

  • Miss the deadline and the engine records StepFailed("lease expired"), then feeds it to the normal retry machinery — attempts remaining means backoff and reassign; exhausted means the workflow sees the error and can compensate.
  • A "zombie" worker that comes back late has a stale token and is rejected, so two workers can never both record a result for the same attempt.
  • Because a dead worker counts as a failed attempt, a task that kills every worker it touches can't loop forever.

With the default no-retry policy, one dead worker fails the workflow explicitly. Set MaxAttempts > 1 on activities that should tolerate worker crashes.

Limitations

Known and deliberate, as of Weeks 1–10:

  • Single node. No replication or failover yet — Raft is Phase 4. The data directory is a single point of failure; back it up.
  • No TLS on any port, and auth (when enabled) is a single shared bearer token on the JSON API only — gRPC and metrics are unauthenticated. Keep ports private.
  • Go workflows must be compiled into the engine binary (Go's internal/ rule). Workflows in TypeScript, Python and Java run out of process via the SDK protocol and have no such constraint; an importable Go SDK is Phase 5.
  • Sequential steps. One step in flight per workflow; no parallel fan-out yet.
  • At-least-once activity execution (see above).
  • No workflow versioning yet — changing a running workflow's step sequence breaks in-flight executions. workflow.Version is Phase 5.
  • Listing is a full scan of the data directory; fine for thousands of workflows, not for millions.
  • One global lock and a full history re-read per operation, which caps throughput. Per-workflow locks and a state cache are planned.
  • No cron scheduling yet.

FAQ

How is this different from a job queue?

A queue remembers tasks; this remembers your program's position. With a queue, a crash between "charge succeeded" and "enqueue ship" loses the thread of the process. Here that transition is an event in a durable log, so restarting resumes exactly at the next step.

Do I need a database?

No. The data directory is the database. Each workflow is one append-only file.

What happens if I redeploy while workflows are running?

Workflows resume from history. That is safe as long as the step sequence your code produces hasn't changed. Adding a step in the middle, renaming one, or reordering will be detected as non-determinism and the affected executions will refuse to advance. Until versioning lands, drain in-flight workflows before changing an existing workflow's shape.

How many workers should I run?

Watch workflowd_task_queue_depth. If it grows, add workers. Workers are stateless, so scaling is just running more processes pointed at the same engine.

Can I use it from another language?

Yes — both halves. Activities have always been language-agnostic (any process that speaks the API can execute them). Workflow definitions can now also run outside the engine: TypeScript and Python SDKs ship in sdks/, with a Java reference implementation, and the protocol is three HTTP endpoints if you want to add another language.

What is the license?

Apache-2.0. The engine is open source and stays that way; a managed/hosted product is the intended commercial layer.

Where to go next

  • Source on GitHub — the docs/ folder holds the architecture notes, ADRs and per-session walkthroughs.
  • docs/ARCHITECTURE.md — component-by-component design.
  • docs/ROADMAP.md — the 17-week plan and what's checked off.
  • docs/walkthroughs/ — plain-English explanations of how each subsystem was built.