<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Temporal: Building Long-Running Systems in Software - General</title>
    <link>https://community.hpe.com/t5/software-general/temporal-building-long-running-systems/m-p/7272273#M1586</link>
    <description>&lt;P&gt;Most backend systems start simple. A request comes in, you do some work, you return a response. Then reality shows up:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;A payment must be charged, then a receipt emailed, then inventory reserved — and each step can fail independently.&lt;/LI&gt;&lt;LI&gt;An order must wait 3 days for the customer to confirm, and auto-cancel otherwise.&lt;/LI&gt;&lt;LI&gt;A data pipeline must process 10,000 records, resume from record 6,832 after a crash, and never double-process a record.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The usual answer is a pile of cron jobs, message queues, a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;status&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;column in a database, and retry loops. It works until a process gets OOM-killed halfway through step 3, and you spend the next two hours reconstructing what happened from logs.&lt;/P&gt;&lt;P&gt;Temporal exists to remove that entire category of work.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;1. What Temporal actually is&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Temporal is a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;durable execution&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;platform. You write your business logic as ordinary code — loops, conditionals, function calls,&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;sleep&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;— and Temporal guarantees that this code runs to completion&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;exactly once&lt;/STRONG&gt;, even if the process running it crashes, the machine dies, or the datacenter goes away for an hour.&lt;/P&gt;&lt;P&gt;The core trick: Temporal does not persist your program's&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;EM&gt;memory&lt;/EM&gt;. It persists the&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;history of everything that happened&lt;/STRONG&gt;. When your process dies and a new one picks up the work, Temporal replays that history through your function to rebuild its exact state, then continues from where it left off.&lt;/P&gt;&lt;P&gt;Two things follow from this, and almost everything else in Temporal is a consequence of them:&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;Your workflow code must be&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;deterministic&lt;/STRONG&gt;, because it gets replayed.&lt;/LI&gt;&lt;LI&gt;Anything non-deterministic (network calls, DB reads, random numbers, clock reads) must be moved out into&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;activities&lt;/STRONG&gt;, whose results are recorded in history and returned verbatim on replay.&lt;/LI&gt;&lt;/OL&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;2. The building blocks Workflow&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;The orchestrator. It decides&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;EM&gt;what&lt;/EM&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;happens and&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;EM&gt;in what order&lt;/EM&gt;. It is deterministic and must not perform I/O.&lt;/P&gt;&lt;PRE&gt;package app

import (
	"time"

	"go.temporal.io/sdk/temporal"
	"go.temporal.io/sdk/workflow"
)

type OrderInput struct {
	OrderID string
	UserID  string
	Amount  int64
}

func OrderWorkflow(ctx workflow.Context, in OrderInput) (string, error) {
	ao := workflow.ActivityOptions{
		StartToCloseTimeout: 30 * time.Second,
		RetryPolicy: &amp;amp;temporal.RetryPolicy{
			InitialInterval:    time.Second,
			BackoffCoefficient: 2.0,
			MaximumInterval:    time.Minute,
			MaximumAttempts:    5,
		},
	}
	ctx = workflow.WithActivityOptions(ctx, ao)

	var chargeID string
	if err := workflow.ExecuteActivity(ctx, ChargeCard, in.UserID, in.Amount).Get(ctx, &amp;amp;chargeID); err != nil {
		return "", err
	}

	if err := workflow.ExecuteActivity(ctx, ReserveInventory, in.OrderID).Get(ctx, nil); err != nil {
		// compensate: the money is already taken
		_ = workflow.ExecuteActivity(ctx, RefundCharge, chargeID).Get(ctx, nil)
		return "", err
	}

	if err := workflow.ExecuteActivity(ctx, SendReceipt, in.UserID, chargeID).Get(ctx, nil); err != nil {
		return "", err
	}

	return chargeID, nil
}&lt;/PRE&gt;&lt;P&gt;Read that again: it looks like a plain Go function. There is no state machine, no&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;status&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;column, no retry loop, no resume logic. If the worker running this dies between&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;ChargeCard&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;and&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;ReserveInventory, another worker picks it up, replays history, sees&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;ChargeCard&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;already returned&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;chargeID, and proceeds straight to&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;ReserveInventory. The card is not charged twice.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Activity&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The part that touches the outside world. Activities may be non-deterministic, may fail, and are automatically retried according to their retry policy.&lt;/P&gt;&lt;PRE&gt;package app

import (
	"context"
	"fmt"
)

func ChargeCard(ctx context.Context, userID string, amount int64) (string, error) {
	resp, err := paymentGateway.Charge(ctx, userID, amount)
	if err != nil {
		return "", fmt.Errorf("charge failed: %w", err)
	}
	return resp.ChargeID, nil
}

func ReserveInventory(ctx context.Context, orderID string) error {
	return inventory.Reserve(ctx, orderID)
}&lt;/PRE&gt;&lt;P&gt;Activities receive a normal&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;context.Context. It is cancelled when the activity is cancelled or its timeout fires, so pass it down to your HTTP and DB calls.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Worker&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The process that hosts your code and polls Temporal for work. It is just a long-running Go binary.&lt;/P&gt;&lt;PRE&gt;package main

import (
	"log"

	"go.temporal.io/sdk/client"
	"go.temporal.io/sdk/worker"

	"example.com/app"
)

func main() {
	c, err := client.Dial(client.Options{
		HostPort:  client.DefaultHostPort, // localhost:7233
		Namespace: "default",
	})
	if err != nil {
		log.Fatalln("dial:", err)
	}
	defer c.Close()

	w := worker.New(c, "orders", worker.Options{})

	w.RegisterWorkflow(app.OrderWorkflow)
	w.RegisterActivity(app.ChargeCard)
	w.RegisterActivity(app.ReserveInventory)
	w.RegisterActivity(app.RefundCharge)
	w.RegisterActivity(app.SendReceipt)

	if err := w.Run(worker.InterruptCh()); err != nil {
		log.Fatalln("worker:", err)
	}
}&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;Starting a workflow&lt;/STRONG&gt;&lt;/P&gt;&lt;PRE&gt;opts := client.StartWorkflowOptions{
	ID:        "order-" + in.OrderID, // your idempotency key
	TaskQueue: "orders",
}

run, err := c.ExecuteWorkflow(ctx, opts, app.OrderWorkflow, in)
if err != nil {
	return err
}

var chargeID string
err = run.Get(ctx, &amp;amp;chargeID) // blocks until the workflow completes&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;The workflow ID is your deduplication mechanism.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Calling&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;ExecuteWorkflow&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;twice with the same ID while the first is still running returns a handle to the existing run rather than starting a second one. Derive it from a business key — order ID, user ID, invoice number — never from a UUID generated at call time.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;3. Task queues and namespaces&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Task queue&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;is the routing mechanism. Workers poll a named queue; workflows and activities are dispatched to a queue. That's it — there is no broker configuration, no exchange, no topic.&lt;/P&gt;&lt;P&gt;This gives you useful control for free:&lt;/P&gt;&lt;PRE&gt;// Route heavy GPU/CPU work to a dedicated fleet
ctx = workflow.WithActivityOptions(ctx, workflow.ActivityOptions{
	TaskQueue:           "video-encoding",
	StartToCloseTimeout: 2 * time.Hour,
})&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;Namespace&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;is the isolation boundary — think of it as a virtual cluster. Use separate namespaces for&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;prod,&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;staging, and per-team workloads. Retention, archival, and access controls are configured per namespace.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;4. Timeouts: the four that matter&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;This is where most people get burned, so be explicit about all of them.&lt;/P&gt;&lt;P&gt;Timeout Meaning Guidance ScheduleToStartTimeout How long a task may sit in the queue before a worker picks it up Usually leave unset; a non-zero value here is a capacity alarm, not a correctness control StartToCloseTimeout Max duration of a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;single attempt&lt;/STRONG&gt; &lt;STRONG&gt;Always set this.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Make it slightly larger than your realistic p99 ScheduleToCloseTimeout Total wall-clock budget across all retries Set when there is a real business deadline HeartbeatTimeout Max gap between heartbeats for long activities Required for anything that runs more than ~a minute&lt;/P&gt;&lt;P&gt;And on the workflow itself:&lt;/P&gt;&lt;PRE&gt;opts := client.StartWorkflowOptions{
	ID:                       "order-123",
	TaskQueue:                "orders",
	WorkflowExecutionTimeout: 7 * 24 * time.Hour, // total lifetime, retries included
	WorkflowRunTimeout:       24 * time.Hour,     // single run
	WorkflowTaskTimeout:      10 * time.Second,   // one decision step
}&lt;/PRE&gt;&lt;HR /&gt;&lt;P&gt;&lt;STRONG&gt;&lt;FONT size="4"&gt;5. Retries and error semantics&lt;/FONT&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Activities retry by default (infinitely, with exponential backoff, unless you say otherwise). Workflows do&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;not&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;retry by default — they are already durable, so a workflow "failing" means your business logic decided to fail.&lt;/P&gt;&lt;P&gt;Some errors should never be retried. A&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;400 Bad Request&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;will still be a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;400&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;on the fifth attempt.&lt;/P&gt;&lt;PRE&gt;// Mark an error as non-retryable from inside an activity
return temporal.NewNonRetryableApplicationError(
	"card declined", "CardDeclined", nil,
)&lt;/PRE&gt;&lt;PRE&gt;// Or exclude by type in the retry policy
RetryPolicy: &amp;amp;temporal.RetryPolicy{
	MaximumAttempts:        5,
	NonRetryableErrorTypes: []string{"CardDeclined", "InvalidInput"},
}&lt;/PRE&gt;&lt;P&gt;Handling failures in the workflow:&lt;/P&gt;&lt;PRE&gt;var appErr *temporal.ApplicationError
err := workflow.ExecuteActivity(ctx, ChargeCard, in.UserID, in.Amount).Get(ctx, &amp;amp;chargeID)
if errors.As(err, &amp;amp;appErr) &amp;amp;&amp;amp; appErr.Type() == "CardDeclined" {
	return workflow.ExecuteActivity(ctx, NotifyDeclined, in.UserID).Get(ctx, nil)
}&lt;/PRE&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;&lt;STRONG&gt;Retries make at-least-once the default.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Design activities to be idempotent. Pass an idempotency key (ActivityInfo.WorkflowExecution.ID&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;+ activity ID is a good one) to any external API that supports it.&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;6. Heartbeats and long-running activities&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;If an activity takes minutes or hours, heartbeat so Temporal can detect a dead worker quickly instead of waiting out the full&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;StartToCloseTimeout. Heartbeats can also carry progress, which lets a retry resume instead of restart.&lt;/P&gt;&lt;PRE&gt;func ProcessRecords(ctx context.Context, batchID string) error {
	start := 0
	if activity.HasHeartbeatDetails(ctx) {
		_ = activity.GetHeartbeatDetails(ctx, &amp;amp;start) // resume where we died
	}

	records, err := loadRecords(ctx, batchID)
	if err != nil {
		return err
	}

	for i := start; i &amp;lt; len(records); i++ {
		if err := handle(ctx, records[i]); err != nil {
			return err
		}
		activity.RecordHeartbeat(ctx, i+1)

		select {
		case &amp;lt;-ctx.Done():
			return ctx.Err() // cancelled or timed out
		default:
		}
	}
	return nil
}&lt;/PRE&gt;&lt;P&gt;Set&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;HeartbeatTimeout&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;in the activity options or the heartbeats are ignored.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;7. Timers: sleeping for days is normal&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;PRE&gt;// Wait three days. Costs nothing while waiting — no thread, no memory, no cron.
if err := workflow.Sleep(ctx, 72*time.Hour); err != nil {
	return err
}&lt;/PRE&gt;&lt;P&gt;A sleeping workflow is not resident in any process. Temporal stores a timer and re-dispatches the workflow when it fires. Sleeping for a year is as cheap as sleeping for a second.&lt;/P&gt;&lt;P&gt;Never use&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;time.Sleep,&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;time.Now(),&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;rand,&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;os.Getenv, or map iteration order in workflow code. Use the deterministic equivalents:&lt;/P&gt;&lt;PRE&gt;now := workflow.Now(ctx)
id := workflow.GetInfo(ctx).WorkflowExecution.ID

var uuid string
_ = workflow.SideEffect(ctx, func(ctx workflow.Context) interface{} {
	return newUUID()
}).Get(&amp;amp;uuid)&lt;/PRE&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;8. Signals: sending data into a running workflow&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Signals are how the outside world talks to a workflow in flight.&lt;/P&gt;&lt;PRE&gt;func SubscriptionWorkflow(ctx workflow.Context, userID string) error {
	cancelCh := workflow.GetSignalChannel(ctx, "cancel")
	cancelled := false

	for month := 0; month &amp;lt; 12; month++ {
		timerCtx, cancelTimer := workflow.WithCancel(ctx)
		timer := workflow.NewTimer(timerCtx, 30*24*time.Hour)

		sel := workflow.NewSelector(ctx)
		sel.AddFuture(timer, func(workflow.Future) {})
		sel.AddReceive(cancelCh, func(c workflow.ReceiveChannel, _ bool) {
			c.Receive(ctx, nil)
			cancelled = true
			cancelTimer()
		})
		sel.Select(ctx)

		if cancelled {
			return workflow.ExecuteActivity(ctx, SendCancellationEmail, userID).Get(ctx, nil)
		}
		if err := workflow.ExecuteActivity(ctx, ChargeMonthly, userID).Get(ctx, nil); err != nil {
			return err
		}
	}
	return nil
}&lt;/PRE&gt;&lt;P&gt;From a client:&lt;/P&gt;&lt;PRE&gt;err := c.SignalWorkflow(ctx, "subscription-user-42", "", "cancel", nil)&lt;/PRE&gt;&lt;P&gt;SignalWithStartWorkflow&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;starts the workflow if it isn't running and signals it if it is — the standard pattern for "append to a session that may or may not exist yet".&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Human-in-the-loop&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;falls out of this naturally:&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;workflow.Await&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;on an approval signal with a timeout, and escalate if nobody responds.&lt;/P&gt;&lt;PRE&gt;approved := false
approvalCh := workflow.GetSignalChannel(ctx, "approval")
workflow.Go(ctx, func(ctx workflow.Context) {
	var decision bool
	approvalCh.Receive(ctx, &amp;amp;decision)
	approved = decision
})

ok, _ := workflow.AwaitWithTimeout(ctx, 48*time.Hour, func() bool { return approved })
if !ok {
	return workflow.ExecuteActivity(ctx, EscalateToManager, in).Get(ctx, nil)
}&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;Queries: reading state out&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Queries are read-only, synchronous, and must not mutate state or schedule work.&lt;/P&gt;&lt;PRE&gt;func OrderWorkflow(ctx workflow.Context, in OrderInput) error {
	status := "pending"
	if err := workflow.SetQueryHandler(ctx, "status", func() (string, error) {
		return status, nil
	}); err != nil {
		return err
	}
	// ... status = "charged" ... status = "shipped" ...
	return nil
}&lt;/PRE&gt;&lt;PRE&gt;resp, _ := c.QueryWorkflow(ctx, "order-123", "", "status")
var status string
_ = resp.Get(&amp;amp;status)&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;Updates: signal + query in one round trip&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;An update sends data in&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;and&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;returns a result, with optional validation that rejects bad input without writing to history.&lt;/P&gt;&lt;PRE&gt;_ = workflow.SetUpdateHandlerWithOptions(ctx, "addItem",
	func(ctx workflow.Context, item Item) (int, error) {
		cart = append(cart, item)
		return len(cart), nil
	},
	workflow.UpdateHandlerOptions{
		Validator: func(ctx workflow.Context, item Item) error {
			if item.Qty &amp;lt;= 0 {
				return errors.New("quantity must be positive")
			}
			return nil
		},
	},
)&lt;/PRE&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;9. Child workflows and fan-out&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Use a child workflow when a sub-process deserves its own ID, history, retry policy, and lifecycle. Use an activity when it's just a unit of work.&lt;/P&gt;&lt;PRE&gt;// Sequential child
var res ShipmentResult
err := workflow.ExecuteChildWorkflow(ctx, ShipmentWorkflow, orderID).Get(ctx, &amp;amp;res)&lt;/PRE&gt;&lt;PRE&gt;// Parallel fan-out over activities
futures := make([]workflow.Future, 0, len(itemIDs))
for _, id := range itemIDs {
	futures = append(futures, workflow.ExecuteActivity(ctx, ReserveItem, id))
}
for _, f := range futures {
	if err := f.Get(ctx, nil); err != nil {
		return err
	}
}&lt;/PRE&gt;&lt;P&gt;ParentClosePolicy&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;controls what happens to children when the parent finishes — terminate them (default), abandon them, or request cancellation.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;10. Continue-As-New: keeping history bounded&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Event history is not infinite. Practically, keep it under a few thousand events. For workflows that loop forever (subscriptions, monitors, per-entity actors), atomically restart with fresh history:&lt;/P&gt;&lt;PRE&gt;func MonitorWorkflow(ctx workflow.Context, state State) error {
	for i := 0; i &amp;lt; 500; i++ {
		if err := workflow.Sleep(ctx, time.Hour); err != nil {
			return err
		}
		if err := workflow.ExecuteActivity(ctx, Poll, state.Target).Get(ctx, &amp;amp;state); err != nil {
			return err
		}
	}
	return workflow.NewContinueAsNewError(ctx, MonitorWorkflow, state)
}&lt;/PRE&gt;&lt;P&gt;The workflow ID stays the same; only the run ID changes. Signals received while continuing-as-new are carried over.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;11. Determinism and versioning&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Because history is replayed,&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;you cannot freely change workflow code that has in-flight executions&lt;/STRONG&gt;. Reordering activities, adding a step in the middle, or removing a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Sleep&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;will make replay diverge from history and produce a non-determinism error.&lt;/P&gt;&lt;P&gt;Two safe options:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Patching&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;— for incremental changes:&lt;/P&gt;&lt;PRE&gt;if workflow.GetVersion(ctx, "add-fraud-check", workflow.DefaultVersion, 1) == workflow.DefaultVersion {
	// old path: existing runs continue as before
} else {
	// new path: only new runs take this
	if err := workflow.ExecuteActivity(ctx, FraudCheck, in).Get(ctx, nil); err != nil {
		return err
	}
}&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;Worker Versioning / new workflow type&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;— for large rewrites: pin old runs to old workers and route new runs to the new build.&lt;/P&gt;&lt;P&gt;Safe changes that never need a patch: activity implementation bodies, timeout and retry policy values, and log statements.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;12. Testing&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Temporal's Go SDK ships a test framework that runs workflows in-process with a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;skipping clock&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;— a workflow that sleeps 30 days finishes in milliseconds.&lt;/P&gt;&lt;PRE&gt;func TestOrderWorkflow(t *testing.T) {
	var s testsuite.WorkflowTestSuite
	env := s.NewTestWorkflowEnvironment()

	env.OnActivity(ChargeCard, mock.Anything, "u1", int64(500)).Return("ch_1", nil)
	env.OnActivity(ReserveInventory, mock.Anything, "o1").Return(nil)
	env.OnActivity(SendReceipt, mock.Anything, "u1", "ch_1").Return(nil)

	env.ExecuteWorkflow(OrderWorkflow, OrderInput{OrderID: "o1", UserID: "u1", Amount: 500})

	require.True(t, env.IsWorkflowCompleted())
	require.NoError(t, env.GetWorkflowError())

	var chargeID string
	require.NoError(t, env.GetWorkflowResult(&amp;amp;chargeID))
	require.Equal(t, "ch_1", chargeID)
	env.AssertExpectations(t)
}&lt;/PRE&gt;&lt;P&gt;Test compensation paths by returning errors from mocks, and test signals with&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;env.RegisterDelayedCallback.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Replay tests&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;are your safety net against accidental non-determinism: download the history of a production workflow and replay it against your new code in CI.&lt;/P&gt;&lt;PRE&gt;replayer := worker.NewWorkflowReplayer()
replayer.RegisterWorkflow(OrderWorkflow)
err := replayer.ReplayWorkflowHistoryFromJSONFile(nil, "testdata/order_history.json")
require.NoError(t, err)&lt;/PRE&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;13. Schedules and cron&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Recurring work is first-class — no external cron, no missed-fire mystery.&lt;/P&gt;&lt;PRE&gt;_, err := c.ScheduleClient().Create(ctx, client.ScheduleOptions{
	ID: "nightly-reconciliation",
	Spec: client.ScheduleSpec{
		CronExpressions: []string{"0 2 * * *"},
		TimeZoneName:    "Asia/Kolkata",
	},
	Action: &amp;amp;client.ScheduleWorkflowAction{
		ID:        "reconcile",
		Workflow:  ReconcileWorkflow,
		TaskQueue: "batch",
	},
	Overlap: enumspb.SCHEDULE_OVERLAP_POLICY_SKIP,
})&lt;/PRE&gt;&lt;P&gt;Schedules can be paused, backfilled, and triggered manually. The overlap policy answers "what if the previous run is still going?" — a question plain cron never asks.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;14. Search attributes and visibility&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Beyond fetching a workflow by ID, you can index and query workflows by business fields.&lt;/P&gt;&lt;PRE&gt;_ = workflow.UpsertSearchAttributes(ctx, map[string]interface{}{
	"CustomerId":  in.UserID,
	"OrderStatus": "charged",
})&lt;/PRE&gt;&lt;PRE&gt;resp, _ := c.ListWorkflow(ctx, &amp;amp;workflowservice.ListWorkflowExecutionsRequest{
	Query: `WorkflowType = 'OrderWorkflow' AND OrderStatus = 'charged' AND ExecutionStatus = 'Running'`,
})&lt;/PRE&gt;&lt;P&gt;This is how you answer "which orders are stuck in payment right now?" without adding a reporting table.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;15. Cancellation and cleanup&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Cancellation is cooperative and propagates to running activities and child workflows. Cleanup must run in a disconnected context, since the normal one is already cancelled.&lt;/P&gt;&lt;PRE&gt;defer func() {
	if errors.Is(ctx.Err(), workflow.ErrCanceled) {
		newCtx, _ := workflow.NewDisconnectedContext(ctx)
		_ = workflow.ExecuteActivity(newCtx, ReleaseInventory, in.OrderID).Get(newCtx, nil)
	}
}()&lt;/PRE&gt;&lt;P&gt;Note the difference:&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;cancel&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;is graceful and lets the workflow run its cleanup;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;terminate&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;is immediate and runs nothing. Prefer cancel.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;16. Running it locally&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;PRE&gt;brew install temporal          # or: https://temporal.io/setup/install-temporal-cli
temporal server start-dev      # server on :7233, Web UI on http://localhost:8233&lt;/PRE&gt;&lt;PRE&gt;temporal workflow start \
  --task-queue orders \
  --type OrderWorkflow \
  --workflow-id order-123 \
  --input '{"OrderID":"o1","UserID":"u1","Amount":500}'

temporal workflow show   --workflow-id order-123
temporal workflow signal --workflow-id order-123 --name cancel
temporal workflow query  --workflow-id order-123 --type status&lt;/PRE&gt;&lt;P&gt;The Web UI shows the full event history of every execution — every activity attempt, every failure, every retry, with inputs and outputs. This alone replaces a large amount of custom logging and admin tooling.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;17. Production notes&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Payload size.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Inputs, outputs, and signals go into history. Keep them small (rule of thumb: under ~100 KB, hard limit 2 MB). Pass S3 keys or row IDs, not blobs. Use a Data Converter for encryption or compression.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Worker tuning.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;MaxConcurrentActivityExecutionSize&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;and&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;MaxConcurrentWorkflowTaskExecutionSize&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;are your throughput knobs. Run workflow and activity workers separately when activity work is heavy, so a saturated activity pool never starves workflow progress.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Retention.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Closed workflow histories are deleted after the namespace retention period (default 3 days on the dev server, commonly 30 days in production). Ship anything you need long-term to your own store.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Metrics.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;The SDK exports Prometheus metrics. Watch&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;temporal_workflow_task_schedule_to_start_latency&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;(workers under-provisioned) and&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;temporal_activity_execution_failed.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Sticky execution.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Workers cache workflow state so they don't replay full history on every task. A cache miss just means a replay — correct, but slower. Size&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;WorkflowCacheSize&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;accordingly.&lt;/LI&gt;&lt;/UL&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;18. When not to use Temporal&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Temporal is not free. It adds a server (plus Cassandra/PostgreSQL/MySQL and Elasticsearch), a programming model with real constraints, and per-step latency measured in milliseconds rather than microseconds.&lt;/P&gt;&lt;P&gt;Skip it when:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;The work is a single, fast, idempotent operation. Just do it inline.&lt;/LI&gt;&lt;LI&gt;You need sub-millisecond latency per step.&lt;/LI&gt;&lt;LI&gt;You need high-throughput stream processing — that's Kafka/Flink territory, not Temporal's.&lt;/LI&gt;&lt;LI&gt;The whole system is one service with one database and no multi-step failure modes worth modelling.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;Reach for it when:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;A process spans multiple services or multiple external APIs.&lt;/LI&gt;&lt;LI&gt;A process spans minutes, days, or months.&lt;/LI&gt;&lt;LI&gt;Partial failure leaves the system in a state someone has to fix by hand.&lt;/LI&gt;&lt;LI&gt;You need auditability of what happened, step by step.&lt;/LI&gt;&lt;LI&gt;You've already written a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;status&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;column, a retry table, and a reconciliation cron for the same flow.&lt;/LI&gt;&lt;/UL&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;19. The mental model to keep&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;&lt;STRONG&gt;Temporal turns a distributed, failure-prone, multi-step process into a single function you can read top to bottom.&lt;/STRONG&gt;&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;Everything else — the event history, replay, task queues, timers, heartbeats — is machinery in service of that one property. Once you internalise "workflows orchestrate and must be deterministic; activities do the work and may fail", the rest of the API stops feeling like magic and starts feeling obvious.&lt;/P&gt;&lt;P&gt;Start with one flow. Pick the one that currently has a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;status&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;column, a retry cron, and a runbook. Port it. You'll delete more code than you write.&lt;/P&gt;</description>
    <pubDate>Tue, 01 Sep 2026 11:08:22 GMT</pubDate>
    <dc:creator>Prithviraj-</dc:creator>
    <dc:date>2026-09-01T11:08:22Z</dc:date>
    <item>
      <title>Temporal: Building Long-Running Systems</title>
      <link>https://community.hpe.com/t5/software-general/temporal-building-long-running-systems/m-p/7272273#M1586</link>
      <description>&lt;P&gt;Most backend systems start simple. A request comes in, you do some work, you return a response. Then reality shows up:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;A payment must be charged, then a receipt emailed, then inventory reserved — and each step can fail independently.&lt;/LI&gt;&lt;LI&gt;An order must wait 3 days for the customer to confirm, and auto-cancel otherwise.&lt;/LI&gt;&lt;LI&gt;A data pipeline must process 10,000 records, resume from record 6,832 after a crash, and never double-process a record.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;The usual answer is a pile of cron jobs, message queues, a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;status&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;column in a database, and retry loops. It works until a process gets OOM-killed halfway through step 3, and you spend the next two hours reconstructing what happened from logs.&lt;/P&gt;&lt;P&gt;Temporal exists to remove that entire category of work.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;1. What Temporal actually is&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Temporal is a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;durable execution&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;platform. You write your business logic as ordinary code — loops, conditionals, function calls,&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;sleep&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;— and Temporal guarantees that this code runs to completion&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;exactly once&lt;/STRONG&gt;, even if the process running it crashes, the machine dies, or the datacenter goes away for an hour.&lt;/P&gt;&lt;P&gt;The core trick: Temporal does not persist your program's&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;EM&gt;memory&lt;/EM&gt;. It persists the&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;history of everything that happened&lt;/STRONG&gt;. When your process dies and a new one picks up the work, Temporal replays that history through your function to rebuild its exact state, then continues from where it left off.&lt;/P&gt;&lt;P&gt;Two things follow from this, and almost everything else in Temporal is a consequence of them:&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;Your workflow code must be&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;deterministic&lt;/STRONG&gt;, because it gets replayed.&lt;/LI&gt;&lt;LI&gt;Anything non-deterministic (network calls, DB reads, random numbers, clock reads) must be moved out into&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;activities&lt;/STRONG&gt;, whose results are recorded in history and returned verbatim on replay.&lt;/LI&gt;&lt;/OL&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;2. The building blocks Workflow&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;The orchestrator. It decides&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;EM&gt;what&lt;/EM&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;happens and&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;EM&gt;in what order&lt;/EM&gt;. It is deterministic and must not perform I/O.&lt;/P&gt;&lt;PRE&gt;package app

import (
	"time"

	"go.temporal.io/sdk/temporal"
	"go.temporal.io/sdk/workflow"
)

type OrderInput struct {
	OrderID string
	UserID  string
	Amount  int64
}

func OrderWorkflow(ctx workflow.Context, in OrderInput) (string, error) {
	ao := workflow.ActivityOptions{
		StartToCloseTimeout: 30 * time.Second,
		RetryPolicy: &amp;amp;temporal.RetryPolicy{
			InitialInterval:    time.Second,
			BackoffCoefficient: 2.0,
			MaximumInterval:    time.Minute,
			MaximumAttempts:    5,
		},
	}
	ctx = workflow.WithActivityOptions(ctx, ao)

	var chargeID string
	if err := workflow.ExecuteActivity(ctx, ChargeCard, in.UserID, in.Amount).Get(ctx, &amp;amp;chargeID); err != nil {
		return "", err
	}

	if err := workflow.ExecuteActivity(ctx, ReserveInventory, in.OrderID).Get(ctx, nil); err != nil {
		// compensate: the money is already taken
		_ = workflow.ExecuteActivity(ctx, RefundCharge, chargeID).Get(ctx, nil)
		return "", err
	}

	if err := workflow.ExecuteActivity(ctx, SendReceipt, in.UserID, chargeID).Get(ctx, nil); err != nil {
		return "", err
	}

	return chargeID, nil
}&lt;/PRE&gt;&lt;P&gt;Read that again: it looks like a plain Go function. There is no state machine, no&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;status&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;column, no retry loop, no resume logic. If the worker running this dies between&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;ChargeCard&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;and&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;ReserveInventory, another worker picks it up, replays history, sees&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;ChargeCard&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;already returned&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;chargeID, and proceeds straight to&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;ReserveInventory. The card is not charged twice.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Activity&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The part that touches the outside world. Activities may be non-deterministic, may fail, and are automatically retried according to their retry policy.&lt;/P&gt;&lt;PRE&gt;package app

import (
	"context"
	"fmt"
)

func ChargeCard(ctx context.Context, userID string, amount int64) (string, error) {
	resp, err := paymentGateway.Charge(ctx, userID, amount)
	if err != nil {
		return "", fmt.Errorf("charge failed: %w", err)
	}
	return resp.ChargeID, nil
}

func ReserveInventory(ctx context.Context, orderID string) error {
	return inventory.Reserve(ctx, orderID)
}&lt;/PRE&gt;&lt;P&gt;Activities receive a normal&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;context.Context. It is cancelled when the activity is cancelled or its timeout fires, so pass it down to your HTTP and DB calls.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Worker&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The process that hosts your code and polls Temporal for work. It is just a long-running Go binary.&lt;/P&gt;&lt;PRE&gt;package main

import (
	"log"

	"go.temporal.io/sdk/client"
	"go.temporal.io/sdk/worker"

	"example.com/app"
)

func main() {
	c, err := client.Dial(client.Options{
		HostPort:  client.DefaultHostPort, // localhost:7233
		Namespace: "default",
	})
	if err != nil {
		log.Fatalln("dial:", err)
	}
	defer c.Close()

	w := worker.New(c, "orders", worker.Options{})

	w.RegisterWorkflow(app.OrderWorkflow)
	w.RegisterActivity(app.ChargeCard)
	w.RegisterActivity(app.ReserveInventory)
	w.RegisterActivity(app.RefundCharge)
	w.RegisterActivity(app.SendReceipt)

	if err := w.Run(worker.InterruptCh()); err != nil {
		log.Fatalln("worker:", err)
	}
}&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;Starting a workflow&lt;/STRONG&gt;&lt;/P&gt;&lt;PRE&gt;opts := client.StartWorkflowOptions{
	ID:        "order-" + in.OrderID, // your idempotency key
	TaskQueue: "orders",
}

run, err := c.ExecuteWorkflow(ctx, opts, app.OrderWorkflow, in)
if err != nil {
	return err
}

var chargeID string
err = run.Get(ctx, &amp;amp;chargeID) // blocks until the workflow completes&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;The workflow ID is your deduplication mechanism.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Calling&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;ExecuteWorkflow&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;twice with the same ID while the first is still running returns a handle to the existing run rather than starting a second one. Derive it from a business key — order ID, user ID, invoice number — never from a UUID generated at call time.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;3. Task queues and namespaces&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Task queue&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;is the routing mechanism. Workers poll a named queue; workflows and activities are dispatched to a queue. That's it — there is no broker configuration, no exchange, no topic.&lt;/P&gt;&lt;P&gt;This gives you useful control for free:&lt;/P&gt;&lt;PRE&gt;// Route heavy GPU/CPU work to a dedicated fleet
ctx = workflow.WithActivityOptions(ctx, workflow.ActivityOptions{
	TaskQueue:           "video-encoding",
	StartToCloseTimeout: 2 * time.Hour,
})&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;Namespace&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;is the isolation boundary — think of it as a virtual cluster. Use separate namespaces for&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;prod,&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;staging, and per-team workloads. Retention, archival, and access controls are configured per namespace.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;4. Timeouts: the four that matter&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;This is where most people get burned, so be explicit about all of them.&lt;/P&gt;&lt;P&gt;Timeout Meaning Guidance ScheduleToStartTimeout How long a task may sit in the queue before a worker picks it up Usually leave unset; a non-zero value here is a capacity alarm, not a correctness control StartToCloseTimeout Max duration of a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;single attempt&lt;/STRONG&gt; &lt;STRONG&gt;Always set this.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Make it slightly larger than your realistic p99 ScheduleToCloseTimeout Total wall-clock budget across all retries Set when there is a real business deadline HeartbeatTimeout Max gap between heartbeats for long activities Required for anything that runs more than ~a minute&lt;/P&gt;&lt;P&gt;And on the workflow itself:&lt;/P&gt;&lt;PRE&gt;opts := client.StartWorkflowOptions{
	ID:                       "order-123",
	TaskQueue:                "orders",
	WorkflowExecutionTimeout: 7 * 24 * time.Hour, // total lifetime, retries included
	WorkflowRunTimeout:       24 * time.Hour,     // single run
	WorkflowTaskTimeout:      10 * time.Second,   // one decision step
}&lt;/PRE&gt;&lt;HR /&gt;&lt;P&gt;&lt;STRONG&gt;&lt;FONT size="4"&gt;5. Retries and error semantics&lt;/FONT&gt;&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Activities retry by default (infinitely, with exponential backoff, unless you say otherwise). Workflows do&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;not&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;retry by default — they are already durable, so a workflow "failing" means your business logic decided to fail.&lt;/P&gt;&lt;P&gt;Some errors should never be retried. A&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;400 Bad Request&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;will still be a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;400&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;on the fifth attempt.&lt;/P&gt;&lt;PRE&gt;// Mark an error as non-retryable from inside an activity
return temporal.NewNonRetryableApplicationError(
	"card declined", "CardDeclined", nil,
)&lt;/PRE&gt;&lt;PRE&gt;// Or exclude by type in the retry policy
RetryPolicy: &amp;amp;temporal.RetryPolicy{
	MaximumAttempts:        5,
	NonRetryableErrorTypes: []string{"CardDeclined", "InvalidInput"},
}&lt;/PRE&gt;&lt;P&gt;Handling failures in the workflow:&lt;/P&gt;&lt;PRE&gt;var appErr *temporal.ApplicationError
err := workflow.ExecuteActivity(ctx, ChargeCard, in.UserID, in.Amount).Get(ctx, &amp;amp;chargeID)
if errors.As(err, &amp;amp;appErr) &amp;amp;&amp;amp; appErr.Type() == "CardDeclined" {
	return workflow.ExecuteActivity(ctx, NotifyDeclined, in.UserID).Get(ctx, nil)
}&lt;/PRE&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;&lt;STRONG&gt;Retries make at-least-once the default.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Design activities to be idempotent. Pass an idempotency key (ActivityInfo.WorkflowExecution.ID&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;+ activity ID is a good one) to any external API that supports it.&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;6. Heartbeats and long-running activities&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;If an activity takes minutes or hours, heartbeat so Temporal can detect a dead worker quickly instead of waiting out the full&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;StartToCloseTimeout. Heartbeats can also carry progress, which lets a retry resume instead of restart.&lt;/P&gt;&lt;PRE&gt;func ProcessRecords(ctx context.Context, batchID string) error {
	start := 0
	if activity.HasHeartbeatDetails(ctx) {
		_ = activity.GetHeartbeatDetails(ctx, &amp;amp;start) // resume where we died
	}

	records, err := loadRecords(ctx, batchID)
	if err != nil {
		return err
	}

	for i := start; i &amp;lt; len(records); i++ {
		if err := handle(ctx, records[i]); err != nil {
			return err
		}
		activity.RecordHeartbeat(ctx, i+1)

		select {
		case &amp;lt;-ctx.Done():
			return ctx.Err() // cancelled or timed out
		default:
		}
	}
	return nil
}&lt;/PRE&gt;&lt;P&gt;Set&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;HeartbeatTimeout&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;in the activity options or the heartbeats are ignored.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;7. Timers: sleeping for days is normal&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;PRE&gt;// Wait three days. Costs nothing while waiting — no thread, no memory, no cron.
if err := workflow.Sleep(ctx, 72*time.Hour); err != nil {
	return err
}&lt;/PRE&gt;&lt;P&gt;A sleeping workflow is not resident in any process. Temporal stores a timer and re-dispatches the workflow when it fires. Sleeping for a year is as cheap as sleeping for a second.&lt;/P&gt;&lt;P&gt;Never use&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;time.Sleep,&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;time.Now(),&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;rand,&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;os.Getenv, or map iteration order in workflow code. Use the deterministic equivalents:&lt;/P&gt;&lt;PRE&gt;now := workflow.Now(ctx)
id := workflow.GetInfo(ctx).WorkflowExecution.ID

var uuid string
_ = workflow.SideEffect(ctx, func(ctx workflow.Context) interface{} {
	return newUUID()
}).Get(&amp;amp;uuid)&lt;/PRE&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;8. Signals: sending data into a running workflow&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Signals are how the outside world talks to a workflow in flight.&lt;/P&gt;&lt;PRE&gt;func SubscriptionWorkflow(ctx workflow.Context, userID string) error {
	cancelCh := workflow.GetSignalChannel(ctx, "cancel")
	cancelled := false

	for month := 0; month &amp;lt; 12; month++ {
		timerCtx, cancelTimer := workflow.WithCancel(ctx)
		timer := workflow.NewTimer(timerCtx, 30*24*time.Hour)

		sel := workflow.NewSelector(ctx)
		sel.AddFuture(timer, func(workflow.Future) {})
		sel.AddReceive(cancelCh, func(c workflow.ReceiveChannel, _ bool) {
			c.Receive(ctx, nil)
			cancelled = true
			cancelTimer()
		})
		sel.Select(ctx)

		if cancelled {
			return workflow.ExecuteActivity(ctx, SendCancellationEmail, userID).Get(ctx, nil)
		}
		if err := workflow.ExecuteActivity(ctx, ChargeMonthly, userID).Get(ctx, nil); err != nil {
			return err
		}
	}
	return nil
}&lt;/PRE&gt;&lt;P&gt;From a client:&lt;/P&gt;&lt;PRE&gt;err := c.SignalWorkflow(ctx, "subscription-user-42", "", "cancel", nil)&lt;/PRE&gt;&lt;P&gt;SignalWithStartWorkflow&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;starts the workflow if it isn't running and signals it if it is — the standard pattern for "append to a session that may or may not exist yet".&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Human-in-the-loop&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;falls out of this naturally:&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;workflow.Await&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;on an approval signal with a timeout, and escalate if nobody responds.&lt;/P&gt;&lt;PRE&gt;approved := false
approvalCh := workflow.GetSignalChannel(ctx, "approval")
workflow.Go(ctx, func(ctx workflow.Context) {
	var decision bool
	approvalCh.Receive(ctx, &amp;amp;decision)
	approved = decision
})

ok, _ := workflow.AwaitWithTimeout(ctx, 48*time.Hour, func() bool { return approved })
if !ok {
	return workflow.ExecuteActivity(ctx, EscalateToManager, in).Get(ctx, nil)
}&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;Queries: reading state out&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Queries are read-only, synchronous, and must not mutate state or schedule work.&lt;/P&gt;&lt;PRE&gt;func OrderWorkflow(ctx workflow.Context, in OrderInput) error {
	status := "pending"
	if err := workflow.SetQueryHandler(ctx, "status", func() (string, error) {
		return status, nil
	}); err != nil {
		return err
	}
	// ... status = "charged" ... status = "shipped" ...
	return nil
}&lt;/PRE&gt;&lt;PRE&gt;resp, _ := c.QueryWorkflow(ctx, "order-123", "", "status")
var status string
_ = resp.Get(&amp;amp;status)&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;Updates: signal + query in one round trip&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;An update sends data in&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;and&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;returns a result, with optional validation that rejects bad input without writing to history.&lt;/P&gt;&lt;PRE&gt;_ = workflow.SetUpdateHandlerWithOptions(ctx, "addItem",
	func(ctx workflow.Context, item Item) (int, error) {
		cart = append(cart, item)
		return len(cart), nil
	},
	workflow.UpdateHandlerOptions{
		Validator: func(ctx workflow.Context, item Item) error {
			if item.Qty &amp;lt;= 0 {
				return errors.New("quantity must be positive")
			}
			return nil
		},
	},
)&lt;/PRE&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;9. Child workflows and fan-out&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Use a child workflow when a sub-process deserves its own ID, history, retry policy, and lifecycle. Use an activity when it's just a unit of work.&lt;/P&gt;&lt;PRE&gt;// Sequential child
var res ShipmentResult
err := workflow.ExecuteChildWorkflow(ctx, ShipmentWorkflow, orderID).Get(ctx, &amp;amp;res)&lt;/PRE&gt;&lt;PRE&gt;// Parallel fan-out over activities
futures := make([]workflow.Future, 0, len(itemIDs))
for _, id := range itemIDs {
	futures = append(futures, workflow.ExecuteActivity(ctx, ReserveItem, id))
}
for _, f := range futures {
	if err := f.Get(ctx, nil); err != nil {
		return err
	}
}&lt;/PRE&gt;&lt;P&gt;ParentClosePolicy&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;controls what happens to children when the parent finishes — terminate them (default), abandon them, or request cancellation.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;10. Continue-As-New: keeping history bounded&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Event history is not infinite. Practically, keep it under a few thousand events. For workflows that loop forever (subscriptions, monitors, per-entity actors), atomically restart with fresh history:&lt;/P&gt;&lt;PRE&gt;func MonitorWorkflow(ctx workflow.Context, state State) error {
	for i := 0; i &amp;lt; 500; i++ {
		if err := workflow.Sleep(ctx, time.Hour); err != nil {
			return err
		}
		if err := workflow.ExecuteActivity(ctx, Poll, state.Target).Get(ctx, &amp;amp;state); err != nil {
			return err
		}
	}
	return workflow.NewContinueAsNewError(ctx, MonitorWorkflow, state)
}&lt;/PRE&gt;&lt;P&gt;The workflow ID stays the same; only the run ID changes. Signals received while continuing-as-new are carried over.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;11. Determinism and versioning&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Because history is replayed,&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;you cannot freely change workflow code that has in-flight executions&lt;/STRONG&gt;. Reordering activities, adding a step in the middle, or removing a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Sleep&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;will make replay diverge from history and produce a non-determinism error.&lt;/P&gt;&lt;P&gt;Two safe options:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Patching&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;— for incremental changes:&lt;/P&gt;&lt;PRE&gt;if workflow.GetVersion(ctx, "add-fraud-check", workflow.DefaultVersion, 1) == workflow.DefaultVersion {
	// old path: existing runs continue as before
} else {
	// new path: only new runs take this
	if err := workflow.ExecuteActivity(ctx, FraudCheck, in).Get(ctx, nil); err != nil {
		return err
	}
}&lt;/PRE&gt;&lt;P&gt;&lt;STRONG&gt;Worker Versioning / new workflow type&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;— for large rewrites: pin old runs to old workers and route new runs to the new build.&lt;/P&gt;&lt;P&gt;Safe changes that never need a patch: activity implementation bodies, timeout and retry policy values, and log statements.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;12. Testing&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Temporal's Go SDK ships a test framework that runs workflows in-process with a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;skipping clock&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;— a workflow that sleeps 30 days finishes in milliseconds.&lt;/P&gt;&lt;PRE&gt;func TestOrderWorkflow(t *testing.T) {
	var s testsuite.WorkflowTestSuite
	env := s.NewTestWorkflowEnvironment()

	env.OnActivity(ChargeCard, mock.Anything, "u1", int64(500)).Return("ch_1", nil)
	env.OnActivity(ReserveInventory, mock.Anything, "o1").Return(nil)
	env.OnActivity(SendReceipt, mock.Anything, "u1", "ch_1").Return(nil)

	env.ExecuteWorkflow(OrderWorkflow, OrderInput{OrderID: "o1", UserID: "u1", Amount: 500})

	require.True(t, env.IsWorkflowCompleted())
	require.NoError(t, env.GetWorkflowError())

	var chargeID string
	require.NoError(t, env.GetWorkflowResult(&amp;amp;chargeID))
	require.Equal(t, "ch_1", chargeID)
	env.AssertExpectations(t)
}&lt;/PRE&gt;&lt;P&gt;Test compensation paths by returning errors from mocks, and test signals with&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;env.RegisterDelayedCallback.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Replay tests&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;are your safety net against accidental non-determinism: download the history of a production workflow and replay it against your new code in CI.&lt;/P&gt;&lt;PRE&gt;replayer := worker.NewWorkflowReplayer()
replayer.RegisterWorkflow(OrderWorkflow)
err := replayer.ReplayWorkflowHistoryFromJSONFile(nil, "testdata/order_history.json")
require.NoError(t, err)&lt;/PRE&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;13. Schedules and cron&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Recurring work is first-class — no external cron, no missed-fire mystery.&lt;/P&gt;&lt;PRE&gt;_, err := c.ScheduleClient().Create(ctx, client.ScheduleOptions{
	ID: "nightly-reconciliation",
	Spec: client.ScheduleSpec{
		CronExpressions: []string{"0 2 * * *"},
		TimeZoneName:    "Asia/Kolkata",
	},
	Action: &amp;amp;client.ScheduleWorkflowAction{
		ID:        "reconcile",
		Workflow:  ReconcileWorkflow,
		TaskQueue: "batch",
	},
	Overlap: enumspb.SCHEDULE_OVERLAP_POLICY_SKIP,
})&lt;/PRE&gt;&lt;P&gt;Schedules can be paused, backfilled, and triggered manually. The overlap policy answers "what if the previous run is still going?" — a question plain cron never asks.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;14. Search attributes and visibility&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Beyond fetching a workflow by ID, you can index and query workflows by business fields.&lt;/P&gt;&lt;PRE&gt;_ = workflow.UpsertSearchAttributes(ctx, map[string]interface{}{
	"CustomerId":  in.UserID,
	"OrderStatus": "charged",
})&lt;/PRE&gt;&lt;PRE&gt;resp, _ := c.ListWorkflow(ctx, &amp;amp;workflowservice.ListWorkflowExecutionsRequest{
	Query: `WorkflowType = 'OrderWorkflow' AND OrderStatus = 'charged' AND ExecutionStatus = 'Running'`,
})&lt;/PRE&gt;&lt;P&gt;This is how you answer "which orders are stuck in payment right now?" without adding a reporting table.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;15. Cancellation and cleanup&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Cancellation is cooperative and propagates to running activities and child workflows. Cleanup must run in a disconnected context, since the normal one is already cancelled.&lt;/P&gt;&lt;PRE&gt;defer func() {
	if errors.Is(ctx.Err(), workflow.ErrCanceled) {
		newCtx, _ := workflow.NewDisconnectedContext(ctx)
		_ = workflow.ExecuteActivity(newCtx, ReleaseInventory, in.OrderID).Get(newCtx, nil)
	}
}()&lt;/PRE&gt;&lt;P&gt;Note the difference:&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;cancel&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;is graceful and lets the workflow run its cleanup;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;STRONG&gt;terminate&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;is immediate and runs nothing. Prefer cancel.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;16. Running it locally&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;PRE&gt;brew install temporal          # or: https://temporal.io/setup/install-temporal-cli
temporal server start-dev      # server on :7233, Web UI on http://localhost:8233&lt;/PRE&gt;&lt;PRE&gt;temporal workflow start \
  --task-queue orders \
  --type OrderWorkflow \
  --workflow-id order-123 \
  --input '{"OrderID":"o1","UserID":"u1","Amount":500}'

temporal workflow show   --workflow-id order-123
temporal workflow signal --workflow-id order-123 --name cancel
temporal workflow query  --workflow-id order-123 --type status&lt;/PRE&gt;&lt;P&gt;The Web UI shows the full event history of every execution — every activity attempt, every failure, every retry, with inputs and outputs. This alone replaces a large amount of custom logging and admin tooling.&lt;/P&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;17. Production notes&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Payload size.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Inputs, outputs, and signals go into history. Keep them small (rule of thumb: under ~100 KB, hard limit 2 MB). Pass S3 keys or row IDs, not blobs. Use a Data Converter for encryption or compression.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Worker tuning.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;MaxConcurrentActivityExecutionSize&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;and&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;MaxConcurrentWorkflowTaskExecutionSize&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;are your throughput knobs. Run workflow and activity workers separately when activity work is heavy, so a saturated activity pool never starves workflow progress.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Retention.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Closed workflow histories are deleted after the namespace retention period (default 3 days on the dev server, commonly 30 days in production). Ship anything you need long-term to your own store.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Metrics.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;The SDK exports Prometheus metrics. Watch&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;temporal_workflow_task_schedule_to_start_latency&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;(workers under-provisioned) and&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;temporal_activity_execution_failed.&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Sticky execution.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;Workers cache workflow state so they don't replay full history on every task. A cache miss just means a replay — correct, but slower. Size&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;WorkflowCacheSize&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;accordingly.&lt;/LI&gt;&lt;/UL&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;18. When not to use Temporal&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;P&gt;Temporal is not free. It adds a server (plus Cassandra/PostgreSQL/MySQL and Elasticsearch), a programming model with real constraints, and per-step latency measured in milliseconds rather than microseconds.&lt;/P&gt;&lt;P&gt;Skip it when:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;The work is a single, fast, idempotent operation. Just do it inline.&lt;/LI&gt;&lt;LI&gt;You need sub-millisecond latency per step.&lt;/LI&gt;&lt;LI&gt;You need high-throughput stream processing — that's Kafka/Flink territory, not Temporal's.&lt;/LI&gt;&lt;LI&gt;The whole system is one service with one database and no multi-step failure modes worth modelling.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;Reach for it when:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;A process spans multiple services or multiple external APIs.&lt;/LI&gt;&lt;LI&gt;A process spans minutes, days, or months.&lt;/LI&gt;&lt;LI&gt;Partial failure leaves the system in a state someone has to fix by hand.&lt;/LI&gt;&lt;LI&gt;You need auditability of what happened, step by step.&lt;/LI&gt;&lt;LI&gt;You've already written a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;status&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;column, a retry table, and a reconciliation cron for the same flow.&lt;/LI&gt;&lt;/UL&gt;&lt;HR /&gt;&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;19. The mental model to keep&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;&lt;STRONG&gt;Temporal turns a distributed, failure-prone, multi-step process into a single function you can read top to bottom.&lt;/STRONG&gt;&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;Everything else — the event history, replay, task queues, timers, heartbeats — is machinery in service of that one property. Once you internalise "workflows orchestrate and must be deterministic; activities do the work and may fail", the rest of the API stops feeling like magic and starts feeling obvious.&lt;/P&gt;&lt;P&gt;Start with one flow. Pick the one that currently has a&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;status&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;column, a retry cron, and a runbook. Port it. You'll delete more code than you write.&lt;/P&gt;</description>
      <pubDate>Tue, 01 Sep 2026 11:08:22 GMT</pubDate>
      <guid>https://community.hpe.com/t5/software-general/temporal-building-long-running-systems/m-p/7272273#M1586</guid>
      <dc:creator>Prithviraj-</dc:creator>
      <dc:date>2026-09-01T11:08:22Z</dc:date>
    </item>
    <item>
      <title>Re: Temporal: Building Long-Running Systems</title>
      <link>https://community.hpe.com/t5/software-general/temporal-building-long-running-systems/m-p/7272439#M1590</link>
      <description>&lt;P&gt;Thanks for sharing this informative topic. It is very helpful.&lt;/P&gt;</description>
      <pubDate>Mon, 07 Sep 2026 11:26:47 GMT</pubDate>
      <guid>https://community.hpe.com/t5/software-general/temporal-building-long-running-systems/m-p/7272439#M1590</guid>
      <dc:creator>Thaufique_Mod</dc:creator>
      <dc:date>2026-09-07T11:26:47Z</dc:date>
    </item>
  </channel>
</rss>

