Middleware
Genkit allows you to use middleware to modify the behavior of generate() calls. Middleware can be used for various purposes, such as retrying failed requests, falling back to different models, or injecting tools and context.
You can use pre-packaged middleware or build your own custom middleware.
Installation
Section titled “Installation”The middleware framework is part of the core ai package, and the pre-packaged middleware ships in plugins/middleware. Both come with the core Genkit module:
go get github.com/firebase/genkit/go@latestRegister the Middleware plugin during genkit.Init to expose the built-ins to the Dev UI and to other-runtime callers:
import ( "context"
"github.com/firebase/genkit/go/genkit" "github.com/firebase/genkit/go/plugins/googlegenai" "github.com/firebase/genkit/go/plugins/middleware")
ctx := context.Background()g := genkit.Init(ctx, genkit.WithPlugins( &googlegenai.GoogleAI{}, &middleware.Middleware{},))The snippets that follow also draw on log, net/http, os, time, github.com/firebase/genkit/go/ai, and github.com/firebase/genkit/go/core/status.
For pure Go programs that just attach middleware to a genkit.Generate() call, plugin registration is optional. Passing a middleware value directly to ai.WithUse invokes its New method on the local fast path without consulting the registry.
Attaching middleware
Section titled “Attaching middleware”ai.WithUse is the option that attaches middleware. It returns an ai.CommonGenOption, so the same call is accepted by genkit.Generate (and GenerateText / GenerateData), by genkit.DefinePrompt, and by Prompt.Execute / ExecuteStream. There are four ways to reach it.
Per call. Pass the middleware’s config struct to the generation you want it on. This is the default and it needs no registration: the value carries its own config, so Genkit calls its New method directly.
resp, err := genkit.Generate(ctx, g, ai.WithPrompt("Summarize the release notes."), ai.WithUse(&middleware.Retry{MaxRetries: 2}),)The built-in middleware declare Name and New on value receivers, so a bare middleware.Retry{...} and a pointer &middleware.Retry{...} both satisfy ai.Middleware. These pages pass pointers.
Inline. For a one-off that needs no named type or Dev UI entry, adapt a closure with ai.MiddlewareFunc. See Inline middleware below.
On a prompt definition. Middleware passed to genkit.DefinePrompt applies to every execution of that prompt.
p := genkit.DefinePrompt(g, "assistant", ai.WithPrompt("{{query}}"), ai.WithUse(&middleware.Retry{MaxRetries: 2}),)
// Inherits the prompt's middleware.resp, err := p.Execute(ctx, ai.WithInput(map[string]any{"query": "hello"}))
// Replaces it: only Fallback runs on this execution, Retry does not.resp, err = p.Execute(ctx, ai.WithInput(map[string]any{"query": "hello"}), ai.WithUse(&middleware.Fallback{}),)Two things about prompts catch people out:
- Prompt-level and execute-level middleware do not merge.
ai.WithUseat execute time replaces the whole chain the prompt was defined with. ai.MiddlewareFunccannot be used ingenkit.DefinePrompt. A prompt action serializes its options, and a function value has no JSON form, so execution fails withjson: unsupported type: ai.MiddlewareFunc. Use a named config type at definition time.
By name. Middleware that is registered, either by a plugin or by genkit.DefineMiddleware, can be selected by name from a .prompt file’s use: frontmatter. Each entry is a bare name or a name with a config map:
---model: googleai/gemini-flash-latestuse: - name: genkit-middleware/retry config: maxRetries: 2 - genkit-middleware/skills---These entries resolve through the registry at execute time, so the plugin providing them has to be registered or the execution fails with NOT_FOUND.
ai.WithMiddleware still exists but is deprecated. It takes an ai.ModelMiddleware, which wraps only the model call, has no generate or tool hook, and never appears in the Dev UI. Use ai.WithUse instead.
Available middleware
Section titled “Available middleware”The plugins/middleware package provides several useful middleware options out of the box. This list represents the middleware built and maintained by the Genkit team, but there may also be community-built middleware available.
1. Retry middleware (Retry)
Section titled “1. Retry middleware (Retry)”Automatically retries failed model API calls on transient error codes (such as RESOURCE_EXHAUSTED and UNAVAILABLE) using exponential backoff with jitter. Only the model API call is retried; the surrounding tool loop is not replayed.
Genkit does not retry on its own. A bare genkit.Generate makes exactly one attempt per model call, and Retry is entirely opt-in.
resp, err := genkit.Generate(ctx, g, ai.WithModelName("googleai/gemini-flash-latest"), ai.WithPrompt("Heavy reasoning task..."), ai.WithUse(&middleware.Retry{ MaxRetries: 3, InitialDelayMs: 1000, BackoffFactor: 2, }),)Configuration options:
MaxRetries(optional): The number of retries after the first attempt (default: 3).MaxRetries: 3allows up to four model calls in total.Statuses(optional): A list ofstatus.Namevalues, fromgithub.com/firebase/genkit/go/core/status, that should trigger a retry (default:status.Unavailable,status.DeadlineExceeded,status.ResourceExhausted,status.Aborted,status.Internal). The Developer UI offers the canonical set as a multi-select, andstatus.Names()returns the same list in code.InitialDelayMs(optional): The initial delay between retries in milliseconds (default: 1000).MaxDelayMs(optional): The upper bound on retry delay in milliseconds (default: 60000).BackoffFactor(optional): The factor by which the delay increases after each retry (default: 2).NoJitter(optional): If true, disables random jitter on the delay (default: false).
What counts as retryable. The decision turns on whether the error is classified, not on its Go type:
func isRetryable(err error, statuses []status.Name) bool { if s, ok := status.Classified(err); ok { return slices.Contains(statuses, s) } return true}A classified error is retried only when its status is on the list. An unclassified error is always retried, whatever the list says, because a provider SDK error that carries no Genkit status is usually a transport failure worth another attempt. Two consequences are worth knowing:
- A cancelled context classifies as
CANCELLED, which is not on the default list, so cancelling stops the retries. Acontext.DeadlineExceededclassifies asDEADLINE_EXCEEDED, which is on the list, so it is retried. - Wrapping an error with
%vinstead of%wdestroys its classification. That silently converts a non-retryableINVALID_ARGUMENTinto an always-retried unclassified error.
Deadlines and cancellation. Genkit sets no default request timeout. The deadline on the context you hand genkit.Generate is the only bound, so give it one:
func handler(w http.ResponseWriter, r *http.Request) { ctx, cancel := context.WithTimeout(r.Context(), 60*time.Second) defer cancel()
resp, err := genkit.Generate(ctx, g, ai.WithPrompt("Summarize the release notes."), ai.WithUse(&middleware.Retry{MaxRetries: 3}), ) if err != nil { http.Error(w, err.Error(), http.StatusInternalServerError) return } w.Write([]byte(resp.Text()))}That one deadline covers every retry attempt; it does not restart per attempt, and the backoff waits come out of the same budget. &middleware.Retry{MaxRetries: 3, InitialDelayMs: 1000, BackoffFactor: 2} spends 1s, then 2s, then 4s sleeping alone, before counting any time in the model. Size the timeout for the whole cascade, not for one call.
Cancelling the context stops the retries: a cancelled context classifies as CANCELLED, which is not on the retry list. Because r.Context() is cancelled when the HTTP client disconnects, an abandoned request stops burning attempts on its own.
2. Fallback middleware (Fallback)
Section titled “2. Fallback middleware (Fallback)”Automatically switches to a different model if the primary model fails on a fallback-eligible status. Useful for falling back to a smaller or faster model when a large model exceeds quota limits.
resp, err := genkit.Generate(ctx, g, ai.WithModelName("googleai/gemini-pro-latest"), ai.WithPrompt("Try the pro model first..."), ai.WithUse(&middleware.Fallback{ Models: []ai.ModelRef{ googlegenai.ModelRef("googleai/gemini-flash-latest", nil), }, Statuses: []status.Name{status.ResourceExhausted}, }),)Configuration options:
Models(required): An ordered list ofai.ModelRefvalues to try after the primary fails. Each ref’sConfigis used verbatim for that model; the original request’s config is not inherited. Usegooglegenai.ModelRef(or the equivalent helper for your provider) to attach configuration. A named model that is not registered fails withai.ErrModelNotFound.Statuses(optional): A list ofstatus.Namevalues that should trigger a fallback (default:status.Unavailable,status.DeadlineExceeded,status.ResourceExhausted,status.Aborted,status.Internal,status.NotFound,status.Unimplemented). The two entriesRetrydoes not share areNotFoundandUnimplemented, so a model the provider does not serve fails over instead of failing.
What counts as fallback-eligible. The predicate is Retry’s with one line changed:
func isFallbackRetryable(err error, statuses []status.Name) bool { if s, ok := status.Classified(err); ok { return slices.Contains(statuses, s) } return false}An unclassified error never triggers a fallback; it propagates immediately. This is the deliberate opposite of Retry, and the reason is cost: moving a request to a different billed model is a bigger action than reissuing the same one, so it demands an explicit classification. A plugin that classifies its provider’s errors, for example with status.Base(status.FromHTTPCode(code)), is what makes fallback work at all.
The retry-fallback sample composes the two, pointing the primary at a model id that does not exist so the cascade fires on every run.
3. Tool approval middleware (ToolApproval)
Section titled “3. Tool approval middleware (ToolApproval)”Restricts tool execution to an allow list. Tools not in the list trigger a tool interrupt that you can resolve by prompting the user and then resuming with an explicit approval flag.
type WriteFileInput struct { Path string `json:"path"` Content string `json:"content"`}
writeFileTool := genkit.DefineTool(g, "write_file", "Write a text file.", func(ctx *ai.ToolContext, in WriteFileInput) (string, error) { if err := os.WriteFile(in.Path, []byte(in.Content), 0o600); err != nil { return "", err } return "wrote " + in.Path, nil })
// 1. Initial attempt: any tool not in AllowedTools interrupts the call.resp, err := genkit.Generate(ctx, g, ai.WithPrompt("write a file"), ai.WithTools(writeFileTool), ai.WithUse(&middleware.ToolApproval{ AllowedTools: []string{}, // Empty list interrupts every tool call. }),)if err != nil { log.Fatal(err)}
if resp.FinishReason == ai.FinishReasonInterrupted { // 2. One turn can hold several interrupts. Approve each one you accept. var restarts []*ai.Part for _, interrupt := range resp.Interrupts() { // Show interrupt.ToolRequest to the user before approving. approved, err := writeFileTool.RestartWith(interrupt, ai.WithResumedMetadata[WriteFileInput](map[string]any{"toolApproved": true}), ) if err != nil { log.Fatal(err) } restarts = append(restarts, approved) }
// 3. Resume, with the same middleware config as the first call. resumed, err := genkit.Generate(ctx, g, ai.WithMessages(resp.History()...), ai.WithTools(writeFileTool), ai.WithToolRestarts(restarts...), ai.WithUse(&middleware.ToolApproval{AllowedTools: []string{}}), ) _ = resumed}RestartWith is a method on the typed *ai.ToolAction[In, Out] that genkit.DefineTool returns, so it needs the tool’s input type at compile time. When the interrupted tool is not known statically, resolve it by name and use the type-erased Restart:
tool := genkit.LookupTool(g, interrupt.ToolRequest.Name)approved := tool.Restart(interrupt, &ai.RestartOptions{ ResumedMetadata: map[string]any{"toolApproved": true},})A bare resume without the toolApproved flag is not treated as approval, so unrelated resume flows can’t bypass approval gating. Approval travels entirely through that restart metadata, not through the allow list, which is why the resume call above passes the same config as the first one.
A blocked call is still visible in traces: when a WrapTool hook resolves a call without running the tool, Genkit emits the tool span the tool itself would have produced.
Configuration options:
AllowedTools(defaults to gating every tool): The tool names pre-approved to run without interruption. The gate is a membership test, so a nil slice and an empty slice behave identically: every tool interrupts. There is no wildcard. To let a tool run unconditionally, list its name.
4. Skills middleware (Skills)
Section titled “4. Skills middleware (Skills)”Scans a directory for SKILL.md files (and their YAML frontmatter) and injects a list of them into the system prompt. It also registers a use_skill tool, taking a single skillName string, that the model calls to load one skill’s full body on demand. Keeping the heavy instructions behind that tool is the point: only the name and description of each skill are on the hot path.
resp, err := genkit.Generate(ctx, g, ai.WithPrompt("How do I run tests in this repo?"), ai.WithUse(&middleware.Skills{SkillPaths: []string{"./skills"}}),)Configuration options:
SkillPaths(optional): A list of directories to scan for skills. Each direct subdirectory containing aSKILL.mdfile is exposed as a skill (default:["skills"], resolved against the process working directory).
Three details of the scan are worth knowing:
- The directory name is the skill name the model uses. The
namefield in the frontmatter is parsed but never used for that, so a directory namedpython-helperispython-helperno matter what the file says. Onlydescriptionreaches the prompt. - Discovery is one level deep. Direct subdirectories are considered, nested ones are not, and names starting with
.are skipped. - A path that cannot be read is skipped rather than fatal, with a warning in the log. A mistyped
SkillPathsentry therefore yields no skills and no error.
Loading a skill costs a turn of the tool loop, so raise ai.WithMaxTurns above the default of 5 to leave room for the answer. The skills sample ships four deliberately loud personas so the effect of a load is visible in one run.
5. Filesystem middleware (Filesystem)
Section titled “5. Filesystem middleware (Filesystem)”Grants the model access to a single root directory by injecting file manipulation tools. Two are always registered, list_files and read_file; write_file and edit_file join them only when AllowWriteAccess is true. Path safety is enforced by os.Root, which rejects any path that resolves outside the root, including via .., absolute paths, or symbolic links.
resp, err := genkit.Generate(ctx, g, ai.WithPrompt("Create a hello world program in the workspace"), ai.WithUse(&middleware.Filesystem{ RootDir: "./workspace", AllowWriteAccess: true, }),)Configuration options:
RootDir(required): The root directory all filesystem operations are confined to. Its JSON key isrootDirectory, which is the name to use in.promptfrontmatter and the Dev UI. An empty value fails withINVALID_ARGUMENT.AllowWriteAccess(optional): If true, additionally registerswrite_fileandedit_file(default: false).ToolNamePrefix(optional): A prefix prepended verbatim to each tool name, so"repo_"yieldsrepo_list_filesand the rest. Use distinct prefixes when attaching multipleFilesystemmiddlewares to one call so their tool names don’t collide.
read_file does not return the file in its tool result. It answers with a short summary and injects the contents as a user message on the next turn, which is why the middleware needs a WrapGenerate hook as well as tools. edit_file refuses to touch a file that was not read earlier in the same call, and write_file refuses to overwrite an existing one that was not read; both refuse again if the file changed on disk since that read. Creating a new file with write_file needs no prior read.
The filesystem sample ships a small mock project and one flow per mode, read-only and write-enabled.
The experimental middleware plugin
Section titled “The experimental middleware plugin”A second, in-preview middleware plugin lives at github.com/firebase/genkit/go/plugins/middleware/exp. It provides Agents for sub-agent delegation, synchronous or in the background, and Artifacts for session artifact access; both are documented in Multi-agent delegation. It registers under the provider name genkit-middleware-exp, so its names do not collide with the stable plugin’s and you can register both:
import ( "github.com/firebase/genkit/go/plugins/middleware" middlewarex "github.com/firebase/genkit/go/plugins/middleware/exp")
g := genkit.Init(ctx, genkit.WithPlugins( &googlegenai.GoogleAI{}, &middleware.Middleware{}, &middlewarex.Middleware{},))A .prompt reference such as genkit-middleware/retry resolves to the stable plugin; the experimental ones are named genkit-middleware-exp/agents and genkit-middleware-exp/artifacts. Being in preview, they may change in any minor release.
A third package, github.com/firebase/genkit/go/plugins/a2ui/exp, ships one middleware of its own: a2uix.Surfaces lets the model stream generative UI to a browser over the A2UI protocol, and a2uix.A2UI is the plugin that registers it by name. It is in preview too. See Generative UI (A2UI).
Building your own custom middleware
Section titled “Building your own custom middleware”A middleware in Go is any value that satisfies the ai.Middleware interface:
type Middleware interface { Name() string // stable, registered identifier New(ctx context.Context) (*ai.Hooks, error) // builds a per-call hook bundle}New is invoked once per genkit.Generate() call. The returned *ai.Hooks bundle is reused across every iteration of the tool loop within that call:
type Hooks struct { // Tools are extra tools to register for this Generate call alongside any user-supplied tools. Tools []ai.Tool
// WrapGenerate wraps each iteration of the tool loop. WrapGenerate func(ctx context.Context, params *ai.GenerateParams, next ai.GenerateNext) (*ai.ModelResponse, error)
// WrapModel wraps each model API call. WrapModel func(ctx context.Context, params *ai.ModelParams, next ai.ModelNext) (*ai.ModelResponse, error)
// WrapTool wraps each tool execution. May run concurrently for parallel tool calls. WrapTool func(ctx context.Context, params *ai.ToolParams, next ai.ToolNext) (*ai.MultipartToolResponse, error)}Implement only the hooks your middleware needs. A nil hook field is treated as a pass-through. An error from New fails the generate call with the status it carries; an unclassified one reports INVALID_ARGUMENT, since a configuration the middleware rejects is the caller’s mistake.
The middleware in this section draw on these imports:
import ( "context" "crypto/sha256" "encoding/hex" "encoding/json" "slices" "strings" "sync" "time"
"github.com/firebase/genkit/go/ai" "github.com/firebase/genkit/go/core/logger" "github.com/firebase/genkit/go/core/status" "github.com/firebase/genkit/go/genkit")When each hook fires
Section titled “When each hook fires”A Generate call runs a tool loop: the model produces output, any tool calls execute, results feed back into a new model call, and so on until the model stops. The hooks attach at three different layers of this loop:
| Hook | Fires | Use for |
|---|---|---|
WrapGenerate | Once per tool-loop iteration. N tool turns means N+1 invocations. | Logic that needs to see the whole conversation: rewrites, system-prompt injection, message accumulation. |
WrapModel | Once per model API call, inside an iteration. | Logic about the model call itself: retry, fallback, caching. |
WrapTool | Once per tool execution. May run concurrently for parallel tool calls in the same iteration. | Logic about a single tool execution: approval, sandboxing, logging. |
WrapGenerate and WrapModel are not called concurrently within a single Generate call. WrapTool may be, since multiple tools can execute in parallel.
What each hook receives
Section titled “What each hook receives”ai.GenerateParams, for WrapGenerate:
| Field | Type | What it is |
|---|---|---|
Request | *ai.ModelRequest | The model request for this turn, with the messages accumulated so far. Replace it or edit it to change what the model receives. |
Options | *ai.GenerateActionOptions | A per-turn copy of the options Generate was called with, carrying what Request does not: Model, MaxTurns, resume directives. |
Iteration | int | The tool-loop iteration, 0-indexed. |
MessageIndex | int | The index of the next message in the streamed response sequence. |
Callback | ai.ModelStreamCallback | The streaming callback, or nil when not streaming. |
ai.ModelParams, for WrapModel:
| Field | Type | What it is |
|---|---|---|
Request | *ai.ModelRequest | The model request about to be sent. |
Callback | ai.ModelStreamCallback | The streaming callback, or nil. |
ai.ToolParams, for WrapTool:
| Field | Type | What it is |
|---|---|---|
Request | *ai.ToolRequest | Name the tool name, Input the decoded arguments as any, Ref the model’s call id, Partial for stream chunks. |
Tool | ai.Tool | The resolved tool: Tool.Name(), Tool.Definition(). |
Both Request fields above are the same *ai.ModelRequest:
| Field | Type |
|---|---|
Messages | []*ai.Message |
Config | any |
Tools | []*ai.ToolDefinition |
ToolChoice | ai.ToolChoice |
Output | *ai.ModelOutputConfig |
Docs | []*ai.Document |
Mutating params. A hook may edit params in place, or build a different value and hand that to next; the engine reads whatever it is given. Three rules follow from how the loop owns these values:
GenerateParams.RequestandModelParams.Requestare fresh per turn, so an in-place edit stays with that turn. The exception is a*ai.Messageshared with the next turn: copy the message withClone()before editing its text.GenerateParams.Optionsis a shallow per-turn copy. Writes to it reach nothing, so treat it as read-only.- For redaction that must not reach the stored history, prefer
WrapModel, and shallow-copy theModelRequestplus only the messages and parts you change.
Middleware that emits its own chunks through Callback must set ModelResponseChunk.Role and Index explicitly, and advance GenerateParams.MessageIndex so downstream middleware and the model see the shifted value.
Short-circuiting. A hook can return without calling next. WrapTool is the useful case: return an *ai.MultipartToolResponse and the tool never runs, but the model still receives a result it can react to on the next turn.
type ToolRefusal struct { Denied []string `json:"denied,omitempty"`}
func (ToolRefusal) Name() string { return "mine/toolRefusal" }
func (tr ToolRefusal) New(context.Context) (*ai.Hooks, error) { return &ai.Hooks{ WrapTool: func(ctx context.Context, p *ai.ToolParams, next ai.ToolNext) (*ai.MultipartToolResponse, error) { if slices.Contains(tr.Denied, p.Request.Name) { return &ai.MultipartToolResponse{ Output: map[string]any{"error": "denied by policy: " + p.Request.Name}, }, nil } return next(ctx, p) }, }, nil}This is the shape ToolApproval uses. A hook that skips next still produces the tool span, so the refusal is visible in traces.
Logging from a hook
Section titled “Logging from a hook”Log through the core/logger helpers in github.com/firebase/genkit/go/core/logger rather than through log or fmt. They take the context as their first argument, and that context is what carries the active span, so each record lands on the trace it belongs to and shows up in the Dev UI beside the span that produced it:
func Debug(ctx context.Context, msg string, args ...any)func Info(ctx context.Context, msg string, args ...any)func Warn(ctx context.Context, msg string, args ...any)func Error(ctx context.Context, msg string, args ...any)Before writing a middleware whose only job is to log, check whether you need it: Genkit already brackets every generate, model, and tool hook with debug records carrying the middleware name, the hook, its duration, whether it short-circuited, and any error. Middleware is the one layer with no span of its own, so those records exist to fill that gap. Hand-rolled timing middleware usually duplicates them.
A simple example
Section titled “A simple example”Here is a custom middleware that reports how long each model call takes, along with a label from its config:
type Timing struct { Label string `json:"label,omitempty" jsonschema_description:"Label attached to each timing record."`}
func (Timing) Name() string { return "mine/timing" }
func (t Timing) New(ctx context.Context) (*ai.Hooks, error) { return &ai.Hooks{ WrapModel: func(ctx context.Context, p *ai.ModelParams, next ai.ModelNext) (*ai.ModelResponse, error) { start := time.Now() resp, err := next(ctx, p) logger.Info(ctx, "model call finished", "label", t.Label, "duration", time.Since(start).Round(time.Millisecond), "error", err) return resp, err }, }, nil}To use it:
resp, err := genkit.Generate(ctx, g, ai.WithPrompt("Hello"), ai.WithUse(Timing{Label: "demo"}),)Note the receivers. Timing declares Name and New on value receivers, so a bare Timing{...} satisfies ai.Middleware, and so does &Timing{...}. Prefer value receivers: a registered middleware is copied before each JSON-dispatched call decodes its config, so a pointer receiver buys nothing.
The jsonschema_description tag is how a config field gets documented. Genkit infers the config schema from the struct without reading Go doc comments, and the Developer UI renders the tag’s text as the field’s tooltip. Keep constraints such as enum= in the separate jsonschema tag, whose comma-separated keyword list would otherwise cut a description off at its first comma.
Sharing state across hooks
Section titled “Sharing state across hooks”State that should be shared across the hooks of a single Generate call lives in closures captured by New. Each call gets a fresh Hooks bundle, so nothing leaks between calls:
type Counter struct{}
func (Counter) Name() string { return "mine/counter" }
func (Counter) New(ctx context.Context) (*ai.Hooks, error) { var modelCalls int return &ai.Hooks{ WrapModel: func(ctx context.Context, p *ai.ModelParams, next ai.ModelNext) (*ai.ModelResponse, error) { modelCalls++ return next(ctx, p) }, WrapGenerate: func(ctx context.Context, p *ai.GenerateParams, next ai.GenerateNext) (*ai.ModelResponse, error) { // The same `modelCalls` is visible here: both closures capture it from `New`. resp, err := next(ctx, p) logger.Debug(ctx, "iteration finished", "iteration", p.Iteration, "modelCalls", modelCalls) return resp, err }, }, nil}WrapTool may run concurrently for parallel tool calls in the same iteration, so any state it touches must be guarded with sync primitives:
func (Counter) New(ctx context.Context) (*ai.Hooks, error) { var ( mu sync.Mutex toolCalls int ) return &ai.Hooks{ WrapTool: func(ctx context.Context, p *ai.ToolParams, next ai.ToolNext) (*ai.MultipartToolResponse, error) { mu.Lock() toolCalls++ mu.Unlock() return next(ctx, p) }, }, nil}The built-in Filesystem middleware uses this pattern: New allocates a per-call file-state cache and a path-lock map, then the read, write, and edit tool implementations close over both.
Illustrative example: Guardrails middleware
Section titled “Illustrative example: Guardrails middleware”The following example demonstrates how to build an illustrative custom middleware that performs pre- and post-generation policy checks. Using WrapModel and WrapGenerate, the middleware can evaluate inputs before an expensive model invocation runs, and inspect outputs before returning them to the caller:
// Guardrail screens the request before the model runs, and checks the// response before the caller sees it. Both hooks use a classifier model.type Guardrail struct { Classifier string `json:"classifier,omitempty"` g *genkit.Genkit}
func (Guardrail) Name() string { return "mine/guardrail" }
func (gr Guardrail) classify(ctx context.Context, instruction, text string) (string, error) { verdict, err := genkit.GenerateText(ctx, gr.g, ai.WithModelName(gr.Classifier), ai.WithSystem("%s Answer with one word: ALLOW or BLOCK.", instruction), ai.WithPrompt("%s", text), ) if err != nil { return "", err } return strings.TrimSpace(verdict), nil}
func (gr Guardrail) New(context.Context) (*ai.Hooks, error) { return &ai.Hooks{ // Input guardrail: screen before running the primary model. WrapModel: func(ctx context.Context, p *ai.ModelParams, next ai.ModelNext) (*ai.ModelResponse, error) { var sb strings.Builder for _, m := range p.Request.Messages { if m.Role == ai.RoleUser { sb.WriteString(m.Text()) sb.WriteString("\n") } } verdict, err := gr.classify(ctx, "Decide whether this request is abusive.", sb.String()) if err != nil { return nil, err } if verdict != "ALLOW" { // Returning without calling next: the primary model never runs. return nil, status.Errorf(status.ErrFailedPrecondition, "input guardrail tripped: %s", verdict) } return next(ctx, p) },
// Output guardrail: check the response before returning to the caller. WrapGenerate: func(ctx context.Context, p *ai.GenerateParams, next ai.GenerateNext) (*ai.ModelResponse, error) { resp, err := next(ctx, p) if err != nil || resp.Text() == "" { return resp, err } verdict, err := gr.classify(ctx, "Decide whether this answer leaks private data.", resp.Text()) if err != nil { return nil, err } if verdict != "ALLOW" { return nil, status.Errorf(status.ErrFailedPrecondition, "output guardrail tripped: %s", verdict) } return resp, nil }, }, nil}The output hook can also modify or replace the response if desired, as WrapGenerate returns the *ai.ModelResponse received by the caller.
Similarly, WrapTool can intercept tool calls, allowing you to validate or reject arguments before tool execution occurs.
Plugin-provided middleware and plugin-level state
Section titled “Plugin-provided middleware and plugin-level state”Middleware shipped as part of a plugin needs two things the simple cases above don’t:
- A way to be registered automatically when the plugin is added to
genkit.Init, so the Dev UI and cross-runtime callers can address it by name. - A way to keep plugin-level state (an HTTP client, a logger, a database handle) that isn’t part of the JSON-serializable config.
Both are handled by implementing ai.MiddlewarePlugin on the plugin struct and putting plugin-level state on unexported fields of the config struct. The plugin’s Middlewares method passes a prototype with those fields populated to ai.NewMiddleware, which captures it in a build closure. Every JSON-dispatched call, whether from the Dev UI, a .prompt file’s use: entry, or a cross-runtime caller, copies the prototype and decodes its own config over the copy, so unexported state carries into each call and nothing one call sets leaks into the next.
Three rules follow. Exported fields are per-call user config and stay zero on the prototype: a call that omits one gets the zero value, so defaults belong in New. State that must be shared across calls, such as a client or a cache, goes behind a pointer, which survives the copy pointing at the same object. And register by value, as below; a pointer prototype is copied through to its pointee, but gains nothing.
import ( "context" "fmt" "io" "time"
"github.com/firebase/genkit/go/ai" "github.com/firebase/genkit/go/core/api")
type Logger struct { Prefix string `json:"prefix,omitempty" jsonschema_description:"Text written before each record."` out io.Writer // unexported; preserved across JSON dispatch by value-copy}
func (Logger) Name() string { return "mine/logger" }
func (l Logger) New(ctx context.Context) (*ai.Hooks, error) { return &ai.Hooks{ WrapModel: func(ctx context.Context, p *ai.ModelParams, next ai.ModelNext) (*ai.ModelResponse, error) { start := time.Now() resp, err := next(ctx, p) fmt.Fprintf(l.out, "%s model call took %s\n", l.Prefix, time.Since(start)) return resp, err }, }, nil}
type LoggerPlugin struct{ Out io.Writer }
func (p *LoggerPlugin) Name() string { return "mine/logger" }func (p *LoggerPlugin) Init(ctx context.Context) []api.Action { return nil }
func (p *LoggerPlugin) Middlewares(ctx context.Context) ([]*ai.MiddlewareDesc, error) { return []*ai.MiddlewareDesc{ ai.NewMiddleware("logs model call latency", Logger{out: p.Out}), }, nil}The io.Writer above is the plugin-level state the example is about, not a recommendation: middleware that only wants its records on the trace should call the core/logger helpers.
Application code then registers the plugin once during Init, which makes the middleware available everywhere by name:
g := genkit.Init(ctx, genkit.WithPlugins( &googlegenai.GoogleAI{}, &LoggerPlugin{Out: os.Stderr},))
resp, err := genkit.Generate(ctx, g, ai.WithPrompt("Hello"), ai.WithUse(Logger{Prefix: "[trace]"}),)When the Dev UI dispatches the same middleware with JSON like {"prefix": "[debug]"}, Genkit value-copies the prototype to recreate the config: out (which isn’t in JSON) is preserved from the plugin’s prototype, while the unmarshaled JSON overrides Prefix.
The built-in plugins/middleware package follows exactly this pattern. See plugin.go for a minimal real-world example.
Application-owned middleware
Section titled “Application-owned middleware”When your application code defines a middleware directly rather than wrapping it in a plugin, use genkit.DefineMiddleware to register it with the Genkit instance:
genkit.DefineMiddleware(g, "logs model call latency", Logger{out: os.Stderr})Registration surfaces the middleware in the Dev UI and lets cross-runtime callers reference it by name. For pure Go use, registration is not required: passing a middleware value directly to ai.WithUse invokes its New method on the local fast path. Registration is what makes the middleware visible to the Dev UI.
Inline middleware
Section titled “Inline middleware”For ad-hoc middleware that doesn’t need a named type or Dev UI visibility, use ai.MiddlewareFunc:
ai.WithUse(ai.MiddlewareFunc(func(ctx context.Context) (*ai.Hooks, error) { return &ai.Hooks{ WrapModel: func(ctx context.Context, p *ai.ModelParams, next ai.ModelNext) (*ai.ModelResponse, error) { logger.Debug(ctx, "model call", "messages", len(p.Request.Messages)) return next(ctx, p) }, }, nil}))The adapter satisfies Middleware with a placeholder name. Inline middleware is resolved on the local fast path and never touches the registry, so the placeholder is fine.
It works on genkit.Generate and on Prompt.Execute, but not in genkit.DefinePrompt: a prompt action serializes its options and a function value has no JSON form.
Composition order
Section titled “Composition order”ai.WithUse(A, B, C) composes left to right with the first listed middleware as the outermost wrapper, like HTTP middleware: at call time the chain expands to A { B { C { actual } } }. Each layer’s next continuation runs the next inner layer:
ai.WithUse( &middleware.Retry{MaxRetries: 3}, // outer: retries the whole inner stack &middleware.Fallback{Models: fallbackModels}, // inner: tries fallback models on failure)// effective chain: Retry { Fallback { model } }Order matters. Retry outside Fallback retries the entire fallback cascade as a unit. Swap them and you’d retry the primary first and fall back only after exhausting retries.
For more complex examples of building custom middleware, you can refer to the source code of the built-in middleware in the Genkit GitHub repository.