Skip to content

Gemini Enterprise plugin

The Gemini Enterprise plugin provides access to Gemini Enterprise, offering capabilities for building, scaling, and governing agents alongside model access, grounding, Vector Search, Model Garden, and evaluation metrics.

Accessing Google GenAI Models via the Gemini Enterprise API

Section titled “Accessing Google GenAI Models via the Gemini Enterprise API”

All languages support accessing Google’s generative AI models (Gemini, Imagen, etc.) via the Gemini Enterprise API with enterprise authentication and features.

The unified Google GenAI plugin provides access to models via the Gemini Enterprise API using the VertexAI initializer:

Terminal window
uv add genkit-google-genai
from genkit import Genkit
from genkit_google_genai import VertexAI
ai = Genkit(
plugins=[
VertexAI(location='us-central1'), # Regional endpoint
# VertexAI(location='global'), # Global endpoint
],
)

Authentication Methods:

  • Application Default Credentials (ADC): The standard method for most Gemini Enterprise use cases, especially in production. It uses the credentials from the environment (e.g., service account on GCP, user credentials from gcloud auth application-default login locally). This method requires a Google Cloud Project with billing enabled and the Gemini Enterprise API (aiplatform.googleapis.com) enabled.
  • Gemini Enterprise Express Mode: A streamlined way to try out many Gemini Enterprise features using just an API key, without needing to set up billing or full project configurations. This is ideal for quick experimentation and has generous free tier quotas. Learn More about Express Mode.
# Using Gemini Enterprise Express Mode (easy to start; some limitations).
# Get an API key from the Gemini Enterprise Express Mode setup.
import os
from genkit import Genkit
from genkit_google_genai import VertexAI
ai = Genkit(
plugins=[
VertexAI(api_key=os.environ['VERTEX_EXPRESS_API_KEY']),
],
)

Note: When using Express Mode, you typically omit project and location on VertexAI (see the Express Mode docs).

Gemini 3 Series - Latest models with state-of-the-art reasoning and multimodal capabilities:

  • gemini-3.8-flash - Most intelligent Flash model, engineered for complex reasoning, coding, and agentic workflows
  • gemini-3.1-pro-preview - Preview of the most capable model for complex reasoning and problem solving
  • gemini-3.5-flash-lite - Fastest, most cost-effective model for high-throughput execution
  • gemini-3.1-flash-image - Fast and efficient image generation and editing
  • gemini-3-pro-image - State-of-the-art image generation and editing for complex visual tasks
from genkit import Genkit
from genkit_google_genai import VertexAI
ai = Genkit(
plugins=[VertexAI(location='us-central1')],
)
response = await ai.generate(
model='vertexai/gemini-pro-latest',
prompt='Explain Gemini Enterprise in simple terms.',
)
print(response.text)
embeddings = await ai.embed(
embedder='vertexai/text-embedding-005',
content='Embed this text.',
)
response = await ai.generate(
model='vertexai/imagen-3.0-generate-002',
prompt='A beautiful watercolor painting of a castle in the mountains.',
)
if response.media:
generated_image = response.media[0].url
response = await ai.generate(
model='vertexai/gemini-3-pro-preview',
prompt='what is heavier, one kilo of steel or one kilo of feathers',
config={
'thinking_config': {
'thinking_level': 'HIGH', # Or 'LOW' or 'MEDIUM'
},
},
)
message = (await ai.generate(
model='vertexai/gemini-pro-latest',
prompt='what is heavier, one kilo of steel or one kilo of feathers',
config={
'thinking_config': {
'thinking_budget': 1024,
'include_thoughts': True,
},
},
)).message

The following advanced features are available in Python. Note that some features require additional plugin packages:

Core Gemini Enterprise features (included in genkit-google-genai):

Terminal window
uv add genkit-google-genai

Model Garden (separate package):

Terminal window
uv add genkit-vertexai

If you want to locally run flows that use these plugins, you also need the Google Cloud CLI tool installed.

from genkit import Genkit
from genkit_google_genai import VertexAI
ai = Genkit(
plugins=[VertexAI(location='us-central1')],
)

The plugin requires you to specify your Google Cloud project ID, the region to which you want to make Gemini Enterprise API requests, and your Google Cloud project credentials.

  • You can specify your Google Cloud project ID either by setting project in the VertexAI() configuration or by setting the GCLOUD_PROJECT environment variable. If you’re running your flow from a Google Cloud environment (Cloud Functions, Cloud Run, and so on), GCLOUD_PROJECT is automatically set to the project ID of the environment.
  • You can specify the API location either by setting location in the VertexAI() configuration or by setting the GCLOUD_LOCATION environment variable.
  • To provide API credentials, you need to set up Google Cloud Application Default Credentials.
    1. To specify your credentials:
      • If you’re running your flow from a Google Cloud environment (Cloud Functions, Cloud Run, and so on), this is set automatically.
      • On your local dev environment, do this by running:
Terminal window
gcloud auth application-default login --project YOUR_PROJECT_ID
  1. In addition, make sure the account is granted the roles/aiplatform.user IAM role. See the Gemini Enterprise access control docs.

This plugin supports grounding Gemini text responses using Google Search.

Important: Gemini Enterprise charges a fee for grounding requests in addition to the cost of making LLM requests. See the Gemini Enterprise pricing page and be sure you understand grounding request pricing before you use this feature.

Example:

ai = Genkit(
plugins=[VertexAI(location='us-central1')],
)
await ai.generate(
model='vertexai/gemini-flash-latest',
prompt='What are the latest developments in quantum computing?',
config={
'google_search_retrieval': {
'disable_attribution': True,
},
}
)

The Gemini Enterprise Genkit plugin supports Context Caching, which allows models to reuse previously cached content to optimize token usage when dealing with large pieces of content. This feature is especially useful for conversational flows or scenarios where the model references a large piece of content consistently across multiple requests.

To enable context caching, ensure your model supports it. For example, gemini-3.8-flash is a generation 3 model that supports context caching. Note that context caching cannot be combined with tool calls or system prompts in the same request.

You can define a caching mechanism in your application like this:

from genkit import Message, Part, TextPart, Role
ai = Genkit(
plugins=[VertexAI(location='us-central1')],
)
llm_response = await ai.generate(
messages=[
Message(
role=Role.USER,
content=[Part(root=TextPart(text='Here is the relevant text from War and Peace.'))],
),
Message(
role=Role.MODEL,
content=[
Part(root=TextPart(text="Based on War and Peace, here is some analysis of Pierre Bezukhov's character.")),
],
metadata={
'cache': {
'ttl_seconds': 300, # Cache this message for 5 minutes
},
},
),
],
model='vertexai/gemini-3.8-flash',
prompt="Describe Pierre's transformation throughout the novel.",
)

In this setup:

  • messages: Allows you to pass conversation history.
  • metadata.cache.ttl_seconds: Specifies the time-to-live (TTL) for caching a specific response.

Example: Leveraging Large Texts with Context

Section titled “Example: Leveraging Large Texts with Context”

For applications referencing long documents, such as War and Peace or Lord of the Rings, you can structure your queries to reuse cached contexts:

from pathlib import Path
from genkit import Message, Part, TextPart, Role
text_content = Path('path/to/war_and_peace.txt').read_text()
llm_response = await ai.generate(
messages=[
Message(
role=Role.USER,
content=[Part(root=TextPart(text=text_content))], # Include the large text as context
),
Message(
role=Role.MODEL,
content=[
Part(root=TextPart(text='This analysis is based on the provided text from War and Peace.')),
],
metadata={
'cache': {
'ttl_seconds': 300, # Cache the response to avoid reloading the full text
},
},
),
],
model='vertexai/gemini-3.8-flash',
prompt='Analyze the relationship between Pierre and Natasha.',
)

Supported models: gemini-3.8-flash

Access third-party models through Model Garden in Gemini Enterprise using the genkit-vertexai package (ModelGarden). The plugin requires a Google Cloud project ID: pass project_id, or set GCLOUD_PROJECT / GOOGLE_CLOUD_PROJECT. Model IDs must use the publisher-qualified names shown in the Google Cloud console (for example meta/... for Llama, anthropic/... for Claude in Model Garden). Pass them to model_garden_name() so Genkit resolves the action as modelgarden/<model-id>.

Installation:

Terminal window
uv add genkit-vertexai
from genkit import Genkit
from genkit_vertexai.model_garden import ModelGarden, model_garden_name
ai = Genkit(
plugins=[
ModelGarden(
project_id='my-gcp-project',
location='us-central1',
),
],
)
response = await ai.generate(
model=model_garden_name('meta/llama-3.1-405b-instruct-maas'),
prompt='Write a function that adds two numbers together',
)

Another identifier shipped in the Python SDK registry is meta/llama-3.2-90b-vision-instruct-maas. Always confirm the exact model resource name for your project in the Model Garden console.

Claude in Model Garden uses anthropic/... model IDs. Version strings often include dates or @ — use the exact ID from the console:

from genkit import Genkit
from genkit_vertexai.model_garden import ModelGarden, model_garden_name
ai = Genkit(
plugins=[
ModelGarden(
project_id='my-gcp-project',
location='us-central1',
),
],
)
response = await ai.generate(
model=model_garden_name('anthropic/claude-haiku-4-5@20251001'),
prompt='What should I do when I visit Melbourne?',
)

Other OpenAI-compatible Model Garden endpoints

Section titled “Other OpenAI-compatible Model Garden endpoints”

For additional publishers (for example Mistral), use the same model_garden_name() pattern with the full Model Garden model ID. Models not in the built-in registry still resolve via the generic OpenAI-compatible Model Garden path.

Gemini Enterprise provides access to various third-party models through Model Garden. Consult the Model Garden documentation for the full list of supported models and their capabilities.

Genkit provides evaluation metrics through the Gemini Enterprise plugin (VertexAI) automatically when a project is configured:

from genkit import Genkit
from genkit_google_genai import VertexAI
# Evaluators are automatically registered when a project ID is provided
ai = Genkit(
plugins=[VertexAI(project='your-project-id', location='us-central1')],
)

Available built-in metrics from the Gemini Enterprise plugin include:

  • BLEU: Translation quality
  • ROUGE: Summarization quality
  • Fluency: Text fluency
  • Safety: Content safety
  • Groundedness: Factual accuracy
  • Summarization Quality/Helpfulness/Verbosity: Summary evaluation

See the evaluation documentation for more details on implementing comprehensive evaluation workflows.

  • Learn about generating content to understand how to use these models effectively
  • Explore evaluation to leverage Gemini Enterprise evaluation metrics
  • See RAG to implement retrieval-augmented generation
  • Check out creating flows to build structured AI workflows
  • For simple API key access, see the Google AI plugin