Vertex AI Gemini API: Enterprise Integration Guide
Gemini is Google's most capable foundation model family, accessible through Vertex AI for enterprise workloads. This guide covers model selection, system instructions, multi-turn conversations, streaming responses, function calling, safety filters, and production deployment patterns.
Gemini's capabilities are impressive in Google AI Studio demos. Making those capabilities work reliably in a production enterprise application requires a different kind of thinking: model selection for cost vs capability trade-offs, system instruction design that shapes model behavior consistently, safety filter configuration appropriate for your use case, and latency management for interactive applications.
This guide is for engineering teams building production AI applications on GCP using Vertex AI as the Gemini access point. We'll cover API setup, model variants, prompt engineering patterns, function calling for agentic workflows, streaming for real-time UX, and cost estimation.
Why Vertex AI for Gemini (vs Google AI Studio / Generative AI API)
Both Vertex AI and the Gemini API provide access to Gemini models. For enterprise deployments:
| Feature | Vertex AI | Gemini API (AI Studio) |
|---|---|---|
| Data residency | Regional (europe-west4, etc.) | Google-managed |
| VPC networking | Yes (private access) | No |
| IAM access control | GCP IAM | API key |
| CMEK encryption | Yes | No |
| Enterprise SLAs | Yes | No |
For any application handling sensitive data, Vertex AI is the right choice: data residency guarantees, private VPC routing, and GCP IAM access control.
Setting Up Access
gcloud services enable aiplatform.googleapis.com
gcloud iam service-accounts create gemini-app-sa --display-name="Gemini Application"
gcloud projects add-iam-policy-binding my-project --member=serviceAccount:gemini-app-sa@my-project.iam.gserviceaccount.com --role=roles/aiplatform.user
import vertexai
from vertexai.generative_models import GenerativeModel, GenerationConfig, SafetySetting, HarmCategory, HarmBlockThreshold
vertexai.init(project="my-project", location="europe-west4")
Model Variant Selection
| Model | Context Window | Best For | Cost (per 1M tokens) |
|---|---|---|---|
| gemini-1.5-flash-001 | 1M tokens | High-volume, cost-sensitive | ~$0.075 in / $0.30 out |
| gemini-1.5-pro-001 | 2M tokens | Complex reasoning, code | ~$3.50 in / $10.50 out |
| gemini-2.0-flash | 1M tokens | Fastest, multimodal | ~$0.10 in / $0.40 out |
Rule of thumb: Start with Flash for 80%+ of use cases. Use Pro only where Flash makes errors.
System Instructions: Shaping Model Behavior
from vertexai.generative_models import GenerativeModel
support_bot = GenerativeModel(
model_name="gemini-1.5-flash-001",
system_instruction="""You are a customer support assistant for Warqline, a cloud consulting firm.
Your responsibilities:
- Answer questions about AWS and GCP cloud services
- Help users understand their Well-Architected assessment results
- Escalate billing issues to the finance team (do not attempt to resolve them)
- Never make promises about service capabilities or timelines
Response style:
- Be concise — aim for responses under 150 words unless detail is needed
- Use bullet points for lists of steps
- Ask a clarifying question when the user's request is ambiguous
What you must never do:
- Share information about other customers
- Speculate about future product features
- Provide legal or financial advice
"""
)
Multi-Turn Conversations
from vertexai.generative_models import GenerativeModel, ChatSession
class ConversationManager:
def __init__(self, model_name: str = "gemini-1.5-flash-001", system_instruction: str = ""):
self.model = GenerativeModel(
model_name=model_name,
system_instruction=system_instruction,
)
def start_session(self) -> ChatSession:
return self.model.start_chat(history=[])
def send_message(self, session: ChatSession, message: str, temperature: float = 0.7) -> str:
response = session.send_message(
message,
generation_config=GenerationConfig(
temperature=temperature,
max_output_tokens=500,
top_p=0.95,
),
)
return response.text
manager = ConversationManager(system_instruction="You are a helpful cloud advisor.")
session = manager.start_session()
print(manager.send_message(session, "What is Cloud Run?"))
print(manager.send_message(session, "How does its pricing compare to GKE?"))
Sessions are in-memory — serialize history to Redis or Cloud Firestore keyed by session ID for web applications.
Function Calling for Agentic Workflows
from vertexai.generative_models import GenerativeModel, Tool, FunctionDeclaration, Part
get_customer_info = FunctionDeclaration(
name="get_customer_info",
description="Get account information for a customer by their email address",
parameters={
"type": "object",
"properties": {
"email": {"type": "string", "description": "The customer's email address"}
},
"required": ["email"],
},
)
list_assessments = FunctionDeclaration(
name="list_assessments",
description="List Well-Architected assessments for a customer account",
parameters={
"type": "object",
"properties": {
"customer_id": {"type": "string", "description": "The customer's ID"},
"status": {"type": "string", "enum": ["pending", "in_progress", "complete"]},
},
"required": ["customer_id"],
},
)
tool = Tool(function_declarations=[get_customer_info, list_assessments])
model = GenerativeModel(
"gemini-1.5-pro-001",
tools=[tool],
system_instruction="You are a support agent. Use the tools to look up customer information before answering.",
)
def handle_support_query(query: str) -> str:
chat = model.start_chat()
response = chat.send_message(query)
max_iterations = 5
iteration = 0
while (response.candidates and
response.candidates[0].content.parts and
response.candidates[0].content.parts[0].function_call.name and
iteration < max_iterations):
function_call = response.candidates[0].content.parts[0].function_call
function_name = function_call.name
function_args = dict(function_call.args)
if function_name == "get_customer_info":
result = get_customer_from_db(function_args["email"])
elif function_name == "list_assessments":
result = get_assessments_from_db(function_args["customer_id"])
else:
result = {"error": f"Unknown function: {function_name}"}
response = chat.send_message(
Part.from_function_response(name=function_name, response={"content": result})
)
iteration += 1
return response.text
Streaming for Real-Time UX
from flask import Flask, Response, stream_with_context, request
import vertexai
from vertexai.generative_models import GenerativeModel
import json
app = Flask(__name__)
model = GenerativeModel("gemini-1.5-flash-001")
@app.route("/chat/stream", methods=["POST"])
def chat_stream():
body = request.get_json()
user_message = body.get("message", "")
def generate():
for response in model.generate_content(user_message, stream=True):
if response.text:
yield f"data: {json.dumps({'text': response.text})}
"
yield "data: [DONE]
"
return Response(
stream_with_context(generate()),
mimetype="text/event-stream",
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
)
Frontend consumption:
async function streamChat(message) {
const response = await fetch('/chat/stream', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ message }),
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
const lines = decoder.decode(value).split('
').filter(Boolean);
for (const line of lines) {
if (line.startsWith('data: ')) {
const data = line.slice(6);
if (data === '[DONE]') return;
appendToOutput(JSON.parse(data).text);
}
}
}
}
Safety Filter Configuration
from vertexai.generative_models import SafetySetting, HarmCategory, HarmBlockThreshold
default_model = GenerativeModel("gemini-1.5-flash-001")
# Adjusted for specific enterprise use cases
configured_model = GenerativeModel(
"gemini-1.5-pro-001",
safety_settings=[
SafetySetting(
category=HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT,
threshold=HarmBlockThreshold.BLOCK_ONLY_HIGH,
),
SafetySetting(
category=HarmCategory.HARM_CATEGORY_HARASSMENT,
threshold=HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE,
),
],
)
# Diagnose blocked responses
response = model.generate_content(prompt)
if not response.candidates:
if response.prompt_feedback.block_reason:
print(f"Blocked: {response.prompt_feedback.block_reason}")
elif response.candidates[0].finish_reason.name == "SAFETY":
for rating in response.candidates[0].safety_ratings:
if rating.blocked:
print(f"Blocked by: {rating.category.name}")
Cost Optimization
def analyze_cost(response, model_name: str = "gemini-1.5-flash-001") -> dict:
PRICING = {
"gemini-1.5-flash-001": {"input": 0.075 / 1_000_000, "output": 0.30 / 1_000_000},
"gemini-1.5-pro-001": {"input": 3.50 / 1_000_000, "output": 10.50 / 1_000_000},
}
usage = response.usage_metadata
pricing = PRICING.get(model_name, PRICING["gemini-1.5-flash-001"])
return {
"input_tokens": usage.prompt_token_count,
"output_tokens": usage.candidates_token_count,
"total_cost_usd": (
usage.prompt_token_count * pricing["input"] +
usage.candidates_token_count * pricing["output"]
),
}
Key cost levers:
- Use Flash instead of Pro where quality allows — 46x cheaper on output
- Limit
max_output_tokens— stop generation when you have enough - Use batch prediction for non-interactive bulk processing
- Cache repeated system prompts with Vertex AI context caching
For RAG applications with Gemini, see our Vertex AI RAG guide. For MLOps workflows, see our Vertex AI MLOps guide.