Vertex AI Gemini API: Enterprise Integration Guide

Gemini is Google's most capable foundation model family, accessible through Vertex AI for enterprise workloads. This guide covers model selection, system instructions, multi-turn conversations, streaming responses, function calling, safety filters, and production deployment patterns.

Gemini's capabilities are impressive in Google AI Studio demos. Making those capabilities work reliably in a production enterprise application requires a different kind of thinking: model selection for cost vs capability trade-offs, system instruction design that shapes model behavior consistently, safety filter configuration appropriate for your use case, and latency management for interactive applications.

This guide is for engineering teams building production AI applications on GCP using Vertex AI as the Gemini access point. We'll cover API setup, model variants, prompt engineering patterns, function calling for agentic workflows, streaming for real-time UX, and cost estimation.

Why Vertex AI for Gemini (vs Google AI Studio / Generative AI API)

Both Vertex AI and the Gemini API provide access to Gemini models. For enterprise deployments:

Feature Vertex AI Gemini API (AI Studio)
Data residency Regional (europe-west4, etc.) Google-managed
VPC networking Yes (private access) No
IAM access control GCP IAM API key
CMEK encryption Yes No
Enterprise SLAs Yes No

For any application handling sensitive data, Vertex AI is the right choice: data residency guarantees, private VPC routing, and GCP IAM access control.

Setting Up Access

gcloud services enable aiplatform.googleapis.com

gcloud iam service-accounts create gemini-app-sa   --display-name="Gemini Application"

gcloud projects add-iam-policy-binding my-project   --member=serviceAccount:gemini-app-sa@my-project.iam.gserviceaccount.com   --role=roles/aiplatform.user
import vertexai
from vertexai.generative_models import GenerativeModel, GenerationConfig, SafetySetting, HarmCategory, HarmBlockThreshold

vertexai.init(project="my-project", location="europe-west4")

Model Variant Selection

Model Context Window Best For Cost (per 1M tokens)
gemini-1.5-flash-001 1M tokens High-volume, cost-sensitive ~$0.075 in / $0.30 out
gemini-1.5-pro-001 2M tokens Complex reasoning, code ~$3.50 in / $10.50 out
gemini-2.0-flash 1M tokens Fastest, multimodal ~$0.10 in / $0.40 out

Rule of thumb: Start with Flash for 80%+ of use cases. Use Pro only where Flash makes errors.

System Instructions: Shaping Model Behavior

from vertexai.generative_models import GenerativeModel

support_bot = GenerativeModel(
    model_name="gemini-1.5-flash-001",
    system_instruction="""You are a customer support assistant for Warqline, a cloud consulting firm.

Your responsibilities:
- Answer questions about AWS and GCP cloud services
- Help users understand their Well-Architected assessment results
- Escalate billing issues to the finance team (do not attempt to resolve them)
- Never make promises about service capabilities or timelines

Response style:
- Be concise — aim for responses under 150 words unless detail is needed
- Use bullet points for lists of steps
- Ask a clarifying question when the user's request is ambiguous

What you must never do:
- Share information about other customers
- Speculate about future product features
- Provide legal or financial advice
"""
)

Multi-Turn Conversations

from vertexai.generative_models import GenerativeModel, ChatSession

class ConversationManager:
    def __init__(self, model_name: str = "gemini-1.5-flash-001", system_instruction: str = ""):
        self.model = GenerativeModel(
            model_name=model_name,
            system_instruction=system_instruction,
        )

    def start_session(self) -> ChatSession:
        return self.model.start_chat(history=[])

    def send_message(self, session: ChatSession, message: str, temperature: float = 0.7) -> str:
        response = session.send_message(
            message,
            generation_config=GenerationConfig(
                temperature=temperature,
                max_output_tokens=500,
                top_p=0.95,
            ),
        )
        return response.text

manager = ConversationManager(system_instruction="You are a helpful cloud advisor.")
session = manager.start_session()

print(manager.send_message(session, "What is Cloud Run?"))
print(manager.send_message(session, "How does its pricing compare to GKE?"))

Sessions are in-memory — serialize history to Redis or Cloud Firestore keyed by session ID for web applications.

Function Calling for Agentic Workflows

from vertexai.generative_models import GenerativeModel, Tool, FunctionDeclaration, Part

get_customer_info = FunctionDeclaration(
    name="get_customer_info",
    description="Get account information for a customer by their email address",
    parameters={
        "type": "object",
        "properties": {
            "email": {"type": "string", "description": "The customer's email address"}
        },
        "required": ["email"],
    },
)

list_assessments = FunctionDeclaration(
    name="list_assessments",
    description="List Well-Architected assessments for a customer account",
    parameters={
        "type": "object",
        "properties": {
            "customer_id": {"type": "string", "description": "The customer's ID"},
            "status": {"type": "string", "enum": ["pending", "in_progress", "complete"]},
        },
        "required": ["customer_id"],
    },
)

tool = Tool(function_declarations=[get_customer_info, list_assessments])

model = GenerativeModel(
    "gemini-1.5-pro-001",
    tools=[tool],
    system_instruction="You are a support agent. Use the tools to look up customer information before answering.",
)

def handle_support_query(query: str) -> str:
    chat = model.start_chat()
    response = chat.send_message(query)

    max_iterations = 5
    iteration = 0

    while (response.candidates and
           response.candidates[0].content.parts and
           response.candidates[0].content.parts[0].function_call.name and
           iteration < max_iterations):

        function_call = response.candidates[0].content.parts[0].function_call
        function_name = function_call.name
        function_args = dict(function_call.args)

        if function_name == "get_customer_info":
            result = get_customer_from_db(function_args["email"])
        elif function_name == "list_assessments":
            result = get_assessments_from_db(function_args["customer_id"])
        else:
            result = {"error": f"Unknown function: {function_name}"}

        response = chat.send_message(
            Part.from_function_response(name=function_name, response={"content": result})
        )
        iteration += 1

    return response.text

Streaming for Real-Time UX

from flask import Flask, Response, stream_with_context, request
import vertexai
from vertexai.generative_models import GenerativeModel
import json

app = Flask(__name__)
model = GenerativeModel("gemini-1.5-flash-001")

@app.route("/chat/stream", methods=["POST"])
def chat_stream():
    body = request.get_json()
    user_message = body.get("message", "")

    def generate():
        for response in model.generate_content(user_message, stream=True):
            if response.text:
                yield f"data: {json.dumps({'text': response.text})}

"
        yield "data: [DONE]

"

    return Response(
        stream_with_context(generate()),
        mimetype="text/event-stream",
        headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
    )

Frontend consumption:

async function streamChat(message) {
  const response = await fetch('/chat/stream', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ message }),
  });

  const reader = response.body.getReader();
  const decoder = new TextDecoder();

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    const lines = decoder.decode(value).split('

').filter(Boolean);
    for (const line of lines) {
      if (line.startsWith('data: ')) {
        const data = line.slice(6);
        if (data === '[DONE]') return;
        appendToOutput(JSON.parse(data).text);
      }
    }
  }
}

Safety Filter Configuration

from vertexai.generative_models import SafetySetting, HarmCategory, HarmBlockThreshold

default_model = GenerativeModel("gemini-1.5-flash-001")

# Adjusted for specific enterprise use cases
configured_model = GenerativeModel(
    "gemini-1.5-pro-001",
    safety_settings=[
        SafetySetting(
            category=HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT,
            threshold=HarmBlockThreshold.BLOCK_ONLY_HIGH,
        ),
        SafetySetting(
            category=HarmCategory.HARM_CATEGORY_HARASSMENT,
            threshold=HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE,
        ),
    ],
)

# Diagnose blocked responses
response = model.generate_content(prompt)
if not response.candidates:
    if response.prompt_feedback.block_reason:
        print(f"Blocked: {response.prompt_feedback.block_reason}")
elif response.candidates[0].finish_reason.name == "SAFETY":
    for rating in response.candidates[0].safety_ratings:
        if rating.blocked:
            print(f"Blocked by: {rating.category.name}")

Cost Optimization

def analyze_cost(response, model_name: str = "gemini-1.5-flash-001") -> dict:
    PRICING = {
        "gemini-1.5-flash-001": {"input": 0.075 / 1_000_000, "output": 0.30 / 1_000_000},
        "gemini-1.5-pro-001": {"input": 3.50 / 1_000_000, "output": 10.50 / 1_000_000},
    }

    usage = response.usage_metadata
    pricing = PRICING.get(model_name, PRICING["gemini-1.5-flash-001"])

    return {
        "input_tokens": usage.prompt_token_count,
        "output_tokens": usage.candidates_token_count,
        "total_cost_usd": (
            usage.prompt_token_count * pricing["input"] +
            usage.candidates_token_count * pricing["output"]
        ),
    }

Key cost levers:

  1. Use Flash instead of Pro where quality allows — 46x cheaper on output
  2. Limit max_output_tokens — stop generation when you have enough
  3. Use batch prediction for non-interactive bulk processing
  4. Cache repeated system prompts with Vertex AI context caching

For RAG applications with Gemini, see our Vertex AI RAG guide. For MLOps workflows, see our Vertex AI MLOps guide.