Firestore vs Bigtable vs Datastore: NoSQL Options on Google Cloud
Google Cloud has three distinct NoSQL databases — Firestore, Cloud Bigtable, and Datastore (now the legacy mode of Firestore). Each optimizes for very different access patterns. Choosing the wrong one leads to either terrible performance or unnecessary complexity. This guide makes the distinction clear.
If you're building on Google Cloud and need a NoSQL database, you have three options that look superficially similar but are architecturally very different: Firestore (the modern, document-oriented option), Cloud Bigtable (the wide-column powerhouse for analytics and time-series), and Cloud Datastore (technically Firestore in Datastore mode — the legacy choice).
The good news: the decision is usually straightforward once you understand the access patterns each database optimizes for. The bad news: teams regularly pick Firestore for time-series workloads that need Bigtable, or use Bigtable for simple document storage where Firestore would be half the complexity.
Datastore: The Legacy Option
Let's dispatch this quickly. Cloud Datastore is now Firestore in Datastore mode. It exists for backwards compatibility with applications built on the original App Engine Datastore.
If you're starting a new project, use Firestore natively. If you have an existing Datastore application, you don't need to migrate urgently — it's fully supported and runs on the same backend as Firestore. But don't choose Datastore for new work.
Key limitation: Firestore in Datastore mode doesn't support Firestore's real-time listeners, and vice versa. Choose at project creation time.
Firestore: Document Database for Applications
Firestore is a serverless, document-oriented database designed for web and mobile applications. Its design philosophy: make it easy to store and sync hierarchical data with low operational overhead.
Firestore Data Model
Firestore organizes data as documents in collections. Documents can contain subcollections, creating hierarchical data structures:
users/ (collection)
user:abc123/ (document)
name: "Alice"
email: "alice@example.com"
preferences: { theme: "dark" } (nested map)
orders/ (subcollection)
order:xyz789/ (document)
amount: 4999
currency: "EUR"
items: [...] (array of maps)
Firestore Strengths
Real-time synchronization: Firestore's killer feature is live data sync. Register a listener, and your application receives updates within milliseconds when data changes:
// Real-time listener — fires immediately and on every change
const unsubscribe = db.collection('orders')
.where('userId', '==', currentUserId)
.where('status', '==', 'active')
.orderBy('createdAt', 'desc')
.onSnapshot((snapshot) => {
snapshot.docChanges().forEach((change) => {
if (change.type === 'added') updateUI(change.doc.data());
if (change.type === 'modified') updateUI(change.doc.data());
if (change.type === 'removed') removeFromUI(change.doc.id);
});
});
Offline support: Firestore's SDKs cache data locally. Mobile and web apps work offline and sync changes when connectivity is restored.
Serverless scaling: Firestore scales reads and writes automatically. No provisioning, no connection limits, no shard management.
Strong consistency: Single-document reads and writes are strongly consistent. Multi-document transactions are supported.
Firestore Querying with Python
from google.cloud import firestore
db = firestore.Client()
# Simple document read
user_ref = db.collection('users').document('user:abc123')
user = user_ref.get()
if user.exists:
print(user.to_dict())
# Query with filters
orders = db.collection('orders').where(
filter=firestore.FieldFilter('userId', '==', 'user:abc123')
).where(
filter=firestore.FieldFilter('status', '==', 'completed')
).order_by('createdAt', direction=firestore.Query.DESCENDING).limit(20)
for order in orders.stream():
print(order.id, order.to_dict())
# Batch write (atomic)
batch = db.batch()
batch.set(db.collection('users').document('new-user'), {'name': 'Bob', 'email': 'bob@example.com'})
batch.update(db.collection('counters').document('user-count'), {'count': firestore.Increment(1)})
batch.commit()
Firestore Limitations
- Limited query flexibility: No joins. No aggregations across documents (COUNT, SUM) without client-side processing or BigQuery export. Queries require composite indexes for multi-field filters.
- Document size limit: 1 MB per document. Arrays are limited to 20,000 items.
- Write throughput: ~1 write/second per document for sustained writes. High-frequency updates to the same document (e.g., a counter) need distributed counters.
- Not for analytics: Firestore is OLTP. For analytical queries, export to BigQuery.
When to Choose Firestore
- User-facing applications with real-time UI requirements
- Mobile apps with offline capability needs
- Hierarchical data (users → orders → items) with document-oriented access
- Low-traffic to moderate-traffic applications (up to millions of documents)
- When you want serverless scaling without capacity planning
Cloud Bigtable: Wide-Column for Extreme Scale
Cloud Bigtable is a distributed wide-column store that powers Google Search indexing, Google Maps, and Google Analytics internally. It's designed for massive-scale time-series, event, and IoT data — not for general-purpose application storage.
Bigtable Data Model
Bigtable organizes data as rows identified by a row key, with column families containing columns, and cells containing multiple timestamped versions of data:
Row Key: device:sensor-123:2025-01-15T10:23:45
Column Family: metrics
temperature: 23.4 (timestamp: 2025-01-15T10:23:45)
temperature: 22.8 (timestamp: 2025-01-15T10:23:44)
humidity: 65.2 (timestamp: 2025-01-15T10:23:45)
Column Family: metadata
location: "building-A/floor-2"
device_type: "env-sensor"
The row key is the only index. Query patterns are row-key based: read a specific row, read a range of rows, or scan a prefix. There are no secondary indexes in the traditional sense.
Bigtable Row Key Design is Everything
Row key design determines performance. The cardinal rule: design row keys to match your access pattern and distribute writes evenly across tablets (Bigtable's distributed storage units).
Good row key designs for time-series IoT:
device:{device_id}#{reversed_timestamp}
Enables: "get last 100 readings for device X"
Bad (creates hotspot at beginning of key space):
{timestamp}#{device_id}
Problem: All recent data goes to same tablet
Sensor reading examples:
Good: sensor:abc123#9999999999999 (reversed epoch)
Bad: 2025-01-15T10:23:45#abc123 (sequential timestamps)
from google.cloud import bigtable
from google.cloud.bigtable import column_family, row_filters
import struct
import time
# Initialize client
client = bigtable.Client(project='my-project', admin=True)
instance = client.instance('my-instance')
table = instance.table('sensor-data')
# Write a sensor reading
# Use reversed timestamp to avoid hotspot
reversed_ts = 9999999999999 - int(time.time() * 1000)
row_key = f"sensor:abc123#{reversed_ts:013d}".encode()
row = table.direct_row(row_key)
row.set_cell(
column_family_id='metrics',
column='temperature',
value=struct.pack('>f', 23.4),
)
row.set_cell(
column_family_id='metrics',
column='humidity',
value=struct.pack('>f', 65.2),
)
row.commit()
# Read the last 100 readings for sensor abc123
prefix = "sensor:abc123#".encode()
rows = table.read_rows(
start_key=prefix,
end_key=prefix + b"ÿ",
limit=100,
filter_=row_filters.CellsColumnLimitFilter(1), # Latest value per column
)
for row in rows:
ts_encoded = row.row_key.decode().split('#')[1]
actual_ts = 9999999999999 - int(ts_encoded)
temp_bytes = row.cells['metrics'][b'temperature'][0].value
temp = struct.unpack('>f', temp_bytes)[0]
print(f"Timestamp: {actual_ts}, Temperature: {temp:.1f}°C")
Bigtable Performance and Scale
- Throughput: 10,000+ writes/second per node; 100,000+ reads/second per node
- Latency: Sub-10ms for single-row reads, with proper row key design
- Scale: Petabytes of data across thousands of nodes
- No index maintenance cost: Adding columns doesn't slow writes
Bigtable Limitations
- No SQL: Uses an HBase-compatible API or the Bigtable client libraries. No JOIN, no WHERE clause, no aggregations at the database level.
- No cross-row transactions: Can't atomically update multiple rows
- Minimum cost: Even the smallest Bigtable cluster (1 SSD node) costs ~$500/month. Not suitable for small workloads.
- Learning curve: Row key design requires careful thought and load testing
When to Choose Bigtable
- Time-series data at scale: IoT sensor readings, application metrics, financial tick data
- AdTech: user event streams, impression logs
- Machine learning feature storage (feature store patterns)
- Any workload needing >10,000 writes/second sustained
- Data that would be queried as "give me the last N values for key X"
Head-to-Head Comparison
| Dimension | Firestore | Bigtable | Datastore (legacy) |
|---|---|---|---|
| Data model | Documents/collections | Wide-column | Entity/kind |
| Query language | Limited SQL-like | Row key/prefix scans | GQL (limited) |
| Secondary indexes | Yes (automated) | No | Yes (limited) |
| Real-time sync | Yes | No | No |
| Transactions | Multi-document | Single-row only | Yes (limited) |
| Minimum cost | ~$0 (serverless) | ~$500/month | ~$0 (serverless) |
| Write throughput | ~10K/second | 100K+/second per node | ~10K/second |
| Time-series native | No | Yes | No |
| Operational overhead | None | Low (managed) | None |
| BigQuery integration | Export available | Direct connector | Export available |
Integration with BigQuery
Both Firestore and Bigtable integrate with BigQuery for analytical queries:
# Firestore → BigQuery export (for analytical queries)
gcloud firestore export gs://my-backup-bucket/firestore-export --collection-ids=users,orders
# Load into BigQuery
bq load --source_format=DATASTORE_BACKUP my_dataset.orders gs://my-backup-bucket/firestore-export/all_namespaces/kind_orders/
# Bigtable → BigQuery via direct connector (no export needed)
bq query --use_legacy_sql=false '
SELECT *
FROM EXTERNAL_QUERY(
"my-project.europe-west4.my-bigtable-connection",
"SELECT * FROM sensor-data WHERE _key LIKE "sensor:abc123%""
)'
For the relational database options on GCP, see our Cloud SQL vs AlloyDB vs Spanner guide. For real-time processing of Bigtable data at scale, see our Cloud Dataflow guide.