Google Cloud Spanner is now half the cost of Amazon DynamoDB

vicpara · on Oct 11, 2023

Just moved our infra from GCP to AWS. Kubernetes clusters, LB, storage, lambdas, KMS and all of it.

Google runs their tech stack as if it's a startup that builds their CV. Everything is immature, tons of hacks, undocumented features. If you are on their k8s there are tons of upcoming new versions and features that force you to revisit key hacks you put in your infra because of their misgivings. Our infra team keeps tinkering around our infra and it never ends. It's 50:50. 50% of time making sure we are prepared for their shit and 50 % our ambitious infra plans. Good luck with that.

With AWS our bill is 60% of what GCP used to be running 3 k8s clusters.

AWS support is so nice, you can't believe it.

Nah, I don't trust Google with anything. It's a scam. Google's support is horrendous. They refer you to idiots that drag you through calls until your will for life dies. And you're back to the mercy of some lost engineer that may comment on a github issue you opened 20 days ago. We have a bug reported back in 2020 that got closed recently without any action because it became stale and the API changed so much it doesn't really matter. It's that bad.

The billing day is a monthly reminder you're paying entitled devs to do subpar work other companies do a lot better.

No, we don't miss them already.

dijit · on Oct 11, 2023

Interesting, if you swap GCP and AWS in your post then thats exactly my experience.

I wonder what makes us different, I work in europe on video games; AWS’s handling of me when I was at Ubisoft left a really sour taste - when I moved into Tencent/Sharkmob I tried really hard to love AWS as it was the defacto industry standard and instead I was left with a feeling that most of it is inconsistent garbage papered over with lambda functions. I referred to these weird gotchas as “3am topics”; things that I don't have the mental capacity to deal with at 3am and convinced the studio to switch to GCP- which, incidentally they are still extremely grateful to me for doing.

cjaybo · on Oct 11, 2023

> I was left with a feeling that most of it is inconsistent garbage papered over with lambda functions

This sounds more like an indictment of the system design than the cloud provider.

What are some of these “3am” topics that made GCP a better choice?

dijit · on Oct 11, 2023

Small examples included (I’m on my phone so these are from memory and you’ll have to forgive the lack of great detail):

1) having the project/account your in visible at the top at all times.

We used SSO for “accounts” which is AWS’s way of completely separating resources; the long string that is returned is not unique in the start and the remainder is cut off: so all accounts/projects looked the same, was impossible to tell at a glance if you were in dev, staging or prod.

2) Autoscaling groups with that had human readable incrementing “names”, in AWS instances have hex slugs as instance names and you can give an instance a special “Name” label: but any new machines created with an ASG will just reuse the same name label making them hard or impossible to tell apart.

The AWS official solution for this is to have a lambda function hook on the scale event and give your new node an incremented name label. Given that AWS is pricy to save me time: I do not personally consider this an elegant solution.

3) having all regions on one page.

We spent €6,000~ on a database we didn't know about until we started digging into the bill. Not knowing what resources are available at a glance feels pretty basic to me tbh.

4) the network implementation overall; in Google you can just make a network and it will work without having to mess with zone routing and configuration of that which is put on the user.

If it’s on the user, it’s a variable that has to be checked during an outage; it is terraform code that has to be grokked and so-on.

master_crab · on Oct 11, 2023

“2) Autoscaling groups with that had human readable incrementing “names”, in AWS instances have hex slugs as instance names and you can give an instance a special “Name” label: but any new machines created with an ASG will just reuse the same name label making them hard or impossible to tell apart. The AWS official solution for this is to have a lambda function hook on the scale event and give your new node an incremented name label. Given that AWS is pricy to save me time: I do not personally consider this an elegant solution”

Why were you even messing with the instance name? This is a ridiculously simple problem to solve with tags on your ASG. And AWS even did the courtesy of propagating those tags across the ASG and all its instances.

https://docs.aws.amazon.com/autoscaling/ec2/userguide/ec2-au...

justin_oaks · on Oct 12, 2023

> 1) having the project/account your in visible at the top at all times.

I agree that this is an annoying issue in the AWS web console.

I assume this is something that could be fixed on your end by a little bit of CSS.

thedougd · on Oct 12, 2023

I believe the solution is to give the account an alias.

https://docs.aws.amazon.com/IAM/latest/UserGuide/console_acc...

robingchan · on Oct 12, 2023

company im working at currently uses Token Vending Machine.

pros: cannot get accounts mixed up.

cons: All sessions are actually 12hr sessions (ASIA not AKIA) and no access to perm keys for cli, security i suppose. Its not too bad though as TVM gives creds for various use cases.

https://aws.amazon.com/blogs/apn/tag/token-vending-machine/

grogenaut · on Oct 12, 2023

we fix that internally by having names for accounts and having stages for accounts in a meta tool. There's a tampermonkey script that pulls that in and shows it on screen and a red banner if it's prod. Could be a json file in a github repo. And yes it could be a console feature but everyone's got different concepts of prod. I think a ton of companies use like 2 total accounts as well.

sparc24 · on Oct 11, 2023

It's amazing how people complain about GCP. We run a massive deployment across 100+ regions cross-cloud GCP, Azure, AWS and oh boy. GCP has good support if you are big enough. Azure though which has a much bigger share than GCP is horrendous. Absolutely garbage all around. Good luck ever getting anyone in Engineering even if you are paying for support. AWS on the other hand - Amazing. We have Ent Support so those guys in our slack channel. The TAMs are amazing. Need to get hold of someone in Route53 no problem they are on the call this week. Feature request for EKS - ok talk to the Product Manager this afternoon.

Azure is a dumpster fire from the ground up.

CommanderData · on Oct 11, 2023

Can you give Azure specifics, as you know Azure has a massive offering.

My experience has been the opposite though not without issues, Azure has some of the best corporate and security features of any cloud and it's only getting better. The zero trust model fits in so nicely with their identity platforms it's a sight to behold compared to other cloud providers which likely use some form of AAD or AD DS anyway.

Their support is responsive and they seem to know what they're talking about. (AKS)

Please provide some specifics on your experience?

numbsafari · on Oct 12, 2023

Azure has had repeated, significant security failures that impact numerous customers. I don't understand how anyone can defend their security except through willful ignorance.

I have friends forced to use Azure and they routinely report issues with provisioning resources, things taking a very long time to spin up or simply being rejected because Azure doesn't have any capacity.

spiffytech · on Oct 12, 2023

A memorable example is when we ran a heavy Azure Functions workload on our App Service Plan, the hosts would devour themselves.

Functions use containers under the hood. Each invocation created a new container, and when enough of them ran long enough, the host disk would fill up. (Pretty sure our workload wrote almost nothing to disk.)

An internal Azure disk clean-up routine kicked in, which deleted image layers for running Functions. This deleted the filesystems for containers that were still running, yanking them out from underneath the running processes. It also meant the host couldn't launch new instances of our Functions.

At this point the host was poisoned and couldn't launch any new work, even after the workload was reduced, It had to be terminated and replaced, after we detected the problem manually.

Azure support never weemed to take the problem seriously, and after we migrated our workload off of Functions they decided the problem must be resolved since we weren't complaining anymore.

g9yuayon · on Oct 11, 2023

> AWS support is so nice, you can't believe it.

This reminds me of the fond days of having weekly customers calls. We develop AWS services, and we answer our customer-support calls directly. No middle man. Just techies to techies. And we made promises to customers on the fly, and customers sometimes project managed us.

kubb · on Oct 11, 2023

Sounds like hell...

oblio · on Oct 11, 2023

The last part is a over the top but direct access to customers as a dev is a plus, not a minus.

User23 · on Oct 12, 2023

Customers aren’t that bad, really.

falcolas · on Oct 11, 2023

We have an old "quiet part out loud" corporate story. It's about how one arm of Google using our service and wondering why it had so much downtime, only for us to point at their GAE arm and say "when they're down, we're down". They went and talked to GAE and - funny enough - were able to correlate the downtime they observed with GAE downtime.

GAE uptime improved, for a little while. Yeah, we're on AWS now too.

nerpderp82 · on Oct 11, 2023

Does Google run on GCP?

lcw · on Oct 11, 2023

From my understanding they don't dogfood a lot of gcp products internally. That's how you end up with janky integrations between their products. It's really frustrating at times to see their cloud architects pitch some grouping of technologies that you should use to find out the integrations aren't well tested at scale. For example, pushing for pubsub to be used with dataflow for near real time processing just to figure out at scale global pubsub has high latency, above 1 minute sometimes 5 minutes, on 1% of messages at scale.

danans · on Oct 12, 2023

Yes in the sense that they use all the services and infrastructure that GCP is built in, but no in the sense of using the vanilla GCP interface.

Instead many aspects of GCP's management console are handled by different internal tools, often command line driven. IME they are often far more unwieldy than GCP.

Sometimes this makes sense (far tighter access controls and configuration change controls than a typical company), and some times it's just because of legacy ways of doing things.

I worked on a team at Google that used the internal GCP to serve some code/content for a specific feature, and it was in some ways it was more frustrating than using just either the normal internal systems or just vanilla GCP.

GabeWeiss_ · on Oct 11, 2023

Parts, yes. In reference to the specifics mentioned in here though, those services run on Infra Spanner, not Cloud Spanner, but they're the same stack. The main reason things like Gmail, Ads, etc haven't swapped into GCP is because of the internal tooling that's built up around the infra spanner relating to those services specific to Google that don't make sense in Cloud Spanner.

QuercusMax · on Oct 11, 2023

It's way WAY more than just Infra Spanner vs Cloud Spanner. Cloud spanner doesn't support protobuf, which is annoying, but that's not a dealbreaker; it's still just a DB. The issue is really all the various internal frameworks (such as Apps Framework for Java), deployment systems (Server Platform, AKA Boq/Pod/Urfin), and so forth.

GabeWeiss_ · on Oct 11, 2023

Of course, I was simplifying. It's always more complicated doing a migration. :)

QuercusMax · on Oct 11, 2023

Not just migrations are hard, either; Google Cloud has put (almost?) zero effort into making it easy to use Cloud from systems running on Borg.

My old team was building a system that was half-GCP and half-Borg, and we had to write our own (extremely bad) Cloud Spanner fake for use in tests. In contrast, Infra Spanner is extremely well supported for tests. Same with BigQuery vs Dremel and many other systems.

bananapub · on Oct 12, 2023

this is maybe the most ill-informed post about google I've ever seen on this site, wow

jezzamon · on Oct 11, 2023

Short answer: No

User23 · on Oct 12, 2023

Borg still? Supposing you can say. Don’t reply if it’s still borg and you can’t.

hot_gril · on Oct 11, 2023

Mostly not.

bananapub · on Oct 12, 2023

pirsquare · on Oct 11, 2023

> AWS support is so nice, you can't believe it.

This! They even custom-coded their support portal better than those off-the-shelf vendor like Zendesk. I say this as a Zendesk paying customer.

GCP on the other hand, is a F-tier in support. Almost feel like I need to beg them to get any level of help.

cameronh90 · on Oct 11, 2023

At one point, Google reached out to me to try and tempt us over from AWS. I had bad experiences with Google support in the past, but liked their AI stuff and was keen to give them another go.

We booked a follow up call in the calendar, I spent good time preparing my notes and requirements for the meeting... and then nobody on their side showed up or contacted me again.

mrtksn · on Oct 11, 2023

A GCP issue was the only time I had a human contact with Google, they did well. However high scale low touch is in their DNA and you can tell it.

Cthulhu_ · on Oct 11, 2023

Wouldn't Zendesk be one of those software things that has had too many features bolted on without oversight and/or a unified philosophy behind it?

I'm of the opinion that focused products created by smaller teams are better.

wiremine · on Oct 11, 2023

Much of the time GCP feels like a science project, and not a real business. AWS (and Azure) seem to be driven by customer requests, instead of Google, which feels very engineering-centric.

Which is on brand with Google. They have no problem launching stuff, and no problem killing stuff. But man, then just get out of the cloud business and focus on what you're good at.

doesnt_know · on Oct 11, 2023

> AWS support is so nice, you can't believe it.

It's actually sort of ridiculous. AWS has the best support I have ever interacted with. I mean, our org certainly pays enough for it but it's so completely unusual in tech, or really any sector to get great support even when you're paying for it.

stouset · on Oct 11, 2023

Every time I run into an issue I’m reluctant to reach out to AWS support, because of my default expectation that it will be a terrible waste of time.

Every time, I am also proven wrong as someone competent on their side both actually understands my issue and finds a resolution.

traceroute66 · on Oct 11, 2023

> We have a bug reported back in 2020 that got closed recently without any action because it became stale

One of my pet hates is the (ab)use by repo maintainers of the auto-close-when-stale feature on Github.

What useful purpose does it serve beyond making the repo maintainers look good because they have a low number of open issues ?

It doesn't actually address the issue. Its the virtual equivalent of brushing under the carpet.

softveda · on Oct 11, 2023

I worked in a Digital team 4 years back where the team was building voice channel apps for our customers on both Amazon Alexa and Google Dialogflow. Alexa NLP engine was less sophisticated we had to give it hundreds of prompts and intents. Dialogflow NLP engine required a handful of prompts for the same thing. But when it came to integration with backend APIs and support Alexa was far ahead. Despite having Dialogflow enterprise Google support would suggest to ask in StackOverflow. Amazon support on the other hand was excellent. We needed support for mTLS with the backend APIs, Amazon supported it as they understood enterprise. Google just shooed us away, their support wouldn’t even escalate this.

mardifoufs · on Oct 11, 2023

I don't know. I like GCP. I have been in an Azure centric corporation for close to two years now and I dearly miss GCP almost every day.

My team has a sort of a sandbox where we can use almost any Azure product we want (our IT is supportive and permissive as far as that sandbox goes, which is a blessing), but even then it's just painful in comparison.

AWS is probably better though.

hot_gril · on Oct 11, 2023

Azure seems like a nightmare, regardless of whether GCP or AWS is better.

catch22p · on Oct 12, 2023

There is no way this is true. Only explanation is you work for AWS :-). GCP strength is it's cost. Yes may be the support could be better. But can you care to explain what "hacks" are you talking about ? And the claim that K8S(from Google) is better on AWS than GCP is absolutely false

vicpara · on Oct 13, 2023

We are a small startup, 12 strong.

Our reason for going all in with GCP was the k8s. We've been using GCP for 2+ years. The trouble we have is with stability and so many of the features being constantly rolled out.

Our experience was that K8s cost more on GCP than AWS.

Just on LoadBalancers alone, you have tons of tricks that are specific to GCP implementation. And we needed a few extra because you couldn't run all the features we wanted on 1-2 per cluster. For example, we have a 3rd party that required all our requests to always originate and respond back from a fixed IP address. We could only pick one not a range, not a list. This was a hard requirement. The service was important so we had to do it.

It took our team several days to find how to do it using online documentation and support. Tech support was useless. We had one guy in our team that spent 2 days on the phone with a paid, local GCP implementation partner trying to get this problem sorted. Nothing came out of it other than being pitched on our dime a lot of services and architecture we didn't need. Eventually we figure it out on our own. I don't even remember speaking about this when we transitioned to AWS.

YetAnotherNick · on Oct 11, 2023

Matches my experience. GCP has many better services than AWS but I am not going to run production workload with them after 2 years of experience in previous company. There are so many undocumented quirks that many times you could find better solution from some random person in stackoverflow than highest tier paid support.

acdha · on Oct 12, 2023

That was my experience, too - a couple of things which were better than AWS but this constant stream of paper cuts hitting all of the problems which weren’t cool enough to get someone promoted.

pojzon · on Oct 11, 2023

Doesnt help when GKEngine has constant issues and recent upgrade feature is on 10 days strike.

Recently was woken up by alert about DNS resolution issues.

GCP rolled out new version of SkyDNS and NodeLocalDNS, SkyDNS reports 99% miss, had to quickly hack it.

This is not the „out-of-the-box” experience you want to have.

vismwasm · on Oct 11, 2023

I generally like GCP, however their sales and customer support just aren't any good. And some services like Vertex AI are extremely buggy while it's hard to actually report these bugs.

I think Google Cloud needs someone like Jeff Bezos as their head: Look what your customers actually want and need and understand their requirements. And they usually want good customer support and want a competent key account manager as well.

When we were looking to migrate our analytics database from on-premise to a cloud alternative we were looking at BigQuery and Snowflake. BigQuery is a great product and we were already deeply invested in GCP as well. However the GCP sales team just couldn't sell BigQuery - they just don't know what old corporations want to hear in a sales pitch. So we went with Snowflake in the end. Not because it's the better product but because their sales team is better.

I'm not sure if the cloud business is actually a priority at Google. If it is then I think they don't understand the mistrust Google is facing when it comes to stable long term support of their products.

insanitybit · on Oct 11, 2023

The horror stories of Google support, across all of their products, is enough for me to never trust GCP. Even if someone told me today "GCP is the exception, they have great support" I probably wouldn't care - they are so organizationally incapable of providing good support that, even if they did so today, I wouldn't believe that it could last.

DevKoala · on Oct 11, 2023

My experience too, GCP is frustrating. However, there is nothing like BigQuery to me, I love that DB.

ExoticPearTree · on Oct 11, 2023

Support wise, GCP is a joke run by entitled people. I had an issue some time ago with a VPN and after doing a lot of troubleshooting and having them agree the problem is on their end (packets would go in their VPN Gateway from the VPC, nothing would come out), the solution was to update my configuration on my end to workaround whatever they did because "it is how is going to be"...

TL;DR: they broke something and wouldn't fix it.

sabellito · on Oct 11, 2023

Doing business with Google is a liability.

pyuser583 · on Oct 12, 2023

Did you consider your own bare-metal?

gregdoesit · on Oct 11, 2023

“ According to the Amazon Prime Day blog post, DynamoDB processes 126 million queries per second at peak. Spanner on the other hand processes 3 billion queries per second at peak, which is more than 20x higher, and has more than 12 exabytes of data under management.”

This comparison seems to be not exactly fair? Amazon’s 126 million queries per second was purely for Amazon-related services serving Prime Day generating this on DynamoDB, and not all of AWS is my read.

What would have perhaps been a more fair comparison is to share the peak load that Google services running Cloud Spanner, and not the sum of all Spanner services across all of GCP and all of Google (Spanner on non-GCP infra).

I will say that it would show a massive of confidence to say that Photos, Gmail and Ads heavily rely on GCP infra: which would be brand new information for me! It would add to confidence to learn more on how they use it, and if Cloud Spanner is on the critical path for those services.

What is confusing, however, is how in this article "Cloud Spanner" is consistently used... except for when talking about Gmail, Ads and Photos, where it's stated that "Spanner" is used by these products, not "Cloud Spanner!". Like if they were not using the Cloud Spanner infra, but their own. It would help to know what is the case, and what the load of Cloud Spanner is: and not Spanner running on internal Google infra that is not GCP.

At Amazon, practically every service is built on top of AWS - a proper vote of confidence! - and my impression was that GCP had historically been far less utilised by Google for their own services. Even in this post, I'm still confused and unable to tell if those Google products listed use Cloud Spanner or their own infra running Spanner.

tedivm · on Oct 11, 2023

From the AWS blog post they referenced-

> DynamoDB powers multiple high-traffic Amazon properties and systems including Alexa, the Amazon.com sites, and all Amazon fulfillment centers. Over the course of Prime Day, these sources made trillions of calls to the DynamoDB API. DynamoDB maintained high availability while delivering single-digit millisecond responses and peaking at 126 million requests per second.

Amazon was very, very clear on this. For Google to use that number without the caveat is just completely underhanded and dishonest. Whoever wrote this is absolutely lacking in integrity.

ljm · on Oct 11, 2023

I used DynamoDB as part of the job a few years ago and never got single-millisecond responses - it was 20ms minimum and 70+ on a cold-start, but I can accept that optimising Dynamo's various indexes is a largely opaque process. We had to add on hacks like setting the request timeout to 5ms and keeping the cluster warm by submitting a no-op query every 500ms to keep it even remotely stable. We couldn't even use DAX because the Ruby client didn't support it. At the start we only had a couple of thousand rows in the table so it would have legit been faster to scan the entire table and do the rest in memory. Postgres did it in 5ms.

If Amazon said they didn't use DAX that day I would say they were lying.

The average consumer or startup is not going to squeeze out the performance of Dynamo that AWS is claiming that they have achieved.

In fact, it might have been fairer in Ruby if they didn't hard-code the net client (Net/HTTP). I imagine performance could have been boosted by injecting an alternative.

iot_devs · on Oct 11, 2023

No need to guess when you can measure.

I am running https://cloud-canary.com a service where I monitor AWS primary services for latency and availability.

It comes with a lot of data.

For instance this is the latency I see doing operations against Dynamo.

https://cloudcanary.grafana.net/public-dashboards/c53e2092d6...

loxias · on Oct 11, 2023

What a cool lil side project/company! Going to circulate this among friends...

Little bit of well meaning advice: This needs copy editing -- inconsistent use of periods, typos, grammar. Little crap that doesn't matter in the big picture, but will block some from opening their wallets. :) ("OpenTeletry", "performances", etc.)

All in all this is quite cool, and I hope you get some customers and gather more data! (a 4k object size in S3 doesn't make sense to measure, but 1MB might be interesting. Also, check out HDRHistogram, it might be relevant to your interests)

iot_devs · on Oct 11, 2023

Thanks!

Any feedback is appreciated!

I pick 4k as a no-op against S3, something that very little time but still does some work.

I will definitely consider to increase it!

bennyg · on Oct 11, 2023

Nice dash - if you don't mind a drive-by recommendation: I use Grafana for work a lot and it's nice to see a table legend with min, max, mean, and last metrics for these kinds of dashboards. Really makes it easy to grok without hovering over data points and guessing.

RedlineTriad · on Oct 12, 2023

What is more important for me when using Grafana (though a summary is as well) is actually units, to know if it's second, millisecond, microsecond, and also if 0.5 is a quantile or what.

Numbers without units are dangerous in my opinion.

iot_devs · on Oct 11, 2023

Thanks a lot!

I'll definitely update it!

qwertox · on Oct 11, 2023

What a cool service. Congratulations!

RhodesianHunter · on Oct 11, 2023

> We had to add on hacks like setting the request timeout to 5ms and keeping the cluster warm by submitting a no-op query every 500ms to keep it even remotely stable.

This sounds like you're blaming dynamo for you/your stack's inability to handle connections / connection pooling.

tedivm · on Oct 11, 2023

Yeah that TLS handshake is an absolute killer if you run it for every request.

iends · on Oct 11, 2023

Been using DynamoDB for years and haven’t had to do any of the hacks you talk about doing. Not using ruby though. TCP keep-alive does help with perf though (which I think you might be suggesting.)

I don’t have p99 times in front of me right this second but it’s definitely lower than 20ms for reads and likely lower for writes. (EC2 in VPC).

mk89 · on Oct 11, 2023

They very well know that people don't read sh* anymore. Just throw numbers there, PowerPoint them and offer an "unbiased" comparison where Google shines - buy Google.

Worst case scenario, it's Google you're buying, not a random startup etc.

azmodeus · on Oct 11, 2023

Google doesn't have a great brand of not killing products. No support and randomly killing stuff is not a good business relationship

GabeWeiss_ · on Oct 11, 2023

Just as a hand in the air...Be careful about what you're comparing here. # of API calls over a period of time is...largely irrelevant in the face of QPS. I can happily write a DDOS script that massively bombards a service, but if that halts my QPS then it doesn't matter. So sure, trillions of API calls were made (still impressive in the scope of the overall network of services, I'm not downplaying that), but ultimately, for DynamoDB and Spanner, it's the QPS that mattered to us in terms of comparisons of DB scaling and performance.

vineyardmike · on Oct 11, 2023

Google calls API calls “queries”… because of their history as a search engine. QPS == API calls/per second == Requests per second

That said, I can’t imagine these numbers mean much to anyone after a certain point. It’s not like either company is running a single service handling them. The scale is limited by their budget and access to servers because my traffic shouldn’t impact yours. I feel like the better number is RPS/QPS per table or per logical database or whatever.

GabeWeiss_ · on Oct 11, 2023

Yes, but QPS vs. "queries to the API". The difference is the time slice. I should have been more explicit. The key here really is the time function between the numbers. That the AWS blog calls out trillions of API calls isn't relevant because there wasn't a specific time denominator. The 126M QPS is the important stat.

forrestbrazeal · on Oct 11, 2023

We shared some details about Gmail's migration to Spanner in this year's developer keynote at Google Cloud Next [0] - to my knowledge, the first time that story has been publicly talked about.

[0] https://www.youtube.com/watch?v=268jdNwH6AM

gregdoesit · on Oct 11, 2023

I tried to find it in this video, but failed. Could you please share a time stamp on where to look?

It’s a pretty big deal if Gmail migrated to GCP-provided Spanner(not to an internal Spanner instance) and sounds like he kind of vote of confidence GCP and Cloud Spanner could benefit from: might I suggest to write about it? It’s easier to digest and harder to miss than an hour-long keynote video with no time stamps.

And so just to confirm: Gmail is on Cloud Spanner for the backend?

dekhn · on Oct 11, 2023

It's almost certainly not the case that Gmail uses Cloud Spanner rather than Internal Spanner. I don't think Cloud Spanner (or most of Google's cloud products) have the featureset required to support loads like Gmail (both in terms of technical capability, and security/privacy features).

When I worked at Google I tried to get more services to migrate to the cloud but the internal environment that was built up over 25 years is much better at supporting billion+ users with private data.

Cthulhu_ · on Oct 11, 2023

And yet, if they do, that's probably one of the best sales pitches they could have - dogfooding. After all, isn't that also how AWS started, just reselling the services and servers they already use themselves?

It doesn't make much sense to have a 'better' version of a product you sell but keep it internal.

xdeepak81 · on Oct 14, 2023

Yet Amazon Retail still don't use DynamoDb for the critical workloads. They still rely on an internal version of DynamoDb (Sable) which is optimized for Retail workload.

jeffbee · on Oct 11, 2023

It makes sense because the public will not use the internal APIs which have non-standard wire protocols, weird authentication schemes, etc.

alphabetting · on Oct 11, 2023

looks like it starts at 50:45. youtube recently made it so you can click "show transcript" in the description then ctrl-f takes you to all the mentions. very helpful for long videos like this.

easton · on Oct 11, 2023

It looks like the Spanner beta dropped to the public in 2017, so < 8 years ago: https://cloud.google.com/spanner/docs/release-notes#February...

I don't think they would've migrated again to GCP Spanner (even if it would've been a show of faith).

dabernathy89 · on Oct 11, 2023

Here's the link with timestamp (note that the speaker says it was a 2 year transition):

https://www.youtube.com/live/268jdNwH6AM?si=WkgnvqaIwFidt-hc...

tazjin · on Oct 11, 2023

Gmail is on Spanner, and Cloud Spanner is on Spanner.

vxNsr · on Oct 11, 2023

In the timestamped video link shared downthread, the speaker does seem to strongly imply that gWorkspace doesn’t manage the infra, when he finishes explaining the migration he declares (around 55:18)“[…]we can focus on the business of gmail and spanner can choose to improve and deliver performance gains automagically[sic]” which would imply, to me at least, that it’s on GCP.

jeffbee · on Oct 11, 2023

That's not what it implied to me. To me, it meant that they adopted an internal managed Spanner with its own SRE team, instead of running their own Spanner. In the past, Gmail ran their own [[redacted]]s and [[redacted]] even though there were company-wide managed services for those things.

eep_social · on Oct 11, 2023

Agree, but with the caveat that [[redacted]] and [[redacted]] were old and originally designed to be run that way. All newer storage systems I can recall were designed to be run by a central team after many years of experience doing it the other way. And many tears shed over migrating to those centralized versions.

Source: I was on the last team running our own [[redacted]].

Vt71fcAqt7 · on Oct 11, 2023

Thanks! Looks really interesting.

link with time-stamp:

https://www.youtube.com/watch?v=268jdNwH6AM?&t=3020

jeffbee · on Oct 11, 2023

Wow, almost content-free presentation! How obnoxious!

This wasn't the first time Gmail has replaced the storage backend in-flight. The last time, around 2011, they didn't hype it up, they called it "a storage software update" in public comms. And that other migration is the origin of the term "spannacle", because during that migration the accounts that resisted moving from [[redacted]] to [[redacted]] we called barnacles.

vxNsr · on Oct 11, 2023

Somehow I thought you were at Amazon/AWS because of how much you push it in your book. Cool to see you’re at GCP.

dmoy · on Oct 11, 2023

> I will say that it does show a vote of confidence to say that Photos, Gmail and Ads use GCP infra,

I'm not sure? I guess I'm mostly not sure what "gcp infra" means there. The blog post says

"Spanner is used ubiquitously inside of Google, supporting services such as; Ads, Gmail and Photos."

But there's google-internal spanner, and gcp spanner. A service using spanner at Google isn't necessarily using gcp. (No clue about photos, Gmail, etc)

Granted, from what I gather, there's a lot more similarity between spanner & gcp spanner than e.g. borg and kubernetes.

bananapub · on Oct 11, 2023

borg and k8s are completely unrelated bits of software with roughly similar goals.

gcp spanner and normal spanner are different deployments of the same code.

marcinzm · on Oct 11, 2023

>different deployments

Which can be the difference between 99.99% availability and 99% availability with data corruption issues. Not saying that's the case here but one should not downplay the difference deployments can make.

gregdoesit · on Oct 11, 2023

Surely in a post about Google Cloud Spanner, all examples mentioned use Google Cloud Spanner? It would be moot listing them as examples if they would not: so my assumption is they are all using GCP infra already for Spanner.

I really want to give Google the benefit of the doubt: but it doesn't help that they did not write that eg Gmail is using "Cloud Spanner." They wrote that it uses Spanner.

ericpauley · on Oct 11, 2023

This is putting a lot of faith in GCP advertising. I strongly doubt the idea that the Google workloads discussed are deployed on GCP instead of internal Borg infrastructure.

kccqzy · on Oct 11, 2023

Years ago they did a reorg and moved all infrastructure services under Cloud even though they are not Cloud products. That would enable this kind of obfuscation because Cloud is literally responsible for both Cloud Spanner and non-Cloud Spanner and they can conflate these two in their marketing copy. They probably feel justified in doing so because they share so much code.

0xbadcafebee · on Oct 11, 2023

Considering that most of Google does not run on GCP, I would not give them the benefit of the doubt.

blueg3 · on Oct 12, 2023

Photos, Gmail, and Ads use Spanner, not Cloud Spanner.

Apparently Cloud Spanner doesn't support protobuf columns? It would be hard for any internal Google product to use it under that restriction.

GabeWeiss_ · on Oct 11, 2023

Infra and Cloud Spanner are the same stack. Having those services run on infra is more about the legacy of tooling to shift it rather than anything around performance or ability to handle it

GabeWeiss_ · on Oct 11, 2023

Infra and Cloud Spanner are the same stack. Having those services run on infra is more about the legacy of tooling to shift it rather than anything around performance or ability to handle it.

tw04 · on Oct 11, 2023

>This comparison seems to be not exactly fair? Amazon’s 126 million queries per second was purely for Amazon-related services serving Prime Day generating this on DynamoDB, and not all of AWS is my read.

There's no indication that google is talking about ALL of spanner either? The examples they list are all internal google services, and they specifically say "inside google".

I'm also dubious that even with all of the AWS usage accounted for that DynamoDB tops Spanner if Amazon themselves are only at 126 million queries per second on Prime Day.

g9yuayon · on Oct 11, 2023

> At Amazon, practically every service is built on top of AWS - a proper vote of confidence!

Not only this, but practically most, if not all, of the AWS services use DynamoDB, including use cases that are usually not for databases, such as multi-tenant job queues (just search "Database as a Queue" to get the sentiment). In fact, it is really really hard to use any relational DB in AWS. I mean, a team would have to go through a CEO approval to get exceptions, which says a lot about the robustness of DDB.

tacozilla · on Oct 11, 2023

Eh, this isn't accurate. Both Redshift and Aurora/RDS are used heavily by a lot of teams internally. If you're talking specifically about the primary data store for live applications, NoSQL was definitely recommended/pushed much harder than SQL, but it by no means required CEO approval to not use DDB

Edit: It's possible you're limiting your statement specifically to AWS teams, which would make it more accurate, but I read the use of "Amazon" in the quote you were replying to as including things like retail as well, etc.

g9yuayon · on Oct 11, 2023

Yeah, within AWS. I'm not sure about other parts of Amazon

sharpy · on Oct 11, 2023

When I was at AWS, towards later part of my tenure, DynamoDB was mandated for control plane. To be fair, it worked, and worked well, but there were times when I wished I could use something else instead.

brunoborges · on Oct 11, 2023

> What would have perhaps been a more fair comparison is to share the peak load that Google services running on GCP generated on Spanner, and not the sum of their cloud platform.

Not necessarily about volume of transactions, but this is similar to one of my pet-peeves with statements that use aggregated numbers of compute power.

"Our system has great performance, dealing 5 billion requests per second" means nothing if you don't break down how many RPS per instance of compute unit (e.g. CPU).

Scales of performance are relative, and on a distributed architecture, most systems can scale just by throwing more compute power.

hangonhn · on Oct 11, 2023

Yeah I've seen some pretty sneaky candidates try that on their resumes. They aggregate the RPS for all the instances of their services even though they don't share any dependencies nor infrastructure. They're just independent instances/clusters running the same code. When I dug into those impressive numbers and asked about how they managed coordination/consensus the truth comes out.

GabeWeiss_ · on Oct 11, 2023

True, but one would hope that both sides in this case would be putting their best foot forward. Getting peak performance out of right sizing your DB is part of that discussion. I can't imagine AWS would put down "126 million QPS" if they COULD have provided a larger instance that could deliver "200 million QPS", right? We have to assume at some point that both sides are putting their best foot forward given the service.

redditor98654 · on Oct 12, 2023

The 126M QPS number was certainly parts of Amazon.com retail that powers Prime Day not all of DDB traffic. If we were to add up all of DDB's volume, it would be way higher. At least a magnitude if not more.

Large parts of AWS itself uses DDB - both control plane and data plane. For instance, every message sent to AWS IoT will internally translate to multiple calls to DDB (reads and writes) as the message flows through the different parts of the system. IoT itself is millions of RPS and that is just one small-ish AWS service.

Source: Worked at AWS for 12 years.

BoorishBears · on Oct 11, 2023

Put yourself in the shoes of who they're targeting with that.

Probably dealing with thousands of requests per seconds, but wants to say they're building something that can scale to billions of requests per second to justify their choices, so there they go.

Rapzid · on Oct 11, 2023

Frankly it's a bit weird to see this kind of dick measuring in a product blog post from the "Director of Engineering" :/

cbarrick · on Oct 11, 2023

s/the "Director of Engineering"/a "Director of Engineering"/

There are many engineering directors at Google.

blueg3 · on Oct 12, 2023

Director is what, L8? There's a ton of those.

Rapzid · on Oct 11, 2023

And only one attributed to the blog post.

swish

ripper1138 · on Oct 11, 2023

True and even worse, inaccurate dick measuring.

jjtheblunt · on Oct 11, 2023

> At Amazon, practically every service is built on top of AWS

is that true finally? It sure wasn't in the 2020-2021 timeframe.

dastbe · on Oct 11, 2023

it does depend on what you mean. By 2020/2021, effectively everything was on top of AWS VMs/VPC and perhaps LBs at that point? Most if not all new services were being built in NAWS.

jjtheblunt · on Oct 11, 2023

SPS was heavily MAWS and I got sick of being the NAWS person from years prior pushing for NAWS in our dysfunctional team, and quit. The good coworkers also quit.

Yet I still see the very deep stack of technically incapable middle manager sorts dutifully posting "come join us" nonsense on LinkedIn.

(I had the luxury of having worked in one of the inner sanctums of Apple hardware for years prior, so was immune to nonsense, and didn't need the job.)

nameless912 · on Oct 11, 2023

And for many projects, Postgres is still cheaper than both. Having used both, I would much, much rather do the work to fit my project in Postgres/CockroachDB than use either Spanner or DynamoDB, which have WAY more footguns. Not to mention sudden cost spikes, vendor lock in, and god knows what else.

AWS and GCP (and Azure, and Oracle cloud, and bare Kubernetes via an operator, and...) support Postgres really well. Just...use Postgres.

bananapub · on Oct 11, 2023

> And for many projects, Postgres is still cheaper than both.

ok? and sqlite3 in memory is even cheaper than postgres!

if you can use (and support correctly) postgres then you should use it, obviously there's no point using a globally scalable P-level database if you can just fit all your data on one machine with posthgres.

airstrike · on Oct 11, 2023

Except for projects for which NoSQL is a better fit than a RDBMS, no?

If I'm writing a chat app with millions of messages and very little in the way of "relationships", should I use Postgres or some flavor of NoSQL? Honest question.

Olreich · on Oct 11, 2023

Postgres. NoSQL databases are specialized databases. They are best-in-class at some things, but generally that specialization came at great cost to their other options. DynamoDB is an amazing key-value store, but is severely limited at everything else. Elasticsearch is an amazing for search and analytics, but is severely limited at everything else. Other specialized databases that are SQL-full are also great at what they do, like Spark is a columnar database that has amazing capabilities for massive datasets where you need lots of cross-joins, but that severely limits it's ability to act in a lot of roles, because they traded latency for throughput and horizontal scalability, and you're restricted in what you can do with it.

The super-power of Postgres is that it supports everything. It's a best-in-class relational database, but it's also a decent key-value store, it's a decent full-text search engine, it's a decent vector database, it's a decent analytics engine. So if there's a chance you want to do something else, Postgres can act as a one-stop-shop and doesn't suck at anything but horizontal scaling. With partitioning improving, you can deal with that pretty well.

If you're writing fresh, there is basically no reason not to use Postgres to start with. It's only when you already know your scale won't work with Postgres that you should reach for a specialized database. And if you think you know because of published wisdom, I'd recommend you set up your own little benchmark, generate the volume of data you want to support, and then query it with Postgres and see if that is fast enough for you. It probably will be.

jpgvm · on Oct 11, 2023

Golden Rule of data: Use PostgreSQL unless you have an extremely good reason not to.

PostgreSQL is extremely good at append-mostly data, i.e like a chat log and has powerful partitioning features that allow you to keep said chat logs for quite some time (with some caveats) while keeping queries fast.

Generally speaking though PostgreSQL has powerful features for pretty much every workload, hence the Golden Rule.

GabeWeiss_ · on Oct 11, 2023

100% this, and even though I work for Google I absolutely agree. BUT, for the folks that need it, PostgreSQL just DOESN'T cut it, so it's why we have databases like DynamoDB, Spanner, etc. Arguing that we should "Just use PG" is kinda a moot point.

nameless912 · on Oct 12, 2023

I think I said this in another comment, but I'm not shitting on Spanner or DDB's right to exist here. Obviously, there are _some_ problems for which a globally distributed ACID compliant SQL-compatible database are useful. However, those problems are few and far between, and many/most of them exist at companies like Google. The fact is your average small to medium size enterprise doesn't need and doesn't benefit from DDB/Spanner, but "enterprise architects" love to push them for some ungodly reason.

pezezin · on Oct 11, 2023

Don't forget PostgreSQL extensions. For something like a chat log, TimescaleDB (https://www.timescale.com/) can be surprisingly efficient. It will handle partitioning for you, with additional features like data reordering, compression, and retention policies.

lost_tourist · on Oct 14, 2023

this is what I've done sqlite3 for my personal stuff, postgres for everything else. I'm far from a "120 million requests per second" level though, so my experience is limited to small to mid-size ops for businesses.

loxias · on Oct 11, 2023

Millions is tiny. Toy even. (I work on what could be called a NoSQL database, unfortunately "NoSQL" is a term without specificity. There's many different ways to be a non-relational database!)

My advise to you is to use Postgresql or, heck, don't over think it, sqlite if it helps you get a MVP done sooner. Do NOT prematurely optimize your architecture. Whatever choice results in you spending less time thinking about this now is the right choice.

In the unlikely event you someday have to deal with billions of messages and scaling problems, a great problem to have, there are people like me who are eager to help in exchange for money.

Lots of people like to throw around the term "big data" just like lots of people incorrectly think that just because google or amazon need XYZ solution that they too need XYZ solution. Lots of people are wrong.

If there exists a motherboard that money can buy, where your entire dataset fits in RAM, it's not "big data".

silisili · on Oct 11, 2023

I've found it's pretty easy to massage data either way, depending on your preference. The one I'm working on now ultimately went from postgres, to mysql, to dynamo, the latter mainly for cost reasons.

You do have to think about how to model the data in each system, but there are very few cases IMO where one is strictly 'better.'

ruuda · on Oct 11, 2023

Postgres; the schema is still structured. But even if you want something less rigid, Postgres has a jsonb type and great operators for querying json.

btown · on Oct 11, 2023

You can also create arbitrary indices on derived functions of your JSONB data, which I think is something that a lot of people don't realize. Postgres is a really, really good NoSQL database.

cooperaustinj · on Oct 11, 2023

Can you expand on this? Documentation or an example so I can learn?

tczMUFlmoNk · on Oct 11, 2023

Sure. Suppose that we have a trivial key-value table mapping integer keys to arbitrary jsonb values:

    example=> CREATE TABLE tab(k int PRIMARY KEY, data jsonb NOT NULL);
    CREATE TABLE

We can fill this with heterogeneous values:

    example=> INSERT INTO tab(k, data) SELECT i, format('{"mod":%s, "v%s":true}', i % 1000, i)::jsonb FROM generate_series(1,10000) q(i);
    INSERT 0 10000
    example=> INSERT INTO tab(k, data) SELECT i, '{"different":"abc"}'::jsonb FROM generate_series(10001,20000) q(i);
    INSERT 0 10000

Now, keys in the range 1–10000 correspond to values with a JSON key "mod". We can create an index on that property of the JSON object:

    example=> CREATE INDEX idx ON tab((data->'mod'));
    CREATE INDEX

Then, we can query over it:

    example=> SELECT k, data FROM tab WHERE data->'mod' = '7';
      k   |           data            
    ------+---------------------------
        7 | {"v7": true, "mod": 7}
     1007 | {"mod": 7, "v1007": true}
     2007 | {"mod": 7, "v2007": true}
     3007 | {"mod": 7, "v3007": true}
     4007 | {"mod": 7, "v4007": true}
     5007 | {"mod": 7, "v5007": true}
     6007 | {"mod": 7, "v6007": true}
     7007 | {"mod": 7, "v7007": true}
     8007 | {"mod": 7, "v8007": true}
     9007 | {"mod": 7, "v9007": true}
    (10 rows)

And we can check that the query is indexed, and only ever reads 10 rows:

    example=> EXPLAIN ANALYZE SELECT k, data FROM tab WHERE data->'mod' = '7';
                                                      QUERY PLAN                                                   
    ---------------------------------------------------------------------------------------------------------------
     Bitmap Heap Scan on tab  (cost=5.06..157.71 rows=100 width=40) (actual time=0.035..0.052 rows=10 loops=1)
       Recheck Cond: ((data -> 'mod'::text) = '7'::jsonb)
       Heap Blocks: exact=10
       ->  Bitmap Index Scan on idx  (cost=0.00..5.04 rows=100 width=0) (actual time=0.026..0.027 rows=10 loops=1)
             Index Cond: ((data -> 'mod'::text) = '7'::jsonb)
     Planning Time: 0.086 ms
     Execution Time: 0.078 ms

If we did not have an index, the query would be slower:

    example=> DROP INDEX idx;
    DROP INDEX
    example=> EXPLAIN ANALYZE SELECT k, data FROM tab WHERE data->'mod' = '7';
                                                QUERY PLAN                                             
    ---------------------------------------------------------------------------------------------------
     Seq Scan on tab  (cost=0.00..467.00 rows=100 width=34) (actual time=0.019..9.968 rows=10 loops=1)
       Filter: ((data -> 'mod'::text) = '7'::jsonb)
       Rows Removed by Filter: 19990
     Planning Time: 0.157 ms
     Execution Time: 9.989 ms

Hence, "arbitrary indices on derived functions of your JSONB data". So the query is fast, and there's no problem with the JSON shapes of `data` being different for different rows.

See docs for expression indices: https://www.postgresql.org/docs/16/indexes-expressional.html

toast0 · on Oct 11, 2023

Either way can work. Getting to millions of messages is going to be the hard part, not storing them.

As with all data storage, the question is usually how do you want to access that data. I don't have experience with Postgres, but a lot of (older) experience with MySQL, and MySQL makes a pretty reasonable key-value storage engine, so I'd expect Postgres to do ok at that too.

I'm a big fan of pushing the messages to the clients, so the server is only holding messages in transit. Each client won't typically have millions of messages or even close, so you have freedom to store things how you want there, and the servers have more of a queue per user than a database --- but you can use a RDBMS as a queue if you want, especially if you have more important things to work on.

sosodev · on Oct 12, 2023

Seems to me like there are still plenty of relationships in a chat app. Postgres can be used like a NoSQL database via JSONB too if you want.

I think the truth is that you should use the simplest, most effective tech possible until you are absolutely certain you need something more niche.

jopsen · on Oct 12, 2023

> If I'm writing a chat app with millions of messages...

Once you have millions of messages, maybe consider moving the data intensive parts out if postgres, if necessary.

The criticism is often that people look for big data solutions, before they have big data.

If you scale out of postgres, you probably have enough users and money that you can fix it :)

But moving to a NoSQL before you have to, might just slow down development velocity -- also you haven't yet learned what patterns users have.

BoorishBears · on Oct 11, 2023

This is going to feel like a non-answer: but if you need to ask this question in this format, save yourself some great pain and use Postgres or MongoDB, doesn't really matter which, just something known and simple.

Normally you'd make a decision like this by figuring out what your peak demand is going to look like, what your latency requirements are, how distributed are the parties, how are you handling attachments, what social graph features will you offer, what's acceptable for message dropping, what is historical retention going to look like...[continues for 10 pages]

But if you don't have anything like that, just use something simple and ergonomic, and focus on getting your first few users. There's a long gap between when the simple choice will stop scaling and those first few users.

airstrike · on Oct 11, 2023

Thanks, I really appreciate this. DynamoDB was pretty simple to setup, all things considered. Some growing pains, but it's a dead simple schema.

I'm using SenseDeep's OneTable which was pretty interesting to learn https://doc.onetable.io/ in case others reading this are curious.

aketchum · on Oct 11, 2023

just migrated off of PG to ddb as the main db for my application (still copying data to SQL for data analytics). Working with distributed functions and code hosted on lambdas, the connection management to SQL became a nightmare with dropped requests all over the place.

jpgvm · on Oct 11, 2023

Sacrificing a real DB to use serverless is exactly why serverless makes zero sense.

makestuff · on Oct 12, 2023

Yeah I have been using Supabase recently and I really like it. You still get the “serverless” benefits but at the end of the day it is just a Postgres database with some plugins. It is super easy to figure out where the data is coming from/going to.

Meanwhile at work I have a cowoker who loves to create AWS soup where they use an assortment of lambdas/api gateways/sqs queues/sns topics to accomplish tasks such as taking files from one s3 bucket and putting them in another s3 bucket owned by a different team. Their justification of this was that it was generic so other teams could use it, but it is a pain to maintain and make changes to.

nameless912 · on Oct 11, 2023

A connection pooling proxy should fix this shouldn't it? I think both Google and AWS have this solved for functions as a service.

aketchum · on Oct 11, 2023

yeah but you run into other issues when you put all of your lambdas on the same VPC

nameless912 · on Oct 11, 2023

Not to be that guy, but why lambdas? I'm genuinely curious. I've never found the "cost savings" (big air quotes) worth it in comparison to the increased configuration/permissions complexity. Especially when Fargate exists, where you can just throw a docker container at AWS, what do Lambdas add? The zero scaling?

qvrjuec · on Oct 11, 2023

With CDK, I can get an ECS service up and running in the same amount of time it'd take to create a lambda function behind API gateway or triggered by SQS/cron. Deploys are easier, cost savings are real, permissions/configuration are the same level of complexity unless you're cutting corners. I'd only use ECS for stuff I know would be high sustained throughput, long duration(>15m) tasks, or things that absolutely need more persistence between executions.

hooverd · on Oct 11, 2023

Serverless is great if you recognize that it's just somebody else's container runtime. I wish there was better tooling for Docker based Lambdas though. I hate whole S3 deployment dance for zip file based Lambdas (yes SAM does it for you now but it's still there).

EC2-backed ECS has a great use case for things that you can run ephemerally in a container but require a persistent data store.

Zanfa · on Oct 11, 2023

Why not? The setup I’m experimenting with for an API right now is basically a single Lambda that’s accessible through a function URL (so no ELB/ALB) + an RDS instance. Spinning up additional environments is a single Cloudformation call and deployment artifacts should work with both Docker containers or S3 (depending on the Lambda execution environment).

Seems like a leaner setup than using ECS/Fargate + LBs to me. Have I overlooked something?

sakopov · on Oct 11, 2023

Your mileage may vary, but I think in majority of the cases Fargate is going to be significantly more expensive than Lambda.

nameless912 · on Oct 11, 2023

But at low request volumes, either is a rounding error to a medium size enterprise, and for personal projects IMO Lambda is a huge PITA.

meowtimemania · on Oct 11, 2023

One of lambda's ideal use cases is personal projects. Personal projects usually serve very few requests so lambda's ability to scale to zero results in cost savings.

hooverd · on Oct 11, 2023

/shill The AWS SAM CLI smooths over a lot of Lambda's rough points. /unshill

nameless912 · on Oct 11, 2023

Gasp, a shill??

I totally believe you, I just can't see how it becomes easier than chucking a container on Fargate or something. Maybe I've just been scarred by lambda rat's nests in the past.

hooverd · on Oct 11, 2023

Yeah, the "proper" way to do Lambdas, shown in so many fancy architecture diagrams, is a rat's nest. I don't like APIs on Lambda unless you can shove them into one container with a catchall proxy on API Gateway. They really shine if you're processing SQS messages or EventBridge events. If you aren't using other AWS services and aren't cost engineering, then Lambdas probably aren't worth the headache.

iends · on Oct 11, 2023

The serverless framework makes lambda for side projects a breeze.

CDK for everything else.

Olreich · on Oct 11, 2023

Lambda is the most expensive thing you can do if you have more than 25% utilization. Fargate is extremely close to modern on-demand EC2 pricing (m7a family).

makestuff · on Oct 12, 2023

Yeah you can run Fargate on EC2 now as well to optimize cost even further.

oarmstrong · on Oct 12, 2023

Fargate on EC2? Sorry you’ve lost me, I understood Fargate to be a layer of AWS managed compute for ECS or EKS deliberately instead of EC2.

makestuff · on Oct 12, 2023

https://lumigo.io/blog/comparing-amazon-ecs-launch-types-ec2... You can now specify your launch type to be an ec2 instance. This has the benefit of lower cost but you are responsible for managing the instances ex: security patches, etc.

oarmstrong · on Oct 14, 2023

Right, running ECS on EC2, not Fargate on EC2. When ECS launched it only had the EC2 launch type (where as you said you must manage your machines). Fargate then came along for both ECS and EKS where Amazon managed the machines for you.

cdelsolar · on Oct 11, 2023

I don't think that's actually true on a GB-second basis (memory * CPU)?

hooverd · on Oct 11, 2023

The cost efficiency of Lambda vs Fargate/EC2 ECS/one of the dozen other ways to run containers on AWS plummets as your RPS goes up.

GabeWeiss_ · on Oct 11, 2023

But that's kind of a moot point. I mean, if you're even looking at the likes of DynamoDB or Spanner, it's because you need the scale of those engines. PostgreSQL is fantastic, and even working for Google, I 100% agree with you. Just use PG...until you can't. Once you're in the realm of Spanner and DynamoDB, that's where this discussion becomes more of a thing.

silisili · on Oct 11, 2023

Not necessarily true. DynamoDB on demand pricing is actually way cheaper than RDS or EC2 based anything for small workloads, especially when you want it replicated.

GabeWeiss_ · on Oct 11, 2023

Fair enough, I'm not (obviously) as familiar with the AWS systems and how the price to performance ratio works out for all the products.

0xbadcafebee · on Oct 11, 2023

Postgres and Spanner do different things, in different ways, with different costs, risks, and implications. You could "just use" anything that is completely different and slightly cheaper. You could use a GitHub repository and just store your records as commits, for free, that's plenty cheap and works for small projects. But not really the same thing, is it?

nameless912 · on Oct 11, 2023

My point is that I've seen very, very few situations (I can think of two in my entire career so far) where a "hyperscale NoSQL database" was actually the right choice to solve the problem. I find that a lot of folks turn to these databases for imagined scale needs, not actual hard problems that need solving.

0xbadcafebee · on Oct 11, 2023

DynamoDB is fantastic for not doing things at scale. It costs a few pennies, there is nothing to set up, it's all managed for you, it is insanely reliable, and it just works. I use it for all kinds of crap that an entire RDBMS is way overkill for.

eklitzke · on Oct 11, 2023

Spanner has a SQL interface, fyi.

nameless912 · on Oct 11, 2023

I don't think the API you interface with fundamentally changes the point that Spanner is hard to recommend from an engineering perspective at anything except the absolute most massive of scales, and even then it will create nearly as many problems as it solves. I'm not saying spanner is _wrong_ or shouldn't exist, but it's very difficult to be in the position where Spanner is the critical key to your application's success and not replaceable by <insert other, cheaper database here>.

endorphine · on Oct 11, 2023

Care to elaborate on what are the problems that it will create? Honest question.

nameless912 · on Oct 12, 2023

Sure. Spanner is expensive, and your primary job as an engineer (if you work for an enterprise like most of us do) is to generate business value. So, if nothing else, you will run into the cost problems of Spanner. There are also other problems; iirc both DynamoDB and Spanner shard their key spaces, and each shard gets the same quota, and the key space shards all have to be the same size. This means that even though you might have paid for 1000rps, for example, that RPS volume is divided across all your shards, so if you have one part of the key space that gets way more volume than another you end up eating up the fractional capacity of that shard way faster than you intend and you have to either overprovision or queue requests, both of which are not ideal.

At a previous job, we ended up creating a very complicated write through cache system in front of spanner that dynamically added memory/CPU capacity as needed to prevent hot shards; our application was extremely read heavy, and writes were relatively low RPS, so this ended up working OK, but we were paying tens of thousands of dollars a month for Spanner plus tens of thousands of dollars a month for all the compute sitting in front of it. I don't think we ended up doing much better than if we had bitten the bullet and run clustered Postgres because our write volume ended up being just a few hundred RPS, even though the read volume was 1000x that. Postgres behind this cache system would have handled the load just as well and cost less than half as much.

The other thing that frustrates me personally about Spanner is that Google's docs are incomplete (as usual); there are lots of performance gotchas like this that exist throughout the entire service, and they aren't clearly documented (unlike, to their credit, AWS with Dynamo, who explains this entire problem very clearly and has an [expensive] prebuilt solution for it in the form of the DynamoDB accelerator).

agonz253 · on Oct 11, 2023

It's not purely a matter of cost, right? Say you want or need a highly available, high performance distributed database with externally consistent semantics. Are you going to handle the sharding of your Postgres data yourself? What replication system will you use for each shard? How will you ensure strong consistency? Will you be able to do transactions across shards? These are problems that systems like Spanner, CockroachDB, etc solve for you.

YetAnotherNick · on Oct 11, 2023

Just curious, why would distributed be design requirement? Is individual machine failure likely in AWS/GCP? The only failure I have seen in region level issues which spanner or dynamo don't help with AFAIK.

agonz253 · on Oct 11, 2023

Individual machine failure is not likely, but we're hypothesizing the need for multiple shards for high performance. So now we have more machines and so the probability of failure increases. So we need to add replication, but then we need to deal with data getting out of sync, etc.... As others have mentioned though, these issues only really become important at a certain scale.

teaearlgraycold · on Oct 11, 2023

Postgres is the database God himself would have made.

nameless912 · on Oct 11, 2023

Like, I don't want to fanboy, but it's _hard_ to find a use case that postgres can't handle at small to medium (i.e. 90% of projects) scale.

0xbadcafebee · on Oct 11, 2023

It's hard to find a use case that a plain old filesystem can't handle at small to medium scale. But there are perhaps more important considerations than just "can it handle it"

diordiderot · on Oct 11, 2023

Like what? (Genuinely curious)

nvm0n2 · on Oct 11, 2023

Uh let's not get carried away. It's fine with enough work maybe. But Postgres has a lot of awkwardness too. HA is a pain, major version upgrades are a pain, JS or JVM stored procs are a pain, configuring auth is a pain. There is a reason so many people are desperate to pay someone else to run Postgres for them instead of just renting a few VMs and doing it themselves.

pier25 · on Oct 11, 2023

Probably more like 99%

hot_gril · on Oct 11, 2023

It's the best, but it's far from perfect. Default mode is non-ACID, and going `serializable` mode makes it very slow. Spanner is always ACID... but always slow.

remus · on Oct 11, 2023

I know the spanner marketing blurb says you can scale down etc. But I think in practice spanner is primarily aimed at use cases where you'd struggle to fit everything in a single postgres instance.

Having said that I guess I broadly agree with your comment. It seems like a lot of people like to plan for massive scale while they have a handful of actual users.

nameless912 · on Oct 11, 2023

I said this in another comment, but I have seen _two_ applications in my career that actually had a request load that might warrant something like one of these databases. One was an application with double digit million MAU and thousands of RPS on a very shardable data set, which fit Spanner's ideal access pattern and performance profile pretty well, but we paid an absolute arm and a leg for the privilege and ended up implementing a distributed cache in front of Spanner to reduce costs. The other just kept the data set in memory and flushed to disk/S3 backup periodically because in that case liveness was more important than completeness.

In the first case, the database created as many problems as it solved (which is true of any large application running at scale; your data store will _always_ be suboptimal). A fancy, expensive NoSQL database won't save you from solving hard engineering problems. At smaller scales (on the order of tens-hundreds of RPS), it's hard to go wrong with any established SQL (or open source NoSQL if that floats your boat) database, and IMO Postgres is the most stable and best bang for your engineering buck feature wise.

ndriscoll · on Oct 11, 2023

Postgres/mysql shouldn't have much trouble doing thousands of RPS on a laptop for basic CRUD queries (i.e. as long as you don't need to do table scans or large index range scans). It's possible to squeeze a lot more than that out of them.

joshuamorton · on Oct 11, 2023

Spanner isn't NoSQL. It's schema'd.

hot_gril · on Oct 11, 2023

My team bought the "scale down" thing and got bit.

Using Spanner is giving up a lot for the scalability, and if you ever reach the scale where a single node DB doesn't make sense anymore, I don't know if Spanner is still the answer, let alone Spanner with your old design still intact. For one, Postgres has scaling options like Citus. Or maybe you don't need a scalable DB even at scale, cause you shard at a higher layer instead.

supportengineer · on Oct 11, 2023

Is there any company that hosts Postgres in the cloud (and does nothing else) and has great customer service?

M3t0r · on Oct 11, 2023

Give https://aiven.io/ a try. But know that I'm biased ;-)

Also https://www.postgresql.org/support/professional_hosting/

LVB · on Oct 11, 2023

I've only kicked the tires, but https://neon.tech is a pure hosted Postgres play. I'd be curious to hear if anyone has used them for a real projects, and how that went.

folmar · on Oct 11, 2023

Elephantsql, I can't say much for the customer service as I never had any trouble.

emodendroket · on Oct 11, 2023

I mean sure, NoSQL gives you more opportunities to screw stuff up because it's doing less for you. But it can be a reasonable tradeoff in some scenarios anyway

paxys · on Oct 11, 2023

Postgres is a piece of software. Cloud Spanner/Dynamo etc are managed services. It makes no sense to directly compare one with the other.

hatsix · on Oct 11, 2023

Sure, You can compare Cloud SQL vs Cloud Spanner and RDS vs Dynamo, but it makes more sense to just say "Postgres" and assume that the reader can figure out that it means "Whatever managed postgres service you want to use".

The entire point is that every cloud provider has a managed postgres offering, and there's no vendor lock-in. Though, technically, Dynamo does have a docker image you could run in other cloud providers if it came down to that, you'd get no support for it.

res0nat0r · on Oct 11, 2023

I don't think it's really relevant to compare plain Postgres to Spanner, most folks have no need for something like this. It is made for folks who need to do millions of ACID type transactions a second/minute from all over the globe, and have a globally consistent database at all times.

There's a reason why Google installed and built their own atomic clocks and put them in their datacenters, it is to facilitate global timekeeping for this type of services. Most likely 99.9% of the time this type of database is overkill, and also likely way more expensive than you need.

https://cloud.google.com/spanner/docs/true-time-external-con...

riku_iki · on Oct 11, 2023

> It is made for folks who need to do millions of ACID type transactions a second/minute from all over the globe, and have a globally consistent database at all times.

I think just doing some (not necessary millions) ACID transactions over the globe and have consistent DB is strong value proposition even for small users.

lijok · on Oct 11, 2023

The dynamodb docker image you’re referring to will get you shot if you try and use it in prod. It’s an API wrapper on top of sqlite and has a ton of missing functionality

There are a couple of databases out there with ddb compatible interfaces, like scylladb

nameless912 · on Oct 11, 2023

> AWS and GCP (and Azure, and Oracle cloud, and bare Kubernetes via an operator, and...) support Postgres really well.

I think they're directly comparable with this context.

esafak · on Oct 11, 2023

Read "Cloud SQL for PostgreSQL"

brianolson · on Oct 11, 2023

"as little as $65 USD/month" for GCP Spanner

vs AWS Free Tier:

"25 GB of data storage ... 2.5 million stream read requests ..."

https://aws.amazon.com/dynamodb/pricing/

So, there's probably somewhere the lines on the graph cross, but Google's headline seems misleading.

hdjjhhvvhga · on Oct 11, 2023

The Free Tier is completely irrelevant here, though. The very reason someone might use Spanner is its excellent scalability. I don't believe there is any reason to use it for smaller projects other than education. The customers who will use Spanners are those for whom CockroachDB is not enough, for example. For everybody with databases that are not that huge PostgreSQL will do just fine.

_ugfj · on Oct 11, 2023

> For everybody with databases that are not that huge PostgreSQL will do just fine.

Ha. Remember Gary Bernhardt of WAT fame? https://twitter.com/garybernhardt/status/600783770925420546

> Consulting service: you bring your big data problems to me, I say "your data set fits in RAM", you pay me $10,000 for saving you $500,000.

jiggawatts · on Oct 11, 2023

AWS will happily rent you a server with 24 TB of memory for about $200/hour.

Columnar databases typically get a 10:1 compression ratio over raw data = 240 TB effectively.

That’s a lot of data.

RhodesianHunter · on Oct 11, 2023

Which columnar databases are doing the above in-memory?

jiggawatts · on Oct 11, 2023

SQL Server supports in-memory columnstore tables. I’m not an expert but I suspect SAP HANA also.

If you squint, any database engine is “in memory” if there is more buffer than data.

Or just use a RAM disk!