I used AI tools to size an OLTP workload on an EC2 system, with DBT-5, a TPC-E-like fair-use implementation. I provided a systematic and mechanical plan for running a series of tests to determine what the appropriate scale factor is on a system.
I pre-configured the database with some settings that are known to be needed to be changed, such as shared_buffers and max_wal_size but more on this at a later time when we characterize the system behavior further to be able to tune some of these settings better. Remember, this is an iterative process when trying to figure it out for any workload.
I decided to use Claude Fable 5.1 for this exercise and fed in the following instructions:
The following chart illustrates the sum of all the testing from starting at a 5000 customer database, up to a 97,000 customer database:
We need to zoom in a little bit to see that the best result for this system is at 32,000 customers with 24 users: It's worth mentioning that during this exercise, Claude also ran some smoke tests at various times to make sure everything was working. There were some minor fixes, but a significant improvement was to actually spend the time to allow multiple Trade Results and Market Feed transactions to be handled concurrently. It's been pointed out at lea[...]The future of Postgres is bright. So bright in fact, that I spent almost a dozen posts expounding on the upcoming features it would bring, with veritable stars in my eyes. Unfortunately, while I was out counting my chickens, it would seem some of these exciting new features failed to hatch.How, and perhaps more importantly why that happened, deserves some investigation.
Postgres events in Scotland now have a permanent address: postgres.scot
It is a deliberately small page. Right now it points at the PostgreSQL Edinburgh User Group (PostgresEDI) on cloomba, where you will find RSVPs, a calendar subscription and an RSS feed, and it will carry other Scottish Postgres events as they come along. The point is to have one address worth bookmarking and sharing, rather than whichever platform we happen to be on this year.
If you are running something Postgres-related in Scotland and want it listed, email info (at) postgres.scot.
postgres.scot is a volunteer website, not affiliated with or endorsed by the PostgreSQL project or the PostgreSQL Community Association.
Thursday, August 13th, Paterson's Land at the University of Edinburgh, two talks, pizza and refreshments sponsored by pgEdge, and the rest of the evening at the Tolbooth Tavern, one of the few pubs nearby that was not hosting an Edinburgh Fringe show that night.
Torsten Förtsch
Torsten Förtsch on replaying the stream of changes into the target database
Torsten started from a move to Aurora and the question of what "your data" actually means once the database is somebody else's service. His answer was to rebuild point-in-time recovery at the logical level: a pg_dump for the base copy, a stream of changes captured with wal2json in place of archived WAL segments, and replay into a database you control, on whatever operating system and Postgres version you like.
He took us through both halves of that. Capture and replay turned out to be 35 lines of jq translating the JSON change stream into SQL statements, with some care over how those statements are written so they find the right row quickly. The harder half is the initial copy: working out which position in the stream a dump corresponds to, so that replay starts in exactly the right place. He finished with where he wants to take it, incl
Two client connections, opened one after the other through PgBouncer transaction mode, asked the same database whose rows they were allowed to see, and the second one got the first one's answer. It had set nothing, it had never met the first caller, and the rows it read belonged to that caller's tenant. Any runbook that keeps a tenant key in a session variable sits one pooler away from this, and since 28 July the MCP protocol carries no session of its own, so a tool call from an agent arrives in this shape by default.
The setup is the one most multi-tenant guides teach. A row level security policy reads current_setting('app.tenant', true), each request or tool call opens with SET app.tenant, and on a connection nobody else is using the rows that come back belong to whoever asked for them. Through the pooler, the first call still behaves as written.
SET app.tenant = 'a';
SET
SELECT current_setting('app.tenant') AS tenant;
tenant
--------
a
(1 row)
SELECT tenant, body FROM docs;
tenant | body
--------+----------------
a | alpha invoice
a | alpha contract
(2 rows)
A second client connection, opened after the first one had closed and setting nothing of its own, then asks the database who it is working for and what it may read.
SELECT current_setting('app.tenant', true) AS inherited_tenant;
inherited_tenant
------------------
a <-- set by the caller before it
(1 row)
SELECT tenant, body FROM docs;
tenant | body
--------+----------------
a | alpha invoice
a | alpha contract
(2 rows)
Both callers ran on one backend, as did the two after them that switched the tenant to b and inherited it. When caller A's implicit transaction ended, PgBouncer released that backend and handed it over without sending anything in between that would have cleared app.tenant. Two changes made that shape the ordinary one. On 28 July 2026 the MCP specification took the session out of the protocol, its announcement stating that "Each request now travels on its
Ask a database vendor about on-prem and you'll usually get a one-line answer: 'we're open source, you can just self-host it.' For a hobby project, fine. For an enterprise, that line rarely survives contact with production. And enterprises want on-prem more than ever right now, largely because of AI. Running open models on hardware you own is far cheaper than renting inference by the token, and in a regulated industry, keeping data where you control it isn't optional. When the models move in-house, the database moves with them, because that's where the data lives and, increasingly, where the AI feeds on it.So a lot of teams are asking a question they thought they had retired: can we run this in our own environment, on our own terms? Plenty of vendors have a ready answer. “Sure, we are open source. Just self-host.” For a solo developer or a small team, that answer holds up fine. For an enterprise, it usually falls apart, because open source, self-hostable, and enterprise-ready are three different promises, and most vendors only deliver the first.
Last week, I spent three days in the Netherlands and gave two talks at two conferences: a lightning talk at PGDay Lowlands in Utrecht on Thursday, September 10, and a session at Percona Live in Amsterdam on Friday, September 11\. In this blog post, I’m going to share my notes from both.
As often happens with conferences (or any big events, really), there was a minor hurdle to overcome before we could get there. On Wednesday, September 9, just one day before PGDay Lowlands, a nationwide 24-hour public transport strike stopped trains, buses, trams and metros across the whole country. Not the ideal warm-up for a conference that draws people from all over the world, but by Thursday morning everything was moving again and the day went ahead as planned. Yay!
PGDay Lowlands, Utrecht {#pgday_lowlands_utrecht}
PGDay Lowlands is a one-day Dutch PostgreSQL conference (although all the talks are in English), organized by PostgreSQL Europe. This was its third edition, and the event moves around: last year, it was held at Blijdorp Zoo in Rotterdam; this year, it took place at TivoliVredenburg, a music venue in the center of Utrecht, with the main track in a hall called Cloud Nine.
As a SQL Server developer and DBA learning Postgres, it’s easy to expect that the information you need for meaningful query and performance tuning will be readily available. For years (decades, maybe) you’ve learned the DMVs, set up Extended Events sessions, relied on Query Store, and regularly run Ola Hallengren’s maintenance scripts and Brent Ozar’s First Responder Kit. Nearly everything you do to find and tune poorly performing queries happens through SQL or through a GUI in SSMS.
Rarely, if ever, do you think about combing through logs to find query performance issues. The error log is where you go when something broke: a failed startup, a corruption message, a login from an IP that shouldn’t exist, a backup that didn’t. It’s an incident destination, not a daily instrument.
Most SQL Server DBAs I talk to have also never had to think hard about log configuration, because there was never a decision to make. Logging is built into the Windows server ecosystem. It just exists, and you get it for free.
It’s no wonder, then, that SQL Server DBAs who are new to Postgres have real confusion about where to find the information they need when there’s a problem. And it’s no wonder so many are shocked when they discover the information isn’t there at all, because Postgres was never configured to record it.
This doesn’t mean Postgres monitoring is worse. In some cases it’s markedly better. Postgres actually gives you significantly more configuration options around what gets tracked and logged. They’re just set conser
[...]
In PostgreSQL, every tuple starts with 23-byte header, and the first eight bytes are two transaction IDs. t_xmin for the transaction that created the row and t_xmax for the one that deleted or updated it. That is the visibility story covered in PostgreSQL MVCC, Byte by Byte. For now we have discussed t_xmax acting as the delete marker.
t_xmax has a second job. When you run SELECT ... FOR UPDATE or an insert checks a foreign key, PostgreSQL has nowhere else to record the row lock. The shared memory lock table is limited by max_locks_per_transaction. Locking a million rows would exceed its capacity. PostgreSQL works around this by storing the locking transaction ID in t_xmax and marking the row as locked with flags in t_infomask, while readers can still see it, so every row lock in PostgreSQL ends up as a write to the page.
The setup is one parent table in the usual shape, plus a child table with a foreign key, since foreign key checks lock parent rows. Everything below was captured on a single PostgreSQL 18.6 cluster using the postgres:18 image. Transaction IDs will be different on your cluster; compare the bits instead.
CREATE EXTENSION IF NOT EXISTS pageinspect;
CREATE TABLE lock_demo (
id integer GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
owner text NOT NULL,
balance numeric(12,2)
);
CREATE TABLE lock_demo_tx (
id integer GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
account_id integer NOT NULL REFERENCES lock_demo (id),
amount numeric(12,2)
);
INSERT INTO lock_demo (owner, balance)
VALUES ('alice', 100.00), ('bob', 200.00), ('carol', 300.00);
SELECT count(*) FROM lock_demo;
SELECT lp, t_xmin, t_xmax, t_ctid,
(heap_tuple_infomask_flags(t_infomask, t_infomask2)).raw_flags
FROM heap_page_items(get_raw_page('lock_demo', 0));
lp | t_xmin | t_xmax | t_ctid | raw_flags
----+--------+--------+--------+-[...]
Postgres 19 won't make the expected release date. PostgreSQL has shipped its major version every fall for the last several years. But this year, the code is in a heavy review cycle, major features have been reverted during beta, and many others are under heavy revision. Beta 4 is scheduled for Sept. 24, 2026. A release of Postgres 19 is certainly delayed by weeks and maybe even months.
Postgres 19 was an ambitious release already, with a lot of large features. With any project, you have to choose priorities. For Postgres, the priorities were quality, followed by shipping within the time window. The team is reducing the scope of the release to get closer to meeting its timeline with the quality it requires.
I'll break down some of the major reversions in Postgres 19. While this is a long list, I want to make it abundantly clear that the PostgreSQL code development process is working beautifully. Code is getting tested at a wide scale and things that aren't ready are getting pulled.
A typical PostgreSQL code development process works like this:
There have been quite a few reverts for version 19 of Postgres: 53 since the beta began in June 2025. Below are some of the more major user-facing features you may be familiar with:
Property graphs would add graph query views on top of tables and a light approach to graph queries in Postgres. The hackers discussion suggests broader design and readiness concerns.
In the final part of this special Postgres in Production deep dive series, Ryan Booz does something the first six episodes rarely did: he queries pg_stat_statements itself. This episode covers why your first stop during an incident should actually be pg_stat_activity, two ways to get a usable window out of cumulative metrics (diffing snapshots, or resetting and re-querying), which columns to order by and why the slowest query is not always your problem, and what to look for when you pick a monitoring tool to keep this history for you.
Share this episode: Click here to share this episode on LinkedIn. Feel free to sign up for our newsletter and subscribe to our YouTube channel.
Transcript
Believe it or not, through these last six episodes we’ve rarely queried the data itself. We’ve talked about what pg_stat_statements is and isn’t (Part 1), how query texts get normalized (Part 2), where the texts themselves are stored and how that can get contentious (Part 3), and we’ve looked at the source code to see exactly what happens once your query finishes executing (Part 4). We covered configuration (Part 5) and, in the la
Over the past few years, I have been working on assessing the strengths and weaknesses of PostgreSQL migration tools. Several articles published by yours truly led to the ambitious project I have been pursuing since 2024 with my colleagues Étienne Bersac and Pierre-Louis Gonon.
… And the fisrt stable version of PostgreSQL Migrator was released on September 4th. This is an opportunity to showcase features I use daily and what advantages they offer over other tools. In this article, I want to focus on one of them, particularly valuable when preparing a migration: the offline catalog.
The catalog of a relational database contains the structure of the data model, table column names and data types, constraint definitions, the definition of a view or a function, and so on. Everything declared by the user with DDL (Data Definition Language) is stored in the catalog as the single source of truth.
In systems like PostgreSQL, MySQL or MSSQL Server, the standard provides a universal catalog, the information_schema schema. For example, table names can be retrieved there with the same query:
SELECT table_name FROM information_schema.tables
WHERE table_schema = 'scott'
ORDER BY table_name;
However, this is not the ideal solution to reconstruct a data model. Each of these systems conforms to the SQL standard as best it can but very often enriches it with language extensions or takes liberties with the implementation of a feature. As a result, the information_schema catalog is not the universal source, little more than a set of views on top of each system’s proprietary system catalog.
Turning to Oracle and Ora2Pg. If we want to recreate the structure of a table in the Oracle ecosystem, several methods exist and they all rely on the catalog views which I cover last.
The DESCRIBE command
By far the least informative of the solutions but the fastest for a first inspection. It is analogous to the \d meta-command in psql or the pragma table_info in SQLite.
D
If you were wondering, it’s the SBOM thing that I mentioned the other day.
Postgres Extensions in containers with full inventory, provenance and attestation. I’ve been using plenty of AI Agents to put this together. This blog is a little scattered (apologies) but as they say, I didn’t have time to write a short letter.
Here’s what I believe to be the exact structure of current CloudNativePG images:
I have a bunch of irons in the fire related to this project:
After working on this for like a week, I realized that the set of commands to validate the security and provenance info was going to be totally different for PGRX than for upstream.
I think this is too confusing to users. There needs to be one simple, consistent command to verify provenance and SBOM material. I didn’t have that at the beginning and the AI Agent enabled racing ahead with the code. I didn’t realize the issue until now.
One thing I had been focused on was not having a bunch of things to copy, if someone needs to mirror images to a private container registry. With that focus, this is what I had i
[...]
There is a conversation that happens in every PostgreSQL shop eventually. A query that has been fine for a year gets slow overnight. Nothing was deployed. The data grew a little, ANALYZE ran, and the planner — entirely reasonably, on the numbers it had — picked a different plan. The old plan was better. You would like it back.
PostgreSQL 19 ships two new modules for exactly this: pg_plan_advice, which can read a plan back out as a string and enforce it later, and pg_stash_advice, which keeps those strings keyed by query id and applies them automatically.
Until now, the only way to tell us about a contribution worth listing was email. There is now a form on the site instead.
You do not need an account, and you do not have to tell us who you are. Write the contribution in one box, in plain text or Markdown, and send it. If you would like us to be able to come back to you with a question, there is an optional line for your name and email address.
Many contributions to and for the PostgreSQL Project happen outside of writing code: talks, meetups, translations, patch reviews, mentoring, advocacy, documentation, event organisation. If you know of one, whether it is yours or somebody else's, we would like to hear about it.
Nothing is published automatically. Everything that comes in gets read by one of us first. Email still works, if you prefer it.
Obviously I would have liked to post this closer to the event, but life got in the way. The talk was in May and the recording has been up for a while. Here it is.
PGConf.dev 2026 took place from May 19th to 22nd, 2026, at Simon Fraser University's Harbour Centre campus in downtown Vancouver, BC, Canada. It was my first one, and I had been meaning to go for years.
If you're familiar with other large Postgres events like PGConf.EU or PostgreSQL@SCaLE, PGConf.dev is a different beast. It is the PostgreSQL development conference, and the whole event is arranged around the work rather than around an audience. An in-person commitfest ran in one of the rooms on the Wednesday and Thursday, so patches were being reviewed in the building while the talks were going on. There were community office hours. The Friday was dedicated entirely to the unconference, with the schedule built on the day from whatever the attendees proposed.
For anyone who wants to get their hands dirty and get closely involved in the project, this is the conference to go to. You are in a room with the people who write and commit the code, and the barrier to walking up and asking them something is about zero. I cannot recommend it highly enough.
On the Wednesday afternoon I presented "PostgreSQL Commitfest Metrics: A Quantitative Analysis" together with Andreas "ads" Scherbaum (EDB). We had been pulling data out of the Commitfest application, and spent a while working out what it says about what happens to a patch after somebody sends it in.
Because this is the sort of subject that is easy to misread, we opened by saying what the talk was not. It is not a critique of any contributor, it is not a critique of any committer, and it is not a claim that anything is broken. It is an observation rather than a diagnosis, and we deliberately stopped short of recommending any fixes. What we wanted was to put the numbers on the table and let the project examine them.
We looked at 58 commit
[...]Every now and then I need a break from writing code. In those cases I like looking at data about a subject I’m interested in - looking for trends, quantifying the expected effects, and so on. I needed just such a break a couple days ago, and I decided to look at statistics about the development activity of the Postgres project. So, here’s a bunch of charts (with a bit of commentary).
On 8 September 2026, PGDay UK 2026 was held in London.
Organized by:
Program Committee:
Code of Conduct Committee:
Speakers:
On 10 September 2026, PGDay Lowlands 2026 was held in Utrecht, NL.
Organized by:
Program Committee:
Code of Conduct Committee:
Speakers:
Debaters:
From 9-11 September, the following community members staffed the PostgreSQL booth at Percona Live, including:
On October 1, I’ll be speaking at PG Summit 2026 about benchmarking hardware with Postgres. I’m looking forward to sharing some tips and learnings that I’ve picked up over the years. In anticipation of the presentation, I just wanted to share a little bit about some of my motivations for the topic.
If you want to know how fast a disk is, use fio. If you want to know how quickly a CPU can perform a particular operation, there are better tools for that too. Those tests can tell you something useful about an individual component, and the numbers on the product page might even be meaningful in that context.
But a Postgres benchmark is not really trying to reproduce the number in the marketing material (A Samsung EVO Plus 990 is marketed at read/write speeds up to 7,150/6,300MB/s, but I don’t think we’ll hit that on a legit Postgres cluster). It’s important to remember that Postgres is a complicated system with memory, concurrency, caching, WAL, checkpoints, background workers, and several kinds of maintenance that can all happen at once. A benchmark that runs for ten seconds might measure a very fast and very warm slice of that system while missing the things that make production interesting.
An experienced DBA or DBRE knows that autovacuum creates work while the workload is running, and checkpoints can create bursts of I/O. A high-concurrency workload can run into lock contention or connection limits before the storage device itself is particularly busy. Even cache state changes the question: are we measuring a workload that fits comfortably in memory, or one that has to keep reading from storage? This is why it is difficult to use Postgres to make a clean statement about a piece of hardware. There are too many other things involved, and they are not noise to be discarded. They are part of the database we are trying to operate.
The useful question is not, “How fast is this SSD?” It is so
[...]In the past couple weeks, I’ve learned more about renovate, SBOMs, provenance and attestations than I ever wanted to know. (But if I’m being honest, I do enjoy learning a bit more about it.)
Backstory is that I decided to make CNPG-Extensions an actually serious project. The original name was “Not-CNPG” as a joke about CNCF’s restrictive licensing policies which forbid hosting open source software with licenses like GPL. https://github.com/cnpg-extensions/
As a “serious” project I wanted to provide provenance info so users can more have assurance about the contents of a container image, and so that scanners can accurately report licenses and compare software versions against vulnerability databases. This week I also started exploring support for pgrx extensions with full rust dependency graphs in the SBOM so that tools like trivy can flag RUSTSEC vulns even on packages buried in the dependency tree.
Example Trivy output for a Debian-based extension:
Example Trivy output for a pgrx-based extension (this is not final):
A few things I’ve learned along the way:
Number of posts in the past two months
Number of posts in the past two months
Get in touch with the Planet PostgreSQL administrators at planet at postgresql.org.