Ask a database vendor about on-prem and you'll usually get a one-line answer: 'we're open source, you can just self-host it.' For a hobby project, fine. For an enterprise, that line rarely survives contact with production. And enterprises want on-prem more than ever right now, largely because of AI. Running open models on hardware you own is far cheaper than renting inference by the token, and in a regulated industry, keeping data where you control it isn't optional. When the models move in-house, the database moves with them, because that's where the data lives and, increasingly, where the AI feeds on it.So a lot of teams are asking a question they thought they had retired: can we run this in our own environment, on our own terms? Plenty of vendors have a ready answer. “Sure, we are open source. Just self-host.” For a solo developer or a small team, that answer holds up fine. For an enterprise, it usually falls apart, because open source, self-hostable, and enterprise-ready are three different promises, and most vendors only deliver the first.
Last week, I spent three days in the Netherlands and gave two talks at two conferences: a lightning talk at PGDay Lowlands in Utrecht on Thursday, September 10, and a session at Percona Live in Amsterdam on Friday, September 11\. In this blog post, I’m going to share my notes from both.
As often happens with conferences (or any big events, really), there was a minor hurdle to overcome before we could get there. On Wednesday, September 9, just one day before PGDay Lowlands, a nationwide 24-hour public transport strike stopped trains, buses, trams and metros across the whole country. Not the ideal warm-up for a conference that draws people from all over the world, but by Thursday morning everything was moving again and the day went ahead as planned. Yay!
PGDay Lowlands, Utrecht {#pgday_lowlands_utrecht}
PGDay Lowlands is a one-day Dutch PostgreSQL conference (although all the talks are in English), organized by PostgreSQL Europe. This was its third edition, and the event moves around: last year, it was held at Blijdorp Zoo in Rotterdam; this year, it took place at TivoliVredenburg, a music venue in the center of Utrecht, with the main track in a hall called Cloud Nine.
As a SQL Server developer and DBA learning Postgres, it’s easy to expect that the information you need for meaningful query and performance tuning will be readily available. For years (decades, maybe) you’ve learned the DMVs, set up Extended Events sessions, relied on Query Store, and regularly run Ola Hallengren’s maintenance scripts and Brent Ozar’s First Responder Kit. Nearly everything you do to find and tune poorly performing queries happens through SQL or through a GUI in SSMS.
Rarely, if ever, do you think about combing through logs to find query performance issues. The error log is where you go when something broke: a failed startup, a corruption message, a login from an IP that shouldn’t exist, a backup that didn’t. It’s an incident destination, not a daily instrument.
Most SQL Server DBAs I talk to have also never had to think hard about log configuration, because there was never a decision to make. Logging is built into the Windows server ecosystem. It just exists, and you get it for free.
It’s no wonder, then, that SQL Server DBAs who are new to Postgres have real confusion about where to find the information they need when there’s a problem. And it’s no wonder so many are shocked when they discover the information isn’t there at all, because Postgres was never configured to record it.
This doesn’t mean Postgres monitoring is worse. In some cases it’s markedly better. Postgres actually gives you significantly more configuration options around what gets tracked and logged. They’re just set conser
[...]
In PostgreSQL, every tuple starts with 23-byte header, and the first eight bytes are two transaction IDs. t_xmin for the transaction that created the row and t_xmax for the one that deleted or updated it. That is the visibility story covered in PostgreSQL MVCC, Byte by Byte. For now we have discussed t_xmax acting as the delete marker.
t_xmax has a second job. When you run SELECT ... FOR UPDATE or an insert checks a foreign key, PostgreSQL has nowhere else to record the row lock. The shared memory lock table is limited by max_locks_per_transaction. Locking a million rows would exceed its capacity. PostgreSQL works around this by storing the locking transaction ID in t_xmax and marking the row as locked with flags in t_infomask, while readers can still see it, so every row lock in PostgreSQL ends up as a write to the page.
The setup is one parent table in the usual shape, plus a child table with a foreign key, since foreign key checks lock parent rows. Everything below was captured on a single PostgreSQL 18.6 cluster using the postgres:18 image. Transaction IDs will be different on your cluster; compare the bits instead.
CREATE EXTENSION IF NOT EXISTS pageinspect;
CREATE TABLE lock_demo (
id integer GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
owner text NOT NULL,
balance numeric(12,2)
);
CREATE TABLE lock_demo_tx (
id integer GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
account_id integer NOT NULL REFERENCES lock_demo (id),
amount numeric(12,2)
);
INSERT INTO lock_demo (owner, balance)
VALUES ('alice', 100.00), ('bob', 200.00), ('carol', 300.00);
SELECT count(*) FROM lock_demo;
SELECT lp, t_xmin, t_xmax, t_ctid,
(heap_tuple_infomask_flags(t_infomask, t_infomask2)).raw_flags
FROM heap_page_items(get_raw_page('lock_demo', 0));
lp | t_xmin | t_xmax | t_ctid | raw_flags
----+--------+--------+--------+-[...]
Postgres 19 won't make the expected release date. PostgreSQL has shipped its major version every fall for the last several years. But this year, the code is in a heavy review cycle, major features have been reverted during beta, and many others are under heavy revision. Beta 4 is scheduled for Sept. 24, 2026. A release of Postgres 19 is certainly delayed by weeks and maybe even months.
Postgres 19 was an ambitious release already, with a lot of large features. With any project, you have to choose priorities. For Postgres, the priorities were quality, followed by shipping within the time window. The team is reducing the scope of the release to get closer to meeting its timeline with the quality it requires.
I'll break down some of the major reversions in Postgres 19. While this is a long list, I want to make it abundantly clear that the PostgreSQL code development process is working beautifully. Code is getting tested at a wide scale and things that aren't ready are getting pulled.
A typical PostgreSQL code development process works like this:
There have been quite a few reverts for version 19 of Postgres: 53 since the beta began in June 2025. Below are some of the more major user-facing features you may be familiar with:
Property graphs would add graph query views on top of tables and a light approach to graph queries in Postgres. The hackers discussion suggests broader design and readiness concerns.
In the final part of this special Postgres in Production deep dive series, Ryan Booz does something the first six episodes rarely did: he queries pg_stat_statements itself. This episode covers why your first stop during an incident should actually be pg_stat_activity, two ways to get a usable window out of cumulative metrics (diffing snapshots, or resetting and re-querying), which columns to order by and why the slowest query is not always your problem, and what to look for when you pick a monitoring tool to keep this history for you.
Share this episode: Click here to share this episode on LinkedIn. Feel free to sign up for our newsletter and subscribe to our YouTube channel.
Transcript
Believe it or not, through these last six episodes we’ve rarely queried the data itself. We’ve talked about what pg_stat_statements is and isn’t (Part 1), how query texts get normalized (Part 2), where the texts themselves are stored and how that can get contentious (Part 3), and we’ve looked at the source code to see exactly what happens once your query finishes executing (Part 4). We covered configuration (Part 5) and, in the la
Over the past few years, I have been working on assessing the strengths and weaknesses of PostgreSQL migration tools. Several articles published by yours truly led to the ambitious project I have been pursuing since 2024 with my colleagues Étienne Bersac and Pierre-Louis Gonon.
… And the fisrt stable version of PostgreSQL Migrator was released on September 4th. This is an opportunity to showcase features I use daily and what advantages they offer over other tools. In this article, I want to focus on one of them, particularly valuable when preparing a migration: the offline catalog.
The catalog of a relational database contains the structure of the data model, table column names and data types, constraint definitions, the definition of a view or a function, and so on. Everything declared by the user with DDL (Data Definition Language) is stored in the catalog as the single source of truth.
In systems like PostgreSQL, MySQL or MSSQL Server, the standard provides a universal catalog, the information_schema schema. For example, table names can be retrieved there with the same query:
SELECT table_name FROM information_schema.tables
WHERE table_schema = 'scott'
ORDER BY table_name;
However, this is not the ideal solution to reconstruct a data model. Each of these systems conforms to the SQL standard as best it can but very often enriches it with language extensions or takes liberties with the implementation of a feature. As a result, the information_schema catalog is not the universal source, little more than a set of views on top of each system’s proprietary system catalog.
Turning to Oracle and Ora2Pg. If we want to recreate the structure of a table in the Oracle ecosystem, several methods exist and they all rely on the catalog views which I cover last.
The DESCRIBE command
By far the least informative of the solutions but the fastest for a first inspection. It is analogous to the \d meta-command in psql or the pragma table_info in SQLite.
D
If you were wondering, it’s the SBOM thing that I mentioned the other day.
Postgres Extensions in containers with full inventory, provenance and attestation. I’ve been using plenty of AI Agents to put this together. This blog is a little scattered (apologies) but as they say, I didn’t have time to write a short letter.
Here’s what I believe to be the exact structure of current CloudNativePG images:
I have a bunch of irons in the fire related to this project:
After working on this for like a week, I realized that the set of commands to validate the security and provenance info was going to be totally different for PGRX than for upstream.
I think this is too confusing to users. There needs to be one simple, consistent command to verify provenance and SBOM material. I didn’t have that at the beginning and the AI Agent enabled racing ahead with the code. I didn’t realize the issue until now.
One thing I had been focused on was not having a bunch of things to copy, if someone needs to mirror images to a private container registry. With that focus, this is what I had i
[...]
There is a conversation that happens in every PostgreSQL shop eventually. A query that has been fine for a year gets slow overnight. Nothing was deployed. The data grew a little, ANALYZE ran, and the planner — entirely reasonably, on the numbers it had — picked a different plan. The old plan was better. You would like it back.
PostgreSQL 19 ships two new modules for exactly this: pg_plan_advice, which can read a plan back out as a string and enforce it later, and pg_stash_advice, which keeps those strings keyed by query id and applies them automatically.
Until now, the only way to tell us about a contribution worth listing was email. There is now a form on the site instead.
You do not need an account, and you do not have to tell us who you are. Write the contribution in one box, in plain text or Markdown, and send it. If you would like us to be able to come back to you with a question, there is an optional line for your name and email address.
Many contributions to and for the PostgreSQL Project happen outside of writing code: talks, meetups, translations, patch reviews, mentoring, advocacy, documentation, event organisation. If you know of one, whether it is yours or somebody else's, we would like to hear about it.
Nothing is published automatically. Everything that comes in gets read by one of us first. Email still works, if you prefer it.
Obviously I would have liked to post this closer to the event, but life got in the way. The talk was in May and the recording has been up for a while. Here it is.
PGConf.dev 2026 took place from May 19th to 22nd, 2026, at Simon Fraser University's Harbour Centre campus in downtown Vancouver, BC, Canada. It was my first one, and I had been meaning to go for years.
If you're familiar with other large Postgres events like PGConf.EU or PostgreSQL@SCaLE, PGConf.dev is a different beast. It is the PostgreSQL development conference, and the whole event is arranged around the work rather than around an audience. An in-person commitfest ran in one of the rooms on the Wednesday and Thursday, so patches were being reviewed in the building while the talks were going on. There were community office hours. The Friday was dedicated entirely to the unconference, with the schedule built on the day from whatever the attendees proposed.
For anyone who wants to get their hands dirty and get closely involved in the project, this is the conference to go to. You are in a room with the people who write and commit the code, and the barrier to walking up and asking them something is about zero. I cannot recommend it highly enough.
On the Wednesday afternoon I presented "PostgreSQL Commitfest Metrics: A Quantitative Analysis" together with Andreas "ads" Scherbaum (EDB). We had been pulling data out of the Commitfest application, and spent a while working out what it says about what happens to a patch after somebody sends it in.
Because this is the sort of subject that is easy to misread, we opened by saying what the talk was not. It is not a critique of any contributor, it is not a critique of any committer, and it is not a claim that anything is broken. It is an observation rather than a diagnosis, and we deliberately stopped short of recommending any fixes. What we wanted was to put the numbers on the table and let the project examine them.
We looked at 58 commit
[...]Every now and then I need a break from writing code. In those cases I like looking at data about a subject I’m interested in - looking for trends, quantifying the expected effects, and so on. I needed just such a break a couple days ago, and I decided to look at statistics about the development activity of the Postgres project. So, here’s a bunch of charts (with a bit of commentary).
On 8 September 2026, PGDay UK 2026 was held in London.
Organized by:
Program Committee:
Code of Conduct Committee:
Speakers:
On 10 September 2026, PGDay Lowlands 2026 was held in Utrecht, NL.
Organized by:
Program Committee:
Code of Conduct Committee:
Speakers:
Debaters:
From 9-11 September, the following community members staffed the PostgreSQL booth at Percona Live, including:
On October 1, I’ll be speaking at PG Summit 2026 about benchmarking hardware with Postgres. I’m looking forward to sharing some tips and learnings that I’ve picked up over the years. In anticipation of the presentation, I just wanted to share a little bit about some of my motivations for the topic.
If you want to know how fast a disk is, use fio. If you want to know how quickly a CPU can perform a particular operation, there are better tools for that too. Those tests can tell you something useful about an individual component, and the numbers on the product page might even be meaningful in that context.
But a Postgres benchmark is not really trying to reproduce the number in the marketing material (A Samsung EVO Plus 990 is marketed at read/write speeds up to 7,150/6,300MB/s, but I don’t think we’ll hit that on a legit Postgres cluster). It’s important to remember that Postgres is a complicated system with memory, concurrency, caching, WAL, checkpoints, background workers, and several kinds of maintenance that can all happen at once. A benchmark that runs for ten seconds might measure a very fast and very warm slice of that system while missing the things that make production interesting.
An experienced DBA or DBRE knows that autovacuum creates work while the workload is running, and checkpoints can create bursts of I/O. A high-concurrency workload can run into lock contention or connection limits before the storage device itself is particularly busy. Even cache state changes the question: are we measuring a workload that fits comfortably in memory, or one that has to keep reading from storage? This is why it is difficult to use Postgres to make a clean statement about a piece of hardware. There are too many other things involved, and they are not noise to be discarded. They are part of the database we are trying to operate.
The useful question is not, “How fast is this SSD?” It is so
[...]In the past couple weeks, I’ve learned more about renovate, SBOMs, provenance and attestations than I ever wanted to know. (But if I’m being honest, I do enjoy learning a bit more about it.)
Backstory is that I decided to make CNPG-Extensions an actually serious project. The original name was “Not-CNPG” as a joke about CNCF’s restrictive licensing policies which forbid hosting open source software with licenses like GPL. https://github.com/cnpg-extensions/
As a “serious” project I wanted to provide provenance info so users can more have assurance about the contents of a container image, and so that scanners can accurately report licenses and compare software versions against vulnerability databases. This week I also started exploring support for pgrx extensions with full rust dependency graphs in the SBOM so that tools like trivy can flag RUSTSEC vulns even on packages buried in the dependency tree.
Example Trivy output for a Debian-based extension:
Example Trivy output for a pgrx-based extension (this is not final):
A few things I’ve learned along the way:
Logical replication has been part of Postgres since version 10, and the syntax page that governs it is almost comically brief. wants a name, a connection string, a list of publications, and then it offers one innocuous line:That single line expands to more than a dozen options, and several of them change how logical replication uses storage resources. After all, subscriptions created with no clause work perfectly fine. Most subscriptions in the wild don't need these tweaks, and nobody thinks about that line again, if they ever knew it existed at all.Imagine a production system boasting several downstream logical replicas. Consider disk monitors lighting up and flagging the directory. It’s suddenly filling with thousands of anonymous artifacts, and nobody seems to know what is writing there or why. Well, it is Postgres writing in that directory, and the "why" is a longer story.So what lives in that directory? What makes it balloon to terrifying proportions seemingly at random? How is logical replication involved? Is there any way to control or even stop this behavior?The answers lie inside that very same innocuous and esoteric WITH clause. Let's see what's going on here.
Being "the database guy" comes with a lot of questions, and over the last eight months those questions changed. The repetitive ones disappeared, nobody asks how to avoid putting things in the database any more, and the code arriving for review got noticeably more polished. Then this summer a schema landed in front of me with twelve proposed index drops on a single table, which is when I started assuming coding agents over-index. The next schema I looked at had the same shape.
Passing this off as AI slop would be too easy, because most of those changes were competent. So I built a harness and measured it.
I loaded 30 model-generated schemas into PostgreSQL and audited 838 indexes across twelve of them. The competence caught me off guard. Only ten served no requirement I could find; the rest showed solid craft. All four models handled GIN and GiST indexes cleanly, built partial indexes with sensible predicates, and got multi-tenant composite keys in the right order. The baseline SQL quality is much better than what agents wrote a year ago.
Indexes are great, until you pile them onto the single table taking all your writes. On quiet tables you will never notice the difference. On hot tables, every index is extra work on every write.
In one support-tool schema, a model created sixteen indexes on tickets alone. Six of them indexed last_activity_at, a column that updates every time an agent touches a ticket. Compared to my hand-written baseline with seven indexes, the generated schema wrote 1.8× the WAL, took 1.9× longer per update, and pushed up VACUUM time just as much.
Those sixteen indexes were not dumb mistakes. For read queries, they run fast. The problem is that coding agents write indexes query by query, without thinking about write traffic.
What actually happens on disk when you touch that row:
Number of posts in the past two months
Number of posts in the past two months
Get in touch with the Planet PostgreSQL administrators at planet at postgresql.org.