top of page

The Year I Stopped Writing About AI

Aug 13
17 min read

What three years of public forecasts got right, and the one that was never made


A spiral notepad headed "Next Year's Predictions" with seven numbered lines, all blank.

In December 2020 I sat down to write a list of things that would matter in real estate technology the following year. I had written the same list in 2017, 2018, and 2019.


It ran under my byline, it took about a week to assemble, and I was reasonably good at it.


The 2021 list had six items.

  • Landlord and tenant relationships.

  • Office space planning.

  • Tenant communications.

  • Digital operations in multifamily.

  • Fraud.

  • Flexibility.


Not one of them was artificial intelligence.


I had written about AI in the two lists before that one. In December 2018 it was item seven. In November 2019 it was item two, under a heading that sounds embarrassing now and did not then: "AI gets (more) real."


Then, in December 2020, I dropped it. Twenty-four months later, ChatGPT was released.


Why you should be skeptical of this

Everyone publishes predictions. Almost nobody publishes the scorecard.

The Four Seats begins here, with one piece and no track record on this masthead, which means the honest thing to do is address why you would spend fifteen minutes reading it.



The industry I work in produces an enormous volume of forward-looking opinion. Vendor blogs, analyst notes, conference keynotes, LinkedIn posts, annual trend lists exactly like the ones I used to write. Almost none of it is ever checked. The pieces are published, they circulate for a week, and then they are gone, and the person who wrote them writes next year's without reference to last year's. The genre has no memory, which is precisely what makes it comfortable to produce.


So rather than open a publication by telling you what I think is coming, I want to open it by showing you my record and grading it in front of you.


Three years of dated predictions, still online, still under my name, none of them edited since publication. If the grading is unflattering, that is useful information about how much weight to give anything I write afterward.


There is a second reason, and it is the more interesting one. Grading the record turned up a failure I did not expect and had not seen described anywhere, and it is a failure that is about to happen to a very large number of people building AI plans right now.


What I learned at Gartner and then stopped doing

The apparatus was not decoration. It was the mechanism that forced revision.

I spent part of my career as a VP of research at Gartner, which is a place that has thought harder about the structure of forward-looking claims than almost anywhere else, and where I learned the craft I later practiced badly on my own.


The core unit there was the Strategic Planning Assumption. An SPA is a claim written to a fixed shape: by a stated year, a stated proportion of a stated population will do a stated thing.


By 2027, over 40% of agentic AI projects will be canceled.

By 2022, 75% of enterprise-generated data will be created outside the traditional data center, up from less than 10% in 2018.


That form is not house style. Every element of it is load-bearing.

  • The date makes the claim expire.

  • The percentage makes it measurable.

  • The baseline makes the trajectory checkable.

  • And SPAs carry probability ratings, which means the analyst is not only making a claim but stating how confident to be in it, which is a separate discipline and a harder one.


Gartner's own definition of an SPA is the part I want you to sit with.


It is described as the best answer to a key issue at a given point in time.


At a given point in time. The impermanence is written into the definition of the thing. An SPA is not a prophecy. It is a position on a moving question, held provisionally, and the entire apparatus around it exists to force the position to be revisited.


Around each SPA sits what I will call the revision apparatus: the standing set of mechanisms that force a published position to be looked at again on a fixed schedule, whether or not anyone feels like looking. It is more elaborate than most outsiders realize, and almost all of it does that one job.


Magic Quadrants are re-run on a cycle, and vendors move between quadrants year over year. Anyone who has been evaluated in one knows the movement is often more informative than the position, because the position is a snapshot and the movement is a trajectory.


Hype Cycle (my favorite framework) placements advance along the curve annually, which means every technology on it gets re-examined on a schedule whether or not anything interesting happened to it that year. That last clause is the important one. Nothing has to warrant attention for the attention to occur. It is on the calendar.


Research goes through internal challenge before it publishes, which means an analyst has to defend a claim to people whose job that day is to break it. And the analyst sits in continuous contact with vendors and clients, which is a steady stream of evidence arriving whether or not it is welcome.


Take the whole structure together and what it amounts to is a set of tripwires. Not one of them guarantees a correct call. Every one of them guarantees that a stale call gets touched.


None of this makes the calls correct. Gartner has published SPAs that aged terribly, in public, with dates on them.


By 2018, more than three million workers globally would be supervised by a robotic boss. That did not happen and there was no ambiguity about it not happening, because the claim was dated and quantified and therefore capable of being wrong in a way that anyone could check.


Which is the point. The method does not produce accuracy. It produces falsifiability and a schedule, and those two things together produce revision.


I left, and I kept making predictions, and I dropped every part of the apparatus.


What I did instead

I did look at the prior year. Looking is not reviewing.

My lists had none of that structure and I did not notice, because the writing felt like the same activity.


There were no dates inside the claims. "AI will continue to gain momentum" cannot expire. There were no proportions, so nothing was measurable. There were no probabilities, so I never had to distinguish between a call I would defend and one I was including because it seemed prudent to mention.


I want to be careful here, because there is a tidier version of this story that is not true. The tidier version is that each December I opened a blank document and started from nothing. That would make this a memory problem, and memory problems are fixable by being more diligent.


What actually happened is that I did look at the prior year's list. Of course I did. I wrote it. It was in my head, and I read it again, and I weighed it against what my own team was building and what the market was clearly demanding that quarter. That is a real input and it is the right input. It is also the entire failure.


Looking at last year's list informally does two things that a review does not. It produces the feeling of having checked without producing a record of what was checked, so nothing distinguishes an item I considered and set aside from an item that never entered my mind at all. Both leave the same trace, which is none.


And it lets recency do the sorting invisibly. When you are holding last year's list in your head alongside a live roadmap and a market with its hair on fire, the things that surface are the things that are loud. That is not a lapse in attention. That is how attention works.


The Gartner version is not more thoughtful than what I did. It is more procedural, and that turns out to be the point. An SPA review is a meeting with an output. Somebody says the words renew, revise, or retire, and the answer gets written down next to a claim that has a date on it. What I did was recall, and recall has no output. Nothing is ever recorded as retired, because nothing is ever formally considered.


An apparatus does not make you smarter. It makes certain questions produce artifacts. Remove it and those questions do not become harder to ask. They stop producing anything when you ask them, which feels identical from the inside.


The lists themselves were not sloppy. They were researched, argued internally, edited, and better than most of what the category produced. Judged as individual documents, each was competent work. That was always the problem. They were only ever judged as individual documents, so each one could be evaluated in full, come back clean, and still be missing the only thing that mattered.


Grading three years

Structure was easy, mechanism was where I failed, and the largest failure left no trace

Three of those lists are still live, still under my byline, and still say what they said when I published them: predictions for 2019, published December 21, 2018. Trends for 2020, published November 21, 2019. Trends for 2021, published December 2, 2020.


These lists are what I am grading, and I want to be exact about why. They are not the complete run. They are the ones you can go and check. An earlier list for 2018 is no longer reachable at the address where it lived, and a piece I cannot point you to is a piece I have no business asking you to take my word about. Grading only what remains verifiable is a smaller claim than grading everything I ever wrote, and in an essay about whether anyone ever checks these things, the smaller claim is the only honest one.


The calls sort into three groups, and the groups have very different reliability.

Group

What I called

Verdict

Structural

Open platforms beat closed suites. Cloud goes mainstream. Digitized transactions attract fraud.

Held, and they were not hard

Mechanism

A downturn inside two years, called in 2018 and again in 2019. The leasing office empties out, called in 2020.

Right outcome, wrong mechanism, three times

Absent

Artificial intelligence, tracked in 2018 and 2019, missing from the 2020 list

No call made, and no way to notice

The structural calls held. In December 2018 I gave two of eight slots to the same argument, that closed all-in-one suites were finished and interoperability was becoming the price of admission rather than a differentiator. My evidence was IBM buying Red Hat and Salesforce buying MuleSoft, plus industry interoperability bodies re-emerging.


That call is now so ordinary it feels like stating that water is wet. In the same list I said cloud had moved into the mainstream and that single sign-on and multi-factor authentication would settle the security objections keeping IT departments on premises. That happened.


The best of them was almost an aside. December 2020, item five: moving transactions online means fraud goes up. By 2025 the National Multifamily Housing Council found 93.3% of rental housing providers had encountered fraud in the prior twelve months, that the average respondent wrote off roughly $4.2 million in bad debt, and that falsified income documentation accounted for 84.3% of it. I called the direction and badly undercalled the magnitude.


Someone will point out that none of these were hard.


Open beats closed was already consensus among people paying attention in 2018. Cloud adoption was a trend line you could extend with a ruler. More online transactions producing more fraud is close to arithmetic.


The objection is correct, and it is most of the point. Structure is legible. Pressure in a system points somewhere, and if you are close enough to the system you can feel which way it points. That is not prophecy, it is proximity, and proximity is available to anyone actually doing the work. Which is why the failures are the interesting part.


The mechanism calls were luck wearing a suit. December 2018, buried at the end of an item about predictive analytics, I wrote that while we aren't economists, it's common knowledge that a downturn is expected to happen, likely in the next two years. Two years from that December is December 2020. The downturn arrived in March.


I said it again in November 2019, hedging on the most anticipated downturn in history and citing a conference-floor poll putting it in the second half of 2020.


Right window, twice, four months out. And completely wrong.


What I was describing was a credit cycle: rates, occupancy softening, the ordinary late-expansion arithmetic every real estate professional could feel in 2018. What arrived was a respiratory virus. My prediction and the event share a date and nothing else.


A forecast that lands for the wrong reason is not a forecast. It is a coincidence with a timestamp, and treating it as a forecast is how people build confidence in methods that do not work.


I would have called that a one-off if it had not happened again, in a place where I could not blame a virus.


In December 2020 I wrote that multifamily operators would reevaluate the leasing office, and I noted that new properties were already being built without a physical one. The outcome was right and it arrived faster than I expected. The on-site leasing office has been shrinking ever since.


But read what I actually wrote about how it would happen. I said operations would shift from a centralized model to a distributed one, meaning the work would come off the single physical desk and spread across digital channels.


The opposite occurred. The industry calls the last five years centralization, and the word is exact: leasing, lease administration, marketing, accounting, and maintenance dispatch were pulled off the property and consolidated into shared service teams covering many communities at once. By 2024, NMHC and One11 Advisors found 48.4% of multifamily respondents had implemented centralization programs with another 51.6% planning to, and the stated benefit was fewer people on site.


The work did not distribute. It concentrated somewhere else, and the property lost the desk because the desk moved to a hub rather than dissolving into a network.


I got the visible result right and the direction of travel backwards. Anyone grading me on outcomes gives me full marks for that call. Anyone grading me on reasoning notices that I would have built the wrong product on the back of it, because a distributed model and a shared-services model need different software, different staffing, and different economics.


This is the failure the SPA structure is built to expose, because a claim with a stated mechanism and a probability attached can be wrong about the mechanism while being right about the outcome, and the review will catch it. My version had no mechanism written down, so there was nothing to catch.


The workplace call sits inside the same failure. In October 2020 I published a framework arguing the pandemic had separated activity from place, and that the separation would hold where an activity was transactional and reverse where it was experiential.

Banking stays remote. Live music comes back. That mechanism held up well. I put work in the second column and predicted collaboration would pull people back into offices.


Wrong, and wrong in an instructive way: the mechanism was right and my classification was not. Culture and collaboration are genuinely experiential. But a large share of knowledge work is producing a deliverable at a desk, and that is a transaction, and transactions do not come back.


The one thing I will take credit for is not sitting on it. Nine months later, on July 2, 2021, I published a piece in GlobeSt on how to make the hybrid model work, and the premise had already moved from whether hybrid would hold to how to operate it. By then we had gone and surveyed landlords, office tenants, and employees rather than continue reasoning from the framework.


Being wrong in public is survivable. Staying wrong is not.


The call I did not make

A wrong prediction files its own correction. An absent one never does.

Now the part that actually keeps me up, and the reason this piece exists.


In December 2018 I wrote about AI. It was hedged, and reading it back it is the writing of a man covering a topic rather than making a claim: these technologies are already being used, we have seen reasonable adoption among leading clients, we will likely see more widespread use as they become easier for mid-market companies to adopt.


In November 2019 I gave it a better slot and a real claim. I argued AI in real estate had moved past the shiny-object phase toward implementation, and I named two places it would land first: lease abstraction, and natural language processing in customer-facing roles.


Both were correct. Lease abstraction became one of the earliest genuinely useful applications in commercial real estate. Customer-facing NLP is now the leasing chatbot, the maintenance triage bot, the tour scheduler.


Three years and one week before ChatGPT was released, I named the two beachheads correctly. Then in December 2020 I left AI off the list entirely.


I know exactly why. Every square inch of that list was claimed by the pandemic. Rent deferrals, lease concessions, space planning, cleaning protocols, online payment adoption, fraud. Those were real problems and my readers had them that month. Nothing about that list was lazy. Every item on it was something an operator would face within ninety days.


And that is precisely the mechanism. I did not conclude AI had stopped mattering. I never made a judgment about AI at all. It lost a competition for space against six things that were on fire, and losing that competition looked exactly like sound editorial judgment, because in every local sense it was.


Remember that I had the prior year's list in front of me, at least in the way I described earlier, which is to say in my head and in the general shape of what we were working on.


So AI did not vanish because I forgot it existed. It was there, in the pile, and it did not make the cut against six things that operators were going to face inside ninety days. Weighed against rent deferrals in December 2020, a technology whose most concrete applications were lease abstraction and a leasing chatbot loses, and it deserves to lose, on any reasonable reading of what my reader needed that month.


The demotion of AI in that list should worry anyone building a roadmap right now, because it was not a bad call. It was a good call, produced by a sound process, reading the market accurately.


Here is what the apparatus would have done differently, and it is narrow.


An SPA has a date on it, so at that date somebody has to say a word out loud: renew, revise, or retire. The word gets written next to the claim.


A Hype Cycle placement gets re-examined annually whether or not the technology did anything interesting that year.


Neither of those would have made me smarter about AI in December 2020. What they would have done is force the demotion to be spoken. I would have had to look at lease abstraction and customer-facing natural language processing, the two things I had correctly identified twelve months earlier, and say the word retire. I would not have said it. Nobody would have said it. And in refusing to say it, I would have kept the claim alive on a schedule that outlasted the news cycle that buried it.


Instead the demotion happened silently, in my head, in the ordinary course of deciding what my reader needed most.


Call it a silent demotion: when a position leaves your working set without ever being evaluated, because something more urgent took its slot. A silent demotion is not a decision. It is an outcome, and it is indistinguishable from a decision until the moment it costs you something.


If you had asked me in January 2021 what I thought about AI in real estate, I would have given you a reasonable answer, because I still held the view. The view had not changed. It had simply stopped being written down, and a view that stops being written down stops being tested, updated, and defended. It becomes a thing you believe rather than a thing you maintain.


Eighteen months later the ground under that unmaintained view moved further and faster than it had in the previous decade, and I had no mechanism to notice, because noticing was something I did once a year in a document I had quietly stopped including it in.


The asymmetry between a wrong call and an absent one is the whole finding of this piece. A wrong prediction is cheap to detect. It is written down, it has a date, and reality eventually files the correction. My workplace call was wrong in October 2020 and substantially fixed by July 2021 because the wrongness was visible and stayed visible until I dealt with it.


An absent prediction has none of those properties. Nothing in my December 2020 list was incorrect. There was no error to notice, no claim to test, no feedback to receive. The failure leaves no residue. I could have graded that list in 2022 on everything it contained and given myself a good score, and the score would have been meaningless, because the thing that mattered was not in the document.


We audit what we said. We do not audit what we stopped saying.


And the reason is structural rather than lazy. Every review process I have sat in, on any side of the table, evaluates artifacts. The deck, the roadmap, the plan, the list. Artifacts contain what someone decided to include, and the review interrogates those contents rigorously. No review I have ever attended asked what was in last year's artifact and is missing from this one. There is no agenda item for it and no owner. The question is nobody's job, which is why nobody asks it.


The exercise

Three steps, one hour, and the third one is the only one that matters

The silent demotion matters in 2026 rather than as autobiography, because a very large number of people are currently building AI roadmaps, and roadmaps are lists. Lists have finite slots. Slots get allocated by urgency, and urgency is not reliably correlated with importance.

  • Pull your list from two years ago. The actual document. Board deck, strategy memo, offsite output, annual planning artifact, whatever your version is. Find it rather than remember it, because memory grades generously.

  • Grade it in three columns, not two. Right, wrong, and right-for-the-wrong-reason. The third column does the work. Anything sitting in it is a place where you have been rewarding a method that did not earn it, and rewarded methods get reused.

  • Then list what fell off. Compare that document to the one from the year before. Put the two side by side, on paper, rather than holding one and reading the other. What was on the earlier list and not the later one? For each item, answer one question: did I decide this stopped mattering, or did it lose a competition for space?


Note the word decide. You almost certainly did think about the item, in the same loose way I did, which is why the honest answer is rarely that you forgot. The question is whether you ever reached a conclusion you would have been willing to write down at the time. If not, you did not retire the item. You silently demoted it, and a silent demotion leaves nothing behind to argue with later.


The second answer is the dangerous one. It means the item left your attention without ever being evaluated, which means you hold no position on it, which means you will not notice when it changes.


Two things make the third step harder than it sounds. It will look unremarkable, because items drop off plans constantly and most of them deserve to. You are looking for the one or two where the honest answer is that something urgent beat something important. In my case there was exactly one, and it mattered more than the six things that beat it combined.


The exercise also implicates whoever built the list, who is usually sitting in the room. Run it on a document old enough that nobody is defending it. Two years back, not last quarter. The point is to test the process, not to grade a colleague.


What this site is for

Claims that can expire, and no pretending I have the apparatus

Which brings me back to why you would spend time here.


I am not going to reproduce Gartner's apparatus. I do not have a research organization, a peer review process, or a vendor briefing calendar, and pretending otherwise would be its own kind of dishonesty.


What I can do is take the part that actually mattered, which was never the quadrants or the curves. It was writing claims that can expire.


So: pieces here will run long and infrequently, four or so a year. Each will make claims specific enough to be wrong, which means dated where a date is honest and quantified where a quantity is honest, rather than the ambient "will continue to gain momentum" that cannot fail because it never said anything.


Notice what that costs you rather than me. A dated claim expires whether or not I come back to it. You do not need my permission or my schedule to check one, which is the entire difference between a position and a mood.


I am not going to promise you a revision cycle. Four pieces a year is already a thin publishing schedule, and a man who has just spent three thousand words explaining that he failed to maintain his own views should be careful about announcing a maintenance program. What I will do instead is narrower and I can actually keep it: when I come back to a subject I have written about before, I will grade the earlier piece first, in the open, before adding anything new to it.


The promise is modest and deliberately so. I am not promising to be right. The record above is a reasonable guide to how often I will be, and it is a mixed record with one embarrassing hole in it. What I am promising is that when I am wrong, you will be able to tell.


That is the only promise worth making here, because checkability is the thing I quietly dropped year after year while believing I was doing the work.


The predictions I got wrong cost me some credibility, which is recoverable.

The prediction I did not make cost me two years, which is not.


There are four seats in the name. The fifth one is yours.

Comments


bottom of page