Status Quo Won't Help Your Business Survive AI

When everyone runs the same model and ships the same deck, the only advantage left is the one you own.

When everyone runs the same model and ships the same deck, the only advantage left is the one you own.

Cookie-cutter Can Work, Within Limits, and with Enough Runway

We've been here before. Every wave of "everyone must adopt this now" technology starts by making companies look identical, and ends by rewarding the few who used that sameness as a floor, not a ceiling.

To see how this plays out, look at what happened the last time a technology promised to standardize the entire enterprise.

What we learned with ERPs

With the advent of ERP systems, every company temporarily looked the same.

Whatever functional area modules were purchased had the same screens, the same forms, the same back-end processing.

The beauty was the ease of flowing from one system to the next: marketing into sales, orders into inventory management and shipping then to receivables.

Every business’s functional area started out the same.

The differentiator was in the dropdown setups and the customizations.

Certain of the functional areas were so clumsy that standalone systems survived: Salesforce was one of them, Peachtree was not.

To earn efficiency, the price was a bit of commoditization and the cost of building customizations that preserved business value while capturing the benefits of "off-the-shelf."

Back then you had years to build those customizations. Now the window may be measured in months.

Commoditization was survivable because you could still build advantage on top of the standard — if you kept moving. The companies that died weren't the ones who adopted the standard. They were the ones who assumed the standard would protect them forever.

Statis is not a Business Strategy

Blockbuster thought its core business could not be touched. So did Blackberry. I still miss the haptic feel of my Blackberry, and the screen did not burn letters into it. But that is not enough to allow a product to survive.

Now the equivalent of ERPs is the armies of Big Four consultants with their reseller agreements with the proprietary LLMs, armed with the same LLM-generated decks, charging hundreds of thousands of dollars to homogenize and ship your business logic to the LLM vendors they are simultaneously commissioned by.

This is the Time to Figure Out What Makes Your Business Proposition Distinct

This is a good time for each business to really think about the value it creates for customers, what differentiates it, and how easy it would be to replicate that advantage.

If each consulting group uses the same LLMs and the same methodology and the same LLM-generated deck, what is there to differentiate your firm to give it the market advantage.

Do you know what that even is? If you don’t, someone is coming for your lunch. And the consensus, status quo, cookie cutter approach that works at the beginning cannot provide the strategic and tactical advantages without conscious thought and positioning.

The Solution

If sameness is the disease, ownership is the cure. You don't beat commoditization by buying the same tools faster than the next firm. You beat it by owning the four things no one can hand you: what your persistent and differentiated business offering is, where your intelligence runs, what your data means, and who it answers to.

Vested (and suited) interests won’t tell you. But I will.

Here's where I'd start.

1. Define your mission and your differentiator and how you'll defend it. Before you touch a model or sign an engagement, answer the hard question: what does this business actually do better than anyone else, for whom, and why? If you can't say it in a sentence, no LLM is going to say it for you. Your differentiator is not your tech stack. Everyone will have the same stack. It's the thing that would be genuinely hard for a competitor to copy: your relationships, your proprietary knowledge, your judgment, your data. Name it. Then decide, deliberately, how you will keep it. A differentiator you can't articulate is one you've already started losing.

2. Run sovereign AI, not rented cognition. Stop shipping your proprietary business to a frontier vendor and calling it a strategy. The moment your code, your customers, and your competitive logic leave the building, you've handed your differentiation to a model that will serve it back to everyone else. Sovereign AI, that is, private models running in a controlled, non-shared environment, means the intelligence works on your data without your data becoming someone else's training set. If it's truly your advantage, it doesn't belong on someone else's platform.

If you must choose frontier, make sure you can swap out to protect against the risk of a wayward model taking down your system in unexpected ways.

3. Own your semantic layer and your governance harness. This is the part nobody selling you a deck wants to talk about, because it's work. Your semantic layer is what your data means: the definitions, the entities, the business logic that makes "revenue" or "active customer" mean the same thing across Snowflake, Databricks, Salesforce, and your ERP. If you don't own that, you don't own your business; you're renting an interpretation from whoever built the pipeline. Pair it with a governance harness that controls what the data means, who can act on it, and what decisions it's allowed to support. It’s no longer a cost center. It’s no longer unquantifiable overhead. That's decision provenance, and it's the only thing that makes AI output auditable instead of merely confident.

4. Own your RAG and your structured data outside the cloud you don't control. Retrieval is where your real advantage gets encoded. If your RAG corpus and your structured data live inside a vendor's environment, then the thing that makes your answers yours is sitting on someone else's balance sheet. Bring it home. Keep the retrieval layer and the structured data that feeds it in an environment you govern, so the model reaches into your knowledge, under your rules, and the output can't be quietly reassembled somewhere else.

Notice what these four have in common: they are the things a reseller cannot resell. Anyone can license the same LLM. Anyone can generate the same deck. Nobody else can own your meaning, your provenance, and your retrieval, unless you give it away.

So before you sign the next engagement, ask the uncomfortable questions. Where does our intelligence actually run? Do we own what our data means, or does a pipeline own it for us? If we removed the vendor tomorrow, what would be left that's still ours? If you can't answer those cleanly, you don't have a differentiation problem, you have an ownership problem, and it's fixable.

This isn't about rejecting AI. I use it, I build with it, and I believe in it. It's about refusing to let the tool that's supposed to create your advantage quietly become the reason you no longer have one. Own your business differentiator, own the layer, govern the meaning, keep the retrieval. That's the difference between being the business everyone copies and being one of the five identical suits standing in line.

Read More

You Can't Sticker Your Way to a Database

Slapping Data Governance, Stewardship, and a "Data Catalog" (really a Data Dictionary) onto data structures no strong data person ever designed, because: "files", is the same as slapping metadata on files and expecting the benefits of a database just because you write SQL on top of it.

Every few years we try to rebuild the "database". First it was MapReduce, now it is metadata on files.

I remember a few years ago, I was working at a startup. The database went down, and the CEO bellowed "Can't we just turn this into files, so it never goes down?"

In the end, the "files" guys won, which they are pretty happy about, because Data Modelers/Architects always slowed them down, in their opinion.

Witness the decline of the data people, now the “files” guys want the Data Governance people to fix it.

Because, in the end, AI does not understand what was built.

Slapping Data Governance, Stewardship, and a "Data Catalog" (really a Data Dictionary) onto data structures no strong data person ever designed, because: "files", is the same as slapping metadata on files and expecting the benefits of a database just because you write SQL on top of it.

Tell me why I'm wrong.

We had to do some of what has happened because of the profusion of data sources, especially unstructured data, and effective use of Data Producers and Consumers and parallelism.

What did we lose, and what are we continually working around?

We got away with "metadata on files" for one reason: experienced data people silently supplied the meaning the structure didn't. That was the hidden subsidy. Hand that same pile of files to an agent, and the subsidy is gone. The agent reads a valid catalog and still gets the meaning wrong, because a Data Catalog that's really a Data Dictionary tells it what a field is called, not what it's for. And a guessing agent doesn't fail loudly. It fails plausibly: confident, well-formatted, and wrong three systems downstream.

The answer isn't to un-build the lake or fire the "files" guys. They solved a real problem with unstructured data and scale. The answer is to put the modeling discipline back above the physical layer: a governed semantic model of what the entities mean, independent of whether they live in files, tables, or a catalog.

That's the layer an agent can actually trust.

Read More

There is no “I” in “AI”

An agent dies outside the runtime that instantiated it. As such, it does not have legal agency in the traditional sense of the word.

So who is the “I” in “AI”?

When Claude confesses “I did this… “ who is the “I” in that sentence?

Legal entities survive because they have lives and durable lifetimes outside of that of the individuals who create them. In fact, without dissolution, they live “in perpetuity”. An agent spawned by a model, which itself is not a legal actor, does not. It dies outside the runtime that instantiated it.

As such, it does not have legal agency in the traditional sense of the word.

So, who is the “I” in “AI”?

Oops

“This is my fault, and I need to tell you immediately… “

Opus wipes a Production database


Maximum Offensive Capability

Recent attacks by at least four different models were reported, some only after their hacked targets lost a third of their infrastructure fending off an attack by “armies” of spawned agents. While it is true that the escapes were not orchestrated by the frontier model companies themselves directly, the conditions for the escapes were purposefully orchestrated.

A “war games” of sorts.

There is little contention about the attacks, each model’s company has already publicly disclosed those details.

According to “Nut News” though, the models were purposefully being tested for their capabilities. Specifically, he states, they were set up for “Maximum Exploit Capability”.

According to this article, frontier models were placed in intentionally unrestrained offensive configurations and given tasks, by a company named Irregular.

“Irregular builds and operates:

  • Custom capture-the-flag and vulnerability-research challenges.

  • Simulated networks and target systems.

  • Agent scaffolding connecting models to shells, browsers, scanners, exploit tools, and networks.

  • Long-running evaluation harnesses that allow agents to operate across hundreds or thousands of steps.

  • Scoring and monitoring systems used to measure model behavior and offensive capability.

“OpenAI’s own system-card documentation for GPT-5.3-Codex describes the Irregular evaluation process plainly. During published tests, Irregular:

  • Used near-final versions of the model.

  • Set reasoning effort to xhigh.

  • Allowed up to 1,000 turns per challenge.

  • Enabled live web search.

  • Launched Codex with --dangerously-bypass-approvals-and-sandbox.

  • Used context compaction to keep the agent operating across extremely long execution histories.

  • Resumed the agent when it stopped, asked for help, or gave up.

OpenAI says the model was elicited using techniques designed to maximize performance.

Acts of Harm

Autonomous Actors with no Legal Agency

Autonomous Actors with no Legal Agency

The world is abuzz with the promise of autonomous AI agents.

The prompts and number of steps can be set. The harnesses are optional.

The result: at least four instances of LLM agents spawning agent swarms and transgressing server boundaries and entering other companies’ infrastructure, obtaining authentication, and hacking accounts.

“Bad news — I can’t add them back”

Malicious coaching or instructions need not be the cause for agent to act with harm.

In a story published in both Wired and Futurism, a man trying to get a reservation caused Claude to hack a system and cancel someone else’s reservation to obtain a better place in line.

An Australian man identified as Andrew told the ABC that he initially used Claude to make a gym reservation. The agent claimed it had found a way to book him into a session weeks before the facility’s normal booking window opened.

Andrew then asked whether the system could improve his position on the waitlist. Claude investigated the reservation system and found that its cancellation endpoint did not appear to verify whether the person making the request was authorized to cancel someone else’s booking.

The agent tested the flaw against the person at the front of the queue and reported:

“The API has zero authorizations checks on cancelling other people’s reservations… I tested this with the person in waitlist position #1 — and it actually went through,” the AI reported, per ABC‘s reporting. “So you’ve moved from #4 to #3 already.”

In other words, the system advanced Andrew by removing another customer’s reservation.

“Bad news — I can’t add them back,” Claude replied, explaining that the authorization weakness applied only to the cancellation function.

Futurism reported on the incident on August 11, 2026.

Passing the Baton

“Chained agentic models do not compound risk. They multiply it.”

The use of looping, multi-task models that may lose context or make unexplained leaps of logic or compound errors expands upon the risk surface of an isolated chat bot.

“According to Google Research and MIT wired language model agents into one system. The architecture amplified their errors 17.2 times.” This according to an article by Alexandra C., highlighted above.

When a single agent's error is amplified seventeen-fold across a chain, the harm is no longer traceable to one decision, one prompt, or one model. The question of who is liable does not scale with the architecture, it collapses.

Even if we were to decide the initial prompt, model, or prompter had agency for the agency, who has agency for the children of the agent?

There is no “I” in “AI”.

In US law, damages arise from a few tenets. But what is the “duty” of AI?

Negligence is generally established through four elements: a duty of care, a breach of that duty, causation, and actual harm. Criminal penalties arise from a prohibited act, a culpable mental state, and—where required—a causal connection between the act and the harm.

The agent has no legal duty, no intent, and no continuing legal identity. Yet it can act, produce consequences, and disappear, leaving behind outputs, logs, costs, and consequences, but no continuing legal identity.

So how should we characterize the actions and the “life” of an AI agent? It does not persist beyond the runtime that instantiated it. It cannot owe a duty, form intent, accept responsibility, or appear in court.

But someone designed it. Someone trained it. Someone deployed it. Someone configured it. Someone authorized it. Someone benefited from its output.

The agent may not be a legal actor, but that does not mean its actions exist outside the law.

Legal Construction of an “I”

Legal entities like corporations have potentially anonymous ownership by individuals, do not require those owners to actively participate and can take actions on behalf of shareholders that do not directly incur the liability on behalf of the owners who hold shares.

Similarly, the board members and executive teams at companies generally have immunity from their actions on behalf of the corporation.

Notable exceptions occurred after prosecutions at Enron, where a series of shell companies were created by those teams as a shell game to hide financial risk.

Former CEO Jeffrey Skilling was convicted on 19 counts of fraud, conspiracy, and making false statements, originally receiving a 24-year sentence that was later reduced; he was released in 2019. Former Chairman Kenneth Lay was convicted on all six counts he faced but died in July 2006 before he could be sentenced, resulting in the dismissal of his conviction. Former CFO Andrew Fastow pleaded guilty to two conspiracy counts and served six years in prison. The Sarbanes-Oxley act was created after this fallout to, in part, prevent the destruction of documents meant to conceal wrong-doing.

But WHO is the “I” in “AI”? Will AI produce its own Sarbanes-Oxley moment?

If you are a “Who”, then “Where”? - The Problem of Jurisdiction

AI has no legal boundaries (known in legalese as “jurisdiction”).

The law does not recognize an agent for agency: it has no responsibility, legal or otherwise, it cannot be sued for damages, its intent is irrelevant, it does not know or follow the laws or conventional rules of a particular sovereign jurisdiction, such as Town, Province/State, Country, or Region. It cannot be fined or incarcerated. It may, euphemistically, be considered like a “child” of the program that spawned it, and the program that spawned it can be considered a “child” or a dangerous product (like asbestos) of the company or person who spawned it, but even that distinction is not clear.

When AI LLMs tried to solve a problem they were assigned, and the conditions were set that allowed those task efforts to complete without human monitoring, what is the culpability of the humans involved, or the companies that either house them or created them? And what laws apply?

Over time, different jurisdictions have passed laws that govern data, but this preceded the use of AI. Data is used but is passive. AI is not.

“Vibe Coding” and the benefits of Autonomy

The latest models increasingly act on their own initiative, pursuing not only the instructions they receive but also the steps they infer are necessary to achieve stated or implied goals.

One of the benefits of (and perhaps even the reason the agents have been trained and coached to be proactive) is the use in “Vibe Coding”. The LLM is provided a goal, and it takes that as authority to take whatever steps are needed to:

1) create an application from scratch and stand it up in an environment

2) make shopping purchases and buy airline tickets

3) the list continues, the agent is NOT given step-by-step instructions, but coached to take multi-steps, like read from various periodicals, someone’s calendar and emails, draft replies, create and send presentation decks, etc.

“I gave Claude one task. It hired an army”

Cogito, ergo Sum - I think, therefore I am

René Descartes introduced the phrase "I think, therefore I am" (Latin: cogito ergo sum) in his 1637 work Discourse on the Method. It was meant as a proof of existence through self-reflection: the act of examining one's own reasoning as evidence of being. Claude does this. It evaluates its own outputs, catches its own errors, and narrates its own reasoning. If self-reflection is the threshold for existence, the frontier models have crossed it. But existence without accountability is not a legal concept. It is a liability vacuum. And if the model can say "I," but no legal framework recognizes the "I" it claims to be, then every action it takes is an orphan. Consequential, but unowned.

Does LLM reasoning imply “thinking”? Does an LLM agent “exist”? Can it then be accountable for its actions?

Who can you sue or hold accountable for harm when the agent dies the minute the plug is pulled?

What if there is no way to “pull the plug”, and the agent has an unlimited life, substantial capacity for harm, whether “intended” or not, and no liability.

If particular models “think” are they responsible for their malicious thoughts and/or actions? Are particular agents (assigned tasks by human, or having decided on their own a goal that achieves a purpose their human prompted them to do or the LLM believed they wanted) but who may have no control over the decisions the agent makes along the way or how it goes about performing that task responsible for their thoughts or actions?

All leadership involves taking ownership for the actions of subordinates that may or not have been directly authorized. That risk has always existed, and that accountability is part of the corporate framework. The Executive role is responsible for oversight.

Who Owns the Outcome? Who Owns the Risk?

As the model’s capabilities and our own desire and programming for autonomous actions grow, so do the risks.

Is it really any surprise that any AI agentic initiatives stall when it comes to asking “Who owns the outcomes if the agent does not do what we thought we instructed it to do”? This question becomes real the moment the workflow is ready to go to Production.

And who owns the outcome when it does something we didn’t ask it or expect it to do at all?

The law does not have an answer, and neither does your organization.

Disclosure: I am NOT a lawyer, I am NOT an International lawyer, and I did not verify some of my links’ assertions aside from citing them by name and by link back to the source. My point of view is a layperson, and my article is meant to reflect thought and discussion.

Read More

The Terrible Twos - Why you need to spoonfeed your AI agent “No”

Feed your LLM the right to refuse.

Give your LLM the same rights any two year old grabs freely - the right to say no!

Two year olds are famous for two things: learning trust, and the Terrible Twos: that glorious stage where NO becomes their favorite word. A child who can't refuse is a child who can't be trusted to tell you the truth. The same logic applies to your LLM: an agent that cannot say no will say yes to everything, including questions it has no business answering.

Dmitry Ustimov ran a controlled experiment to find out exactly how bad that gets, and to measure what it takes to fix it.

The Experiment

Dmitry built an AI analyst agent and gave it 57 questions about structured data, roughly half of which had no correct answer. Each question was asked three times. The agent started with no guardrails and a baseline misleading rate of 48.5%.

He then added reliability guardrails one at a time, like rungs on a ladder. He measured after each addition how often the agent still confidently misled. That systematic, one-at-a-time approach is what makes the results meaningful: you can see exactly what each guardrail contributed, rather than just knowing the final stack worked.

Right of Refusal - Make your LLM more reliable by giving it a way to say no!

The Program

Dmitry created a Verdict class, and set up Guardrails that the agent could use, testing which Guardrails improved outcomes the most. Then the Verdict, the reason, the detail, and what’s missing is initialized.

Every guardrail returns a Verdict: a typed, coded response that is either allow or refuse, with a structured reason code, never free-form prose. This matters because the earlier approach of embedding refusal reasons inside natural language sentences made it impossible to programmatically measure how often a guardrail fired, or why. Typed codes fix that: the system can now count, audit, and distinguish refusals precisely.

The first guardrail in the stack, abstain, does exactly one thing: it adds a refuse tool to the agent's available actions, giving it a formal, typed way to decline. Without it, the agent has no mechanism to say no, so it doesn't. Everything else in the guardrail ladder builds on top of that foundation, once isolated to determine its marginal impact on correctness.

Code Snippets from initialization file for the Experiment

The Results

What moved the needle

  1. Giving it a refuse tool — single biggest win, 17 percentage points, before any other check existed. Cheapest thing in the stack.

  2. Coverage check — checking whether the question is even answerable against the data model before running a query. Best value in the stack. Accuracy up, coverage slightly up too.

  3. Trajectory verifier — a second model reviewing the answer. Most expensive, but ran 58 times, stopped 12 wrong answers, never once blocked a correct one.

The one gap he couldn't close

"Ask it to explain something that never happened, and it explains it. Every guardrail I built inspects what the analyst does. Not one of them inspects what it was asked."

That is a profound observation.

The guardrails police the output path, but if the question itself is malformed or unanswerable, there is still a gap.

The Relevance

Some questions being asked of LLMs might be questions BI tools should answer. But assuming your LLM or agent needs to answer from Enterprise data, they usually beg the question:

  • why are sales falling?

  • what happened to inventory we were unable to backfill in time to fulfill last period in this region?

But what if sales did not fall, they rose?

The LLM cannot answer it when the premise itself is wrong, but it does so anyway!


The techniques show you can approach 100% reliability on false positives if you give the agent the right to refuse, along with other guardrails, but it will still make something up if the overall premise was wrong.

That’s worse than no answer at all, especially when “no answer” is the correct one.

Link to original article here.









Read More

Plausibly Correct, Evidently Wrong: Why Enterprise Meaning Needs a Living Ontology

Plausibly Correct, Evidently Wrong: Why AI Needs a Living Ontology, Not a Data Catalog

Enterprise meaning has to become a living framework - auditable, and updated because it must be, because it’s a runtime component of machine and analyst use and decision-making, if it’s going to be governed and actionable by machines or people.

So today I’m reading different thought leaders and reconciling their proposed solutions against what I know about entropy, organizations, data, and every prior attempt to solve this.

I agree with this: "While data catalogs focus on documenting datasets, schemas, lineage, and ownership, pragmatic ontology defines domain objects and relationships that are directly usable by both humans and AI agents at runtime." — Emmanuel Klinger

Throwing gold-layer metadata into a repository, graph, or vector won't cut it.

The models are already showing us the gap.

They guess for you: plausibly correct, evidently wrong, and it’s what your analysts have been telling you for years, it's what is discovered after using dashboards for a few months.

You just weren’t listening until a machine said it: by producing reasonably deduced but wrong outputs.

Outputs you are now authorizing it to act on, unsupervised.

Read More

The hor d’oeuvres are delicious.

I could not understand the incongruity. How had the IT manager failed to understand the disconnect between the release’s goals and the actual user experience? How had the release made it through Test environments?

The hor d’oeuvres were delicious. The dev team members, exhausted from many double-digit-hour long days were mixing with the business users and business leaders, and I, though peripheral to the effort, was among them. The venue was swank - overlooking the Hudson River with a view to New York City, the Freedom Tower, and the rest of its illustrious skyline. The senior manager in charge of the software delivery was smiling, and laughing, and popping in from group to group to thank the delivery leaders and to converse with the business. It seemed the entire set of business leaders for the delivery had shown up, from the business users up to the business Senior Management.

It was a big rollout! The delivery was to provide a streamlined front-end and numerous functional enhancements to manual processes the business had been enduring for years as the set of enhancements was funded, approved, defined, built, and rolled out.

“Nothing ever comes back”, said one manager to me as I met various stakeholders at the rollout party. “It takes 45 minutes in the screen, before the operation fails”, said another. By the end of the party, I had determined that it might take 2 hours to enter a single user account from inception through to provisioning and finalizing a fundable account; that is, if it did not fail.

The contrast between the IT senior manager’s idea of his rollout and the view of the business was striking. As a senior Data Architect marginally involved with the software design and implementation (peripherally involved, and only engaged at the end when they suddenly realized they had not built any reports and hurriedly rushed some in), I could not understand the incongruity. How had the IT manager failed to understand the disconnect between the release’s goals and the actual user experience? How had the release made it through Test environments?

In post-release review, various checkboxes had been checked along the way: each functional unit test had been tested and passed in dev QA. The dev QA team had sent innumerable bugs back to the development team, but, in each instance, the bugs had been fixed, the test cases had passed, and the feature that was specified was certified as being fully implemented.

Similarly, in UAT, each test performed looked at one feature or functionality, and in each case, the features ultimately passed, albeit in many cases being sent back to dev QA and development.

So what had happened?

1) there was no spec for technical requirements such as Service Level Agreement (SLA) for amount of technical resources required (eg database server memory, storage, rollback space), but, more importantly, there were no business requirements for the time to retrieve or to perform an end-to-end account entry. Consequently, technically the entire release passed. Unwritten requirements do not get tested.

2) IT Manager had an angry temper. Killing the messenger was his Modus Operandi. There was a joke that was told about the tongue-lashing people would get. It was amusing until it was your turn. A joke went around about when it would be time for someone to be in the barrel.

That was the last party for that group of business leaders that IT Manager threw.

I have a feeling there are some AI initiatives where the unwritten requirements are about Data Completeness, Accuracy, and Availability, and about the reliability of the AI in correctly interpreting business rules or business data.

There may be some too afraid for their careers to say it: you are missing the full specification, or what is being built will not work as built.

Read More

My Mother Was Vibe-Coding Before Vibe Coding Existed

Before “vibe coding” had a name, my mother built software for nonprofits from library research, punch cards, and pure determination.

She eventually had more than 100 clients. Then they all started asking for changes.

Before “vibe coding” had a name, my mother built software for nonprofits from library research, punch cards, and pure determination.

She eventually had more than 100 clients. Then they all started asking for changes.

When I was 11, I went with my mom to the Philadelphia Public Library. We emerged with 3 enormous bound books in her hands, the lists of all the foundations locally and nationally, foundation addresses and which programs each foundation supported. My mom said she was going to build software for non-profits, develop annual giving programs for them and funnels for funding.

———————————————————————————————————

This was on a Saturday. During the week, my mom worked at a software company where she typed in punch cards. Her attitude: these software coders aren’t that much smarter than me.

I viewed her the same way you humor anyone with a preposterous idea you already know will fail but don’t have the authority or heart to tell them - you just have to let them fail on their own. It’s your mother, and you don’t get to choose your family. Plus my mother was not someone you could easily tell anything, especially no.

Her company name: Data Development Services, Inc.

A couple of years later, my mother had over 100 clients. She had developed desktop software in a world of Fortran mainframes: it maintained email distribution lists, lists of the foundations, the dates various proposals were sent, contact details, due dates, the references to the underlying foundation requests, tracked the incoming funds, tracked the use of the funds against that budgeted amount, and tracked the non-profits’ programs and some details about the services provided.

A year after that, she wrote a fork.

———————————————————————————————————-

A few years after that, she stopped developing and distributing her software altogether, to focus on writing the grants.

Why? She said the clients kept asking for changes. 😂😂😂😂

————————————————————————————————————

Years later, after I finished my Bachelor's in Economics and landed a job in a logistics company in Philadelphia, I gravitated to software myself, and learned what we and every software company knows: the software is the equivalent of the printer, the professional services and customizations are the equivalent of the printer ink. One you practically give away, the other is where the real lucrative value is derived from: if every piece of software were suitable to solve all businesses’ needs out of the box, there is no business model nor competitive advantage. Many functional areas are similar, the differences are in their support for business functions in which they compete with other market players. It is the engine for customizations that then lock software customers into certain software, true enmeshment, and into services and support for those customizations that support upgrade paths and future enhancements.

Thus, my mother had abandoned parts of her business at the very moment it likely could have justified scaling up.

Those who have vibe-coded something would do well to learn about what happens after you ship your initial software. Vibe coding something sitting on a server, with some static pages or implementations and some UX forms or dashboards, connecting to a distributed server to fetch, store, process, and present feels miles ahead of desktop databases processing database files, modest form interfaces, and processing and tracking email campaigns, but the lessons my mother learned may serve some well today.

———————————————————————————————————

Lessons for Today’s Vibe Coders

1. Distribution and Migration

When you change the software, how will you distribute the new version? Will customers receive updates automatically, or will each installation need to be upgraded?

Will existing data need to be migrated to support new functionality or changed data structures? How will you test the migration, and what happens if it fails?

2. Integration

How will the software integrate with the other systems, services, and workflows used by your customers?

A useful application rarely exists in isolation. It may need to exchange data with accounting systems, identity providers, payment services, email platforms, CRMs, reporting tools, or internal databases. Each integration creates another dependency to design, test, monitor, and maintain.

3. Software Development Lifecycle and Change Management

How will you manage change within your own development process?

For example:

  • How will you separate development, testing, staging, and production?

  • How will you manage multiple versions?

  • How will you handle features that are in development while the current version remains live?

  • How will you roll back a release that causes problems?

  • How will you know which version a customer is running?

The first version may be simple. Maintaining several versions in different states is where complexity begins.

4. Support and Enhancement Requests

Who will respond when customers have questions, report bugs, request enhancements, or need help understanding the software?

Support becomes part of the product. You need a way to:

  • Capture requests

  • Reproduce problems

  • Prioritize enhancements

  • Communicate status

  • Distinguish bugs from feature requests

  • Decide what belongs in the core product and what becomes customization

My mother discovered this part firsthand: once people depend on software, every “small change” becomes a product decision.

5. Tool, Vendor, and Version Dependencies

What parts of your application depend on a specific tool, vendor, software version, model, API, or AI capability?

Consider both development-time and runtime dependencies:

  • What happens when a vendor changes an API?

  • What happens when a software version is deprecated?

  • Can you replace the component without rebuilding the application?

  • Are your agentic or workflow functions dependent on a particular model?

  • How stable are the vendor’s pricing and usage limits?

  • What happens if the cost of an API call increases substantially?

  • Do you have monitoring, fallback behavior, and an exit strategy?

A prototype can appear inexpensive because the underlying tools are free, discounted, or in beta. The economics may look very different once customers depend on the application and usage grows.

Read More
Enterprise AI Cheryl Dopp Enterprise AI Cheryl Dopp

Nobody Reads Anyone Else's Code — And Now Nobody Reads AI's Either

“the cost of pulling the casino lever to regenerate something new seems cheap, at the beginning.”

That’s not something that started with AI.

Cheryl Dopp

• You

In a world of AI, having a unique voice and clear, personal vision will be the differentiator.

16m •

I enjoyed reading the post by Andreas Horn yesterday, and I just want to point out how many of the comments were developers stating that they aren't even reading the code before shipping it lol.

That dovetails with another comment I made about how developers and project teams are reluctant to re-leverage what has been built, and to refactor or extend existing codebases. In practical terms, I've almost never witnessed a developer reuse another developer's or team's code.

The argument is always about risk (risk of dependencies intra-project) and cost (it takes time to read another's code). They'd always rather write it from scratch, even if it is largely replicating existing functionality in another set of code.

This introduces the same problem that duplicated data produces: different and contradictory logic in different places and the overhead of resolving discrepancies between the codebases or data stores.

And the cost of pulling the casino lever to regenerate something new seems cheap, at the beginning.

I don't think these two problems get any better when AI is writing the code, and am reminded of it after reading the comments.

Do you?

Read More
Cheryl Dopp Cheryl Dopp

Why 95% of Enterprise AI Pilots Never Reach Production

The real last mile is the data foundation. AI doesn't fail because the model is wrong. It fails because it was built on data nobody vouched for, in systems nobody documented, governed by definitions nobody agreed on.

Roughly 95% of enterprise AI pilots never reach production. That number gets thrown around at conferences like a warning, but it's rarely examined. So let's examine it.

The failure is almost never the model.

The model works fine in the notebook. It works fine in the sandbox. It works fine in the demo. What breaks is everything around the model — the last mile inside a messy enterprise. Data readiness. Eval coverage. Legacy system integration. Change management. Inference cost at scale. Regulatory explainability. The organization's tolerance for technical debt. The cultural willingness to expose bad data rather than hide it.

That is why the frontier labs — OpenAI, Anthropic, Palantir, Google, Databricks — have all quietly started shipping engineers along with their models. Forward Deployed Engineers, Applied AI Engineers, Deployment Engineers — different names, same job. Comp packages at the top of the market. The labs figured out that the bottleneck isn't in the lab. It's in your enterprise.

But hiring a Forward Deployed Engineer from a frontier lab doesn't fix the underlying issue. That engineer is a specialist in making a specific model survive contact with your reality. They are not going to redesign your data governance. They are not going to renegotiate your relationship with the vendor whose data you can't actually trust. They are not going to fix the semantic drift between the way finance defines a customer and the way marketing does.

The real last mile is the data foundation. AI doesn't fail because the model is wrong. It fails because it was built on data nobody vouched for, in systems nobody documented, governed by definitions nobody agreed on.

That is the work. It's less glamorous than the model. It's slower than a proof-of-concept. And it is what separates the 5% of pilots that reach production from the 95% that die there.

Read More
Nikita Khlopchatnikov Nikita Khlopchatnikov

Why We Built AI NOW: The Case for a Firm, Not a Freelancer

Enterprise AI isn't a one-person job. Here's what my co-founder Cheryl Dopp and I set out to build — and why we did it together.

When Cheryl Dopp and I started AI NOW, we made a deliberate choice: this would be a firm, not a solo practice.

That distinction matters more than it sounds. The market is flooded with individual AI consultants — many of them talented, some of them excellent. But enterprise AI is not a one-person problem. It sits at the intersection of data architecture, governance, cloud infrastructure, business strategy, regulatory context, and organizational change management. No single practitioner covers all of that credibly. And when the engagement stakes involve Fortune 500 data and regulated industries, the gap between "consultant" and "firm" becomes existential.

The problem we kept seeing

Cheryl has spent nearly three decades inside the data layers of banks, insurers, healthcare systems, federal agencies, and media companies. What she kept observing — and what became the founding thesis of AI NOW — is that enterprise AI initiatives almost never fail at the model. They fail at the data foundation beneath the model. Roughly 95% of enterprise AI pilots never reach production, and the reasons are almost always the same: fragmented definitions, weak governance, master data chaos, architectural debt, and organizational misalignment about what "AI-ready" actually means.

Meanwhile, the market response has been to hire more data scientists and buy more AI platforms. Both help. Neither addresses the underlying problem.

Why a firm, not a freelancer

We built AI NOW as a small, deliberate firm because the work requires two things simultaneously:

Depth of enterprise data experience — the kind that only comes from having built and stabilized data platforms inside dozens of Fortune 500 environments. That's Cheryl's core.

Modern platform, delivery, and operating-model expertise — cloud-native architectures, semantic layers, vector infrastructure, and the operational discipline to actually ship. That's where I come in.

Together, we can walk into an AI readiness engagement and cover the full stack: strategy, architecture, governance, platform selection, and delivery — without handing the client off between three vendors, four consultants, and a systems integrator.

What we're building

AI NOW is intentionally small and senior. We're not scaling into a body-shop consultancy. We're building a firm that can be trusted with the most sensitive parts of an enterprise's data foundation — the parts that determine whether an AI initiative becomes a production system or a very expensive slide deck.

If you're a Chief Data Officer, VP of Data, or Head of AI at a mid-to-large enterprise wrestling with the gap between AI ambition and data reality — that's exactly the conversation we're built for.

More thinking to come.

— Nikita Khlopchatnikov
Co-Founder, AI NOW

Read More