top of page

Why AI Projects Fail at Mid-Market Companies: 6 Patterns and How to Pick a First Project That Works

Sep 28
7 min read

Short answer first. When AI projects fail at mid-market companies, it is rarely the technology. It is almost always which project got picked. Narrow your first project to work that is frequent, already has data, and can be checked by a human before it matters, and the failure rate drops noticeably. This post lays out that test, plus five real projects (names removed) where we applied it.


A sentence we hear a lot

"We built a chatbot last year. Nobody uses it."

Of the mid-market executives we met this year, maybe a third opened with some version of that. Surveys put AI project failure somewhere around 70–80% (the number swings a lot depending on how "failure" is defined, so "more than half miss expectations" is probably the safe reading). Our experience isn't far off. And the diagnosis is consistent across sources: it isn't a lack of technology. It's a lack of operating rules.

The trouble is that at a 30- to 300-person company, one failed project costs more than it does at an enterprise. Not in dollars. In conclusions. The first failure hardens into "AI doesn't work for us," and then two or three years go by.


The six ways it goes wrong

This list came out of a few years of proposals and kickoff meetings. It is ordered by how often we see each one.


1. Picking the tool before defining the problem. "Let's use ChatGPT at work" is a goal, not a project. Without naming which task, which bottleneck, and by how much, there is no yardstick when it comes time to measure.

2. Skipping the data cleanup. Quotes live in PDFs, spreadsheets, and screenshots from a messaging app, and every person calls the same thing something different. Put a model on top of that and what you get is "the AI is wrong," not a working system.

3. Leadership approves, then disappears. If nobody senior looks at the pilot results, a successful pilot doesn't spread. A CEO who looks at the actual output screen every two weeks matters more than most people expect.

4. No usage rules, so every team does its own thing. Without a one-page rule on who may put what data into which tool, shadow AI shows up, and one day customer data is sitting in an external service.

5. Making the first project too big. "Company-wide knowledge chatbot" is almost never a good first project. Once every department is involved, the effort dies in stakeholder alignment before it produces anything.

6. Choosing a partner by demo quality. Anyone can build a good demo. The real questions are whether they can show results on your data in two weeks, and whether they stay through operations after launch.


Three questions for picking a first AI project

This is the table we actually use in kickoff meetings. Score each candidate task 1–5 on each axis and start with the highest total.

Axis

The question we ask

What a 5 looks like

Frequency

How many times a day does this happen?

Dozens of times daily

Data availability

Can you pull the last six months of examples this week?

Straight out of the inbox or a system

Cost of being wrong

When the AI is wrong, who catches it and when?

A human reviews a draft before it goes out


The safest first project scores high on all three: frequent, data already exists, and a person checks the output. For a manufacturer, that is usually classifying order and quote emails and drafting replies. For an IT services firm, it is almost always first-pass triage of support tickets.


Five projects where we picked this way (names removed)

Examples travel faster than theory, so here are five engagements we delivered or scoped. What they share is that the first project was narrowed with the table above. What differs is the industry and the state of the data.


1. Outdoor equipment manufacturer A — from one newsletter to a chatbot

Under 50 people, one marketing person. The initial ask was "automate our marketing with AI," which, taken literally, walks straight into failure patterns 1 and 5.

So we made the first project very small: automatically collect competitor pricing and industry news from US, Korean, and European sources, summarize with an LLM, and send a weekly newsletter. There is no syndicated market data for this category, so this had been a manual job. It ran weekly, the data was public web, and a wrong summary just got skipped by the reader.

It landed well. On that trust, the roadmap moved to content automation (generating background and vehicle variants from existing product photos) and then to a RAG-based chatbot. One thing stuck with us: the CEO evaluated the AI work against the cost of hiring one more person. At a company this size, that seems to be the natural math.

Figure 1. Company A's three steps — newsletter automation → content generation → RAG chatbot
Figure 1. Company A's three steps — newsletter automation → content generation → RAG chatbot

2. Company A's US subsidiary — 66 findings, not all at once

The US side was a different situation. An on-site assessment produced 66 improvement items. The common mistake here is to put all 66 on the roadmap.

Instead, twelve weeks were split in three. The first four weeks went only to verifying how the commerce platform, ERP, support tool, and integration layer were actually wired together, and to measuring a KPI baseline. The next six weeks built only the high-frequency items: order processing, marketing workflows, chatbot integration. The last two weeks were handover and documentation. The remaining items weren't discarded; they went into a prioritized backlog.

The Korean headquarters followed a similar shape. A company-wide assessment turned up 96 items, and a six-month plan was built on three rules: don't replace the ERP, integrate with it; go department by department; hand off to internal operations at the end.

Figure 2. The 12-week plan — 66 findings narrowed to the high-frequency items
Figure 2. The 12-week plan — 66 findings narrowed to the high-frequency items

3. Security software company B — not "translate everything," just security terminology

General-purpose LLM translation kept getting security terms wrong, so the team was living with machine translation plus human post-editing (MTPE). Going for "automate all document translation" here is failure pattern 5.

The first project was a single fine-tuned translation model trained on the company's security glossary and years of past translation pairs. Translation requests came in daily, the pairs already existed, and every output was reviewed anyway. All three axes lined up. A nice side effect: corrections from review fed back into training data, so the model improved with use.

Figure 3. Reviewer corrections flow back into training data
Figure 3. Reviewer corrections flow back into training data

4. Beauty appliance manufacturer C — the database before the platform

A hair dryer manufacturer came in wanting an AI matching platform. The vision was good: a mentor–mentee network for roughly 350,000 hair professionals. But the platform had nothing to recommend yet.

So the first project became a structured content database of styling and hair-care content, with metadata, instead of the platform. Not glamorous. Skip it, though, and you land in failure pattern 2. The roadmap goes database, then matching platform, then B2B content partnerships.

Figure 4. A three-layer roadmap where the database comes first
Figure 4. A three-layer roadmap where the database comes first

5. Our own company — pulling numbers from five systems with agents

It would be unfair to only talk about other people's companies. Inside TecAce, revenue, cost, and cash flow data lived in five separate systems, and answering "how does cash look this quarter" took hours.

Rather than rip out the ERP, we attached agents for cash flow analysis, revenue forecasting, and what-if simulation, and brought them together in one dashboard. Data entry points dropped by more than 60%, and scenario analysis that took hours now takes seconds. We use it every day.

Figure 5. Five systems → integration layer → role-based agents → dashboard
Figure 5. Five systems → integration layer → role-based agents → dashboard

All five on one chart

Figure 6. The five projects plotted by data availability and cost of being wrong
Figure 6. The five projects plotted by data availability and cost of being wrong

What the five have in common: the first project was never the most impressive option. It was the one most likely to succeed first. Company C is the only one on the left, and that is the case where building the data was the first project.


A six-question check before you start

These are the six failure patterns, inverted. If any answer is "no," it may be too early.

  1. Is the success criterion written as a number (turnaround time, volume, error rate)?

  2. Can you pull six months of real data this week?

  3. Has a decision-maker committed to looking at results every two weeks?

  4. Is there a one-page rule on which data must never go into an AI tool?

  5. Do all stakeholders for the first project sit within a single team?

  6. Will the partner show results on your data and stay through operations?


Questions we get asked

Q. We have no AI staff. Can we still start?

Most of our clients are in that position. Once the project is narrowed with the table above, what the company needs to supply is data access and one person from the business side. The model and the build typically sit with the outside partner.


Q. What's the difference between a PoC and a real deployment?

A PoC answers "does it work." Deployment answers "do people use it daily." If the first project is a task someone already does every day, that gap mostly disappears.


Q. How long does a first project usually take?

In the examples above, four to twelve weeks. Longer than that is usually a sign the project is too big.


Q. We're a manufacturer and our data is scattered across the floor.

That is the most common starting point in manufacturing. In that case, do what Company C did: make collecting the data the first project. It is faster than starting with a chatbot.


How TecAce fits in

TecAce has been building software for global companies since 2000. The AX (AI Transformation) practice focuses on mid-market and growing companies. We are an Anthropic Claude partner with 12 Claude-certified architects, and we embed forward deployed engineers (FDEs) with clients from two hubs: Bellevue, WA and Seoul. The model is not remote advice. It is showing results on the client's own data, then staying to set up the operating rules after launch.

If you're weighing where to start, the enterprise AI transformation consulting page walks through how an engagement runs. If the more pressing problem is that nobody can find anything in your internal documents, the AI Knowledge Hub (AXKH) is usually the entry point.

Comments


bottom of page
AI Transformation
How Far Along Is Your AI Transformation?
Start your AI transformation
FREE