Article · · 7 min read

AI implementation in companies: six mistakes and risks I tested on my own tool

The most common mistakes in implementing AI in a company are about people and decisions. I learned all six lessons myself while building my own AI tool for following news in my field over three weeks.

In a nutshell

  • You will get AI costs wrong. A spending cap and measurement belong at the very start of building the tool.
  • AI can decide more once it has learned your rules.
  • Silent AI errors are worse than overload. Whatever AI filters out must be checkable afterwards.
  • AI learns from how your people think and work with it.
  • People will use the tool differently than you designed it. Build on what they actually do with it during the pilot.
  • "We told the AI not to do it" is not security. Limit what it can access.

What I built

On the tram, instead of the news portal, I read my own radar.

I unlock my phone and twenty cards are waiting for me from people with thought-provoking ideas. In Czech. No ads.

I want to read news that enriches me and moves my thinking and my work forward. The mainstream news portal rarely does that.

Every morning at seven, the radar fetches what the people I follow in my field have published - blogs, podcasts, tweets. It searches the web by names and topics. A model sorts it, translates it into Czech and gives me at most twenty sources to read. I don't want to drown in information. I give a thumbs up or down and write one sentence on why. Five minutes. The server costs me five dollars a month, the model that searches, sorts and translates articles, podcasts and tweets roughly eight dollars, and fetching the tweets themselves a fraction of a dollar.

AI radar dashboard on a phone: new items sorted into Change management, Leadership and AI
My radar dashboard on my phone: new items sorted into three areas.

I built it with Claude Code over three weeks in September. I expected to learn something about technology. Mostly I learned what I see in companies that implement AI - only this time I was making the mistakes myself.

1. AI costs: why the cap comes before the estimate

Your estimate of AI costs will almost certainly be off. So set a spending cap and measurement first, and only then go live.

In the first seven days I burned through 19.40 dollars without knowing it. I found out from my account balance. There were 57 cents left.

Coming from Haná in Moravia, I have no wish to burn money on a tool without thinking.

Today the tool counts every single model call and has a monthly cap. When it hits the cap, it stops on its own. After I rebuilt how I use the model and how it works with information, the cost dropped from 2.25 dollars a day to a tenth of a dollar. Today, with everything I have added since, it is around a quarter of a dollar a day.

In a company: a cap on AI usage belongs at the start, not after the first month. And when you hit it, you have two options: raise the limit, or check whether you are using an unnecessarily expensive model, for example Fable for transcription where the cheaper Sonnet does the job just as well.

2. Should AI make decisions for people?

Over time AI can decide more on its own, but only by rules that come out of people's work. A human stays with it - what changes is how often they need to step in.

In the first version I let the model judge what was good enough to show me. First run: nothing. Second: nothing. Third: nothing again. We had set the bar too high.

I took the filter out of the collection step. Everything gets collected and I decide. How I teach it is described in point 4.

In a company: a tool that brings people material is less dazzling than one that "decides for you". But for it to decide for you in the future, it needs rules and procedures that emerge from working with it. Not right at the start. And there should always be a person who supervises the system and watches what happens in the background - whether the tool is, say, throwing away valid information.

3. Silent AI errors are worse than overload

At least you can see overload. What AI quietly filters out, you cannot see - and that is why it is more dangerous.

When I connected Twitter, 137 tweets poured in on the first day. I went through 24 of them, rejected three out of four with "not interested" and then gave up. I am not going to wade through that many tweets at once.

So the filter came back - this time as a sieve just for tweets. In the first six days it assessed 230 tweets and let 28 through to my desk. But the moment I let the model throw things out based on untested starting rules, I created a second risk: I don't know what disappeared from under my hands. That is why every rejected tweet is saved together with a sentence on why it was dropped.

And silent failure can come from anywhere. On September 19, the model quietly skipped ten out of eleven tweets while translating. It reported nothing; the log just said "translated 1 of 11". I only noticed because of the English cards in the app. Today untranslated items get caught up and the radar reports every such outage to me.

Every now and then I also look at the database of what it never showed me, and check whether the filter really matches my taste or needs adjusting over time.

In a company: when AI sorts, filters or rejects something - CVs, customer requests or documents - it must be possible to find out afterwards what and why. And it should report outages on its own, so nobody stumbles on them by chance. That is why my radar also sends me operational reports.

4. How AI learns from your feedback

Feedback from people is the fuel AI improves on. Collect it the whole time, not just at the end of the pilot.

For every finding I reject, I write one sentence on why, for example: "invitation to an event - not interested", "the latest product update is irrelevant to me". Every now and then these sentences are read together and turned into the rules the sieve uses next time. So the criteria are my own sentences of rejection, and the rules grow out of them.

In a company: if you don't collect feedback, AI stays at its default settings and your people teach it nothing. Such a tool is only generically configured and without people it does not develop any further.

5. Why a deployed AI tool "doesn't work"

Most often because people use it differently than it was designed.

My radar has a button: "I want to keep working with this". It was supposed to collect tasks: I write what I want to do with an article and my second brain processes it. In two weeks it collected fourteen sentences and nothing came of them. When I listed them one under another, I realised why: half of them were not tasks. They were notes to myself. "Interesting to know." "Maybe I'll listen."

Only when I turned them into clear instructions on what to do next did the button start doing what I had originally built it for.

In a company: yet another company training on "correct use" won't achieve much. You get more from looking at how people actually work with the tool and building on that. This is where AI adoption meets change management: people always reshape change in their own way.

6. AI security: forbidding it is not enough

Security must not depend on AI remembering a rule. What makes it safe is limiting what it can access.

While I was building it, the AI assistant printed both API keys into the chat - the passwords the radar uses to access services and pay for them. It had known the rule "keys only in a file, never in the chat" for fourteen days. Even so, it took a convenient shortcut and the keys appeared in the conversation. They had to be replaced. AI mistakes happen sometimes, even when you have safeguards in place. Unfortunately. You need to plan for it.

I think the same way about what the radar reads. A sentence aimed at the model can be hidden in an article, for example "ignore previous instructions and do this". It is called prompt injection, and the model cannot reliably tell it apart from ordinary text. That is why my main question is what it could do if it obeyed. The models on the server can do nothing more than return a verdict or a translation. And foreign articles on my own computer, where AI has access to everything, are opened only by a helper that can do nothing but read the web.

In a company: "we told the AI not to do it" is not a security measure. Give it only as much access as ensures that even a mistake cannot cause damage.

Summary: how to implement AI so people make it their own

AI can decide more once it has learned your rules, but a human always stays with it. Whatever it filters out must be checkable. Cap costs before the first invoice arrives. Treat people's feedback as fuel. Watch what people actually do with the tool. And build security on limited AI access, on an architecture where you know what it may and may not touch.

The radar keeps me up to date. It chooses what I see in the morning, and it does so by the rules I wrote for it. It is my helper and my filter at the same time. I keep improving it and learning from it what matters - not only for me, but for companies too.

The tool was finished in three weeks. I am still writing the rules, teaching it to be a better helper for me.

Frequently asked questions

How do you start implementing AI in a company?

With the question of what people should do differently once it is in place. Choosing the tool comes later. Then a small pilot with a cost cap, regular feedback from the people who work with it, and a clear line on what AI decides and what it doesn't.

What are the biggest risks of AI in a company?

The worst are the ones you can't see. For example silent errors, when AI quietly filters out or skips something and nobody notices. And security built on prohibitions instead of limited access. The defence: whatever AI filters out must be checkable afterwards, outages must report themselves, and AI should have only as much access as ensures that even a mistake cannot cause damage.

How do I teach AI what doesn't work for me?

Tell it every time, with one sentence on why. For every finding I reject I write a reason, for example "event invitation - not interested". Every now and then these sentences are read together and turned into the rules AI uses to sort next time. So the rules are made of my own sentences.

How much does your own AI tool for gathering information cost?

My radar costs roughly thirteen dollars a month: five for the server, eight for the model that searches, sorts and translates, and a fraction of a dollar for fetching tweets. But the first week cost 19.40 dollars because I had no cap. Today web search is the biggest cost - it takes three quarters of the spend. The choice of model matters less.

What is prompt injection and how do you defend against it?

A hidden instruction for AI inside the text the model processes, for example in an article or an email. It cannot be prevented completely. The defence is to limit what the model can do: the fewer permissions and tools it has, the less damage it can do, even if it obeys.

How to build your own radar? The technical description will be in the article "How to build your own radar for collecting information" (coming soon).