AI
AI automation: where it works and where it does not
A meaningful share of tasks described as AI problems are better solved with a database query. Here is how to tell which is which before committing budget.
The current wave of AI capability is genuinely useful for a specific class of work. It is also being applied to problems where a rules engine would be cheaper, faster and correct every time.
The test that matters
Ask two questions about the task:
Is it language-shaped? Does it involve reading, writing, summarising, classifying or extracting from unstructured text? If the input is already structured data, you probably want a query, not a model.
Are the rules genuinely fuzzy? If you can write down the logic — "route to team B when amount exceeds 5000 and region is APAC" — write that down. It will be faster, free to run, and correct every time. Models earn their place when enumerating the rules is impractical.
Tasks passing both tests are good candidates. Tasks failing either are usually better served conventionally.
Where it works well
Classification with many categories
Routing incoming requests across thirty categories where the distinctions are contextual rather than keyword-based. Rules get unmanageable; a model with a good evaluation set handles it well.
Extraction from variable formats
Pulling fields from invoices, contracts and forms that arrive in hundreds of layouts. Traditional parsing needs a template per format. A model generalises — with confidence scores and a review queue for uncertain cases.
Question answering over a document corpus
Staff asking questions of policies, documentation and past tickets. Retrieval with citations means answers are grounded and verifiable, rather than depending on someone remembering which document covers what.
First-draft generation
Drafting replies, summaries and descriptions for a person to review and approve. The economics work because reviewing is much faster than writing from scratch.
Where it goes wrong
No evaluation set
The most common failure. Without a set of real examples with known-correct answers, "it seems better now" is your only feedback, and it is worthless for deciding whether to ship. Build the evaluation set *before* building the feature.
Automating a decision that should be reviewed
Models produce confident output whether or not it is correct. If a wrong answer has real consequences — money, legal exposure, safety — a person needs to be in the loop, or the confidence threshold needs to route uncertain cases to one.
Using a model where a query would do
"Show me customers who haven't ordered in 90 days" is a database query. Asking a model produces a slower, more expensive, occasionally wrong version of a solved problem.
Ignoring unit economics
Inference is priced per token. A workflow costing four cents per item is fine at a hundred items a day and a problem at fifty thousand. Model the cost during design.
What a sensible project looks like
- Pick one task with high volume and a measurable current cost.
- Build an evaluation set — a few hundred real examples with agreed-correct answers.
- Try the simplest approach — a good prompt against a strong model. Score it.
- Add complexity only where it scores better. Retrieval, then fine-tuning, each earning its place.
- Design the failure path. What happens when it is wrong, and who notices.
- Ship to a subset, measure against the manual baseline, expand if it holds.
The uncomfortable question
Before any of that, ask what happens if the system is wrong 5% of the time. For draft generation reviewed by a person, that is fine. For automatically issuing refunds, it is not.
If the honest answer is that a 5% error rate is unacceptable and no review step is possible, the task probably is not a good fit — regardless of how impressive the demo looked.
We go into the engineering side further on our AI development and generative AI pages.
MI Technologies Engineering
Software engineering team at Mohansh Innovations Pvt. Ltd.