AI & data

Software for data that first has to be understood.

Some processes are not hard because software is missing. They are hard because the information is not structured in the first place: e‑mails every sender writes differently, ten years of PDFs without labels, supplier documents that keep changing, data whose meaning only becomes clear from context. We build systems that make sense of that information. Conventional software, statistics, machine learning and AI, each where it makes sense.

Fixed price before we start · effort estimate before any commitment
Three levels

Clear rules we automate. Unclear data we make understandable first.

Whether a process needs conventional software, a learning system or no software at all is decided by the process, not by the technology. We check that before the first quote.

01

The process is clear

IfThe rules can be written down: if X, then Y. A person still repeats them by hand every day.

ThenConventional software. Integration, automation, line-of-business application. Traceable, testable, no model involved.

What we build
02

The process is messy, but a person understands it

IfNobody can write the rule down. Yet the clerk knows at a glance which purchase order the e‑mail belongs to and what the PDF says.

ThenThis is where we use data science, machine learning and AI. The system learns what the person understands, and conventional software takes care of the rest.

03

The process itself is the problem

IfIt is not the data that is messy but the process behind it: duplicate paths, unclear responsibilities, three systems for one task.

ThenThen we say so. Automating nonsense only makes it faster. Sometimes the answer is a changed process and no software.

How we work
What we build with it

Five tasks where until now a person had to read.

Not AI features, but steps of work that only a person could do so far, because they first had to understand what was there. Each line is a type of project we deliver this way.

01

Understand documents

TodayEvery supplier, customer or partner sends the same document differently: PDF, scan, Excel, different order, own abbreviations, missing fields.

With JMKA system that brings all variants into one shared structure and learns new variants itself. Order confirmations, delivery notes, quotes, invoices.

02

Connect unsorted mailboxes

TodayA shared mailbox with hundreds of e‑mails in which orders, queries, complaints and advertising are mixed together. Sender and subject say little.

With JMKThe system recognises what an e‑mail is, which order or customer it belongs to, who is responsible and what should happen next. The mailbox stays as it is.

03

Make historical data usable

TodayTen years of documents, exports and folders, unlabelled, in changing formats. The knowledge is there but cannot be queried.

With JMKWe group, classify and structure the material without anyone labelling everything by hand first. The result is a data basis software can work with.

04

Detect changes

TodayPrice lists, data sheets, tenders, regulations or supplier pages change. A keyword search returns either everything or nothing.

With JMKA service that keeps reading sources, compares them with the previous state and only reports what is relevant to your business. With a reason and the evidence.

05

Learn patterns instead of maintaining exceptions

TodayEvery new supplier, every new layout and every new spelling currently needs another rule that someone has to add.

With JMKA system that learns from the existing cases and classifies new variants itself. At the door manufacturer new suppliers are added without JMK having to step in.

Businesses have more structured knowledge than they realise. It sits in product codes, document layouts, naming conventions, old orders and years of repeated decisions. Using statistical methods, machine learning and data analysis we bring those structures out and turn them into explicit, usable data.

Knowledge from historical data

What your business knows is often stored in no database.

Extraction

You know what you are looking for.

“Find order number, customer, price and delivery date.” The target is known, the system has to find it in every document.

Discovery

You do not know the structure yet yourself.

“Work out which characteristics these product codes describe and how that differs between suppliers.” The system asks the data which rules produced it.

Example: product codes at the door manufacturer

VS Material
B Climate class
30 Fire rating
··· further characteristics
Schematic. The characteristics found are real, the code is an example.

About ten years of purchase orders and confirmations contained supplier product codes whose structure was documented nowhere. An experienced clerk reads material, climate class and fire rating from them. The system reconstructed that structure from the correlations in the data: which positions of the code belong together and what they stand for.

And differently for each supplier. One distinguishes tubular chipboard from solid chipboard, another does not carry tubular chipboard at all. A third encodes the same property somewhere else. From these vocabularies a shared internal representation emerges, and only that makes the comparison possible at all.

How that is built at the door manufacturer
From history to automation
  1. Archive ten years, unlabelled
  2. statistics
  3. Structure discovered
  4. model
  5. Knowledge made explicit
  6. rules
  7. Comparison runs automatically
Why this is not an API call

The result should feel simple. The problem was not.

In the end the clerk sees one e‑mail: confirmation checked, one mismatch. Behind it lies a chain of steps, and a different method for each step. Not every problem needs a language model. The criterion is not what is fashionable but what runs reliably and economically.

01

Deterministic rules

Where the layout is stable and the rule is known. Fast, verifiable, no surprises. Always the first candidate.

02

Statistics and clustering

To find structure where nothing is labelled: which documents belong together, which positions of a code are related.

03

Classical machine learning

For classification and matching with many examples: what kind of e‑mail is this, which supplier, which variant.

04

Small, specialised models

Trained or adapted by us for exactly one task. Small enough for your own infrastructure, fast enough for continuous operation.

05

Large language models

Where language has to be understood and examples are missing. One component of several, with checks behind it, never the whole solution.

06

A person decides

The system prepares, the clerk approves. Errors remain hints that someone sees and do not become orders that go into production wrong.

We do not default to the largest model.

For some tasks a call to a large provider is the right solution. For others we train small, specialised models, combine several methods or replace learned components entirely with fixed logic. That keeps systems faster, cheaper and easier to operate. And small models can run on your own infrastructure when the data must not leave the building.

That we can build such models is on record: with a language-identification system of their own that assigns texts to more than 2,000 languages, where common models stop at around 200, Tobias Madlberger and his team took second place at the Austrian federal AI competition in 2025. Organiser’s review

The full chain in the door-manufacturer case
Frequently asked

What we get asked about AI first.

Isn’t this just ChatGPT with a PDF?

For a single question about a single document that is often enough. For a process that runs every day without supervision it is not. That needs matching, checks against fixed rules, handling of new variants and a traceable result. In our systems a language model is one component of several, often not the most important one.

Do we need clean, labelled data for this?

No, that is the point. At the door manufacturer there were about ten years of material without a single label. The system derived the structure from the data itself. What you need is access to the existing material and two or three people who can read it.

Does our data have to go to an AI provider?

Not necessarily. Which steps need a large model and which a small one that can run locally, we decide per task and discuss before we start. Where data must not leave the building, we plan for that from the beginning.

What happens when the system gets it wrong?

The system does not decide, it prepares. At the door manufacturer the result goes to the clerk as an e‑mail, and the clerk gives the approval. Extracted values are checked against fixed rules beforehand. So an error remains a hint that a person sees, not an order that goes into production wrong.

Do we have to change our processes for this?

Usually not, and that is deliberate. At the door manufacturer it would have been cleaner to reorganise supplier communication. Instead the software learns the existing shared mailbox. Only when the process itself is the problem do we recommend changing it instead of automating it.

What does a project like this cost?

Like everything with us: a fixed price for the agreed scope, with an effort estimate beforehand. The door-manufacturer case was a fixed-price project of about two months. Guide prices for the other types of project are on the page about how we work.

Call us Discuss a project