IANA recently published An Intelligent Container Journey, a free guide on how AI applies to intermodal. It breaks AI into four types and follows a container from booking to empty return, showing what each type can do at every step.
I wanted to go deeper on the practical side, so I sat down with Chris Machut, founder and CEO of SiteTrax.io. His team built the guide with IANA and a group of industry experts over 18 months.
SiteTrax is an AI company itself. It sells camera systems that read container and chassis numbers at gates and yards, which is one of the four types below.
In this piece, I go deeper on what each type is, what it looks like in a real operation, and where to start.
Let's start with the four types.
1. The Interpreter: AI that reads
The first type is large language models: ChatGPT, Claude, Copilot.
They read human language and turn it into structured data. In intermodal that matters because so much of a move happens in language nobody organized. Bookings come in as email. Instructions sit in PDFs. Exceptions get explained over the phone and written up later in a spreadsheet.
The guide puts this to work at the start of the move: converting emails, EDI variants and PDFs into a clean order, and flagging any field it is unsure about for a person to check. It comes back later in the journey to draft delay notices for customers and to match proof of delivery against the bill of lading so the order closes without follow-up.
So when I asked Chris where an operator should start, I expected him to say cleaning up that paperwork. He suggested somewhere narrower. Start with customer communication, and with the conversations you are dreading.
"Get the $20 a month subscription and have that advisor. Help you think through what is your next step, be it with your boss, be it with your customers, be it with you internally, the processes," Chris said.
He calls this the metacognition use case, or thinking about thinking. It costs about $20 a month - a ChatGPT or Claude subscription.
A second use sits closer to operations, and he wrote about it in a LinkedIn post earlier this year. The idea is to point a model at operational data you do not have time to read yourself. The catch is that a vague request gets a vague answer, so the question has to be specific.
His example is a night crew that keeps getting called in on Fridays. Trucks are showing up at 11 pm for 2 pm appointments, and the reason is buried in hundreds of rows of gate logs. The request that works is specific: find every move outside standard hours, group them by carrier and time, estimate what the overtime is costing, and show which patterns are worth fixing.
This type carries the least risk because its mistakes stay in the chat window. A wrong draft costs a retry. Nothing is billed and no truck is dispatched.
The guide still tells readers to be careful with it. A warning near the end says AI "can hallucinate, misread context, and produce confident-sounding errors," and to validate against reliable sources before acting. Chris made the same point. "I think we have peaked at the LLM capabilities," he said. "They're going to make mistakes just like human beings do, but they do it even more confidently." He said the money is now going into the second type.
2. The Optimizer: AI that sees
The second type is cameras.
Chris called it applied AI, and said the industry also calls it physical AI. It is machine learning pointed at the physical world, reading container and chassis numbers at gates, yards and ramps. He described what his company does as taking the physical world and digitizing it.
It answers the questions an operator asks all day. When did that container enter the yard? When did it leave? Who dropped it off? Is this move heading toward dwell or detention charges?
The guide spreads this type across the whole journey. Computer vision guides crane moves in the yard to cut mis-spots and re-handles. Cameras photograph containers at outgate so pre-existing damage is on record before a dispute starts. Sensors on the train catch overheating wheel bearings, shifted loads and reefer failures while the container is still en route.
This is the category Chris and his team at SiteTrax.io work in, and by his count the accuracy has jumped in the past decade. General OCR used to read container numbers at about 70%, which put roughly three of every 10 assets into a system wrong. He puts modern AI-trained vision at 98% or better, including on numbers that are damaged, dirty or photographed at bad angles.
He spent most of his time on what happens when the system is unsure. Models he tested would report a read as 72.6% accurate, a number he said means nothing to the person holding the container. The same models would misread an O as a zero and report it with the same confidence as a correct read.
So SiteTrax took the scores out and replaced them with status codes. If the software cannot tell a zero from a one, it throws an exception and hands the case to a person.
Chris called exception handling the holy grail across all four types.
3. The Sentinel: AI that predicts
The third type is the one that looks ahead.
Analytics learn from historical and live data to forecast what is about to go wrong. In the guide, this type keeps updating a container's arrival time against weather, track congestion and crew availability while it rides the train. At the destination ramp, it forecasts chassis demand by type and timing before the train pulls in. On the back end, it projects where empty containers will pile up days before the imbalance appears.
Standard operational dashboards report what already happened. This one projects forward.
He described the work as continuous modeling and simulation rather than a single question and answer, in the same category of AI used to optimize routing.
It is also the type most exposed to everything upstream. A forecast learns from whatever the gates and clerks recorded. If the record is wrong, the forecast is wrong. The models also need history, which starts building only once the record is accurate.
4. The Conductor: AI that acts
The fourth type is the one that does something on its own.
An agent watches for a condition, applies your rules, takes an action and logs it. In the guide, agents match available drivers to loads using hours of service, proximity and appointment windows. They reschedule gate appointments as containers become available at the ramp. On the linehaul, they execute reroutes around a disruption, but only reroutes that were approved in advance.
The guide keeps repeating that framing: your rules, your guardrails, pre-approved moves.
A chatbot produces an answer. An agent produces an action. A wrong answer costs a retry. A wrong action becomes a transaction, and it moves through billing, inventory and customer records before anyone checks it.
Chris said the mistake operators make is scope.
"AI agents are very highly specific. They're not like LLMs, which are amazing generalists. AI agents do one thing and one thing. And that's where people mess up. They try to deploy AI agents and have them do more than they're actually good at."
Giving an agent a broad objective plus access to systems is not autonomy, he said. He calls it "ambiguity with API access."
Some of the people standing up AI agents for supply chain companies can demonstrate what the technology does but have no operating context behind it. He said that is why the deployments fail.
Chris built a multi-agent system for his own use and spent more than $4,000 on model usage in one month. He describes the result as barely working.
Outside research points the same way.
On τ-bench, a test that runs agents through customer service tasks with working tools and stated policies, the strongest models finished fewer than half the tasks. Given the same retail task eight times, they completed all eight correctly less than a quarter of the time. Gartner estimates about 130 of the thousands of vendors selling agents meet the definition. The firm expects more than 40% of agentic AI projects to be canceled by the end of 2027.
That failure rate compounds the longer a workflow runs: at 95% success per step, a 20-step workflow finishes correctly about a third of the time.
Chris said narrow and supervised agents do work. One example he uses: a trailer passes its dwell threshold by two hours. The agent flags the exception, notifies operations, opens a follow-up task and logs each step. A person decides what to do. His rule is that if a person has to approve something today, an agent should not do it tomorrow without a deliberate decision to change that.
The guide draws the same boundary. It describes a loop it calls monitor, assure, iterate: dashboards surface problems as signals, human checkpoints sit in front of high-impact decisions, and every exception becomes training data. Its summary of the split: automation executes, humans govern.
Where to start
Chris draws the framework as a Greek temple. Four columns hold up a roof marked AI. Two words are carved into the base: good data.
Freight moves in real time. The official record gets assembled afterward, out of emails, spreadsheets, phone calls and whatever everyone agrees probably happened.
I asked what intermodal's data looks like now.
"Maybe 70% accurate on a good day. More likely 50%. And then they have 8,000 different exception handling routines that Bob, Mary, and John each individually know how to handle."
He gave an example. In EDI, one shipper enters a container's weight in the dollar field and the dollar amount in the weight field. That is how their data entry works. A person at the ocean carrier knows this and corrects it by hand every time. The correction is not written down anywhere.
These models were trained on the open internet, an even messier dataset than intermodal's own records, and they still work. Chris's explanation is that training and inference are different. During training, errors average out because one wrong document among trillions gets outvoted. When a model runs on one operation, nothing averages out. The gate record for that container is the only record there is.
Another of his examples shows what one bad record does. A container marked CONT-4471 is dirty or badly lit and gets recorded as CONT-4477. The wrong number picks up a timestamp, a location and an accuracy score. Billing uses it. Inventory uses it. Customer service explains it weeks later.
I asked why data standards have not fixed this. "EDI is not a standard. It's a protocol more than anything," Chris said. Standards would work, he said, but adoption needs consequences. Freight has no penalty for bad data and no reward for good data.
His advice on sequence follows from that. Before picking a type of AI, he says to rate the operation on data quality, process standardization and guardrails, then start with whichever scores lowest.
What IANA is doing next
IANA is the trade association for intermodal. Its goal with AI is to get ahead of the technology instead of reacting to it: shared language and preferred practices in place before regulators write the rules.
The guide is Part One, and IANA is asking the industry what Part Two should cover. It also formed an AI working group under its Operations Committee in June.
The place to join that conversation is Intermodal EXPO, Sep 14 to 16 in Long Beach. The working group gives its update on day one, and Chris takes the four types on stage the next morning.
If you move freight on rail, or you are deciding whether to, it is worth being in that room. Register at intermodal.org.






