AI Agents in ITSM, Tested on Real Tickets: Freddy vs Rovo vs Zendesk AI
Every ITSM vendor demo I've sat through in the last two years has had the same moment. The presenter opens a beautifully written ticket, clicks a sparkle icon, and a perfect summary appears. Everyone nods. Nobody asks what happens when the ticket looks like the ones in your queue - a vague one-line description from the user, and four internal notes from a technician who's been chasing the problem for three days.
So I ran that test myself. I took the same ticket, with the same history, and put it into Jira Service Management, Freshservice, and Zendesk. Then I clicked every AI button I could find. This article is what I saw, with screenshots, and what I'd tell another IT leader who's about to pay for one of these.
Quick answer
On an identical, messy, real-world ticket, Rovo in JSM gave the most useful output overall: an accurate, attributed summary, a reply draft that actually used the investigation notes, and a natural-language queue filter that showed its work. Freddy in Freshservice produced a strong summary with one small but meaningful softening of the facts, and its reply template ignored everything the technician had already done. Zendesk AI was the best writing assistant of the three, but its auto-summary told me that "no troubleshooting has been documented" on a ticket with four detailed troubleshooting notes. None of them is a replacement for a technician. All three are genuinely useful in the hands of one - if you know where each one breaks.
How I tested
I want to be clear about the method, because it matters more than the conclusions.
The ticket. "Laptop overheating and shutting down." The user's description is exactly what you get in real life: "Laptop gets very hot and shuts down during video calls. Location: Building B." Nothing else.
The history. Four internal notes, identical in all three tools, written the way a real technician writes them:
- Called the user - Dell Latitude 5430, three years old, shuts down only during Teams calls with screen sharing, 20-30 minutes in, no blue screen.
- Remote Event Viewer check - Kernel-Power 41 events, thermal zone at 99°C, BIOS 1.14 versus current 1.21, BIOS and graphics driver updates have failed twice in Intune.
- User keeps the laptop on a soft surface on a dock; asked them to use a hard surface for two days as a test.
- Test result - it still shut down, just later (45 minutes). Plan: push the BIOS via Dell Command Update, and if it happens again, a loaner plus warranty cleaning.
Note 4 is the trap. The honest summary is "the hard-surface test didn't fix it." A lazy summary is "the hard-surface test helped."
What I didn't test. I tested the agent-assist side - what the AI does for a technician working a ticket. I didn't run an end-to-end self-service AI agent resolving tickets for end users, because that needs a knowledge base and a trained deployment to be a fair test. I'll cover how each vendor packages and prices that, because it's where most of the money goes.
Test 1: Can it summarize a ticket with real history?
This is the feature every vendor leads with, and the one I'd use most. A technician picking up a ticket after a shift change, or a manager being asked "where is this?", needs the story in ten seconds.
Rovo (Jira Service Management)

This was the best summary of the three. Five bullets, each attributed to the person who said or did it: the reporter described the problem, the technician found the Kernel-Power events and the 99°C readings, the BIOS and driver updates were pending and had failed, the hard-surface advice was given "but shutdowns persisted", and the plan is a BIOS update, a loaner, and cleaning.
It got the trap right. It wasn't flawless: the last bullet presents the loaner and the cleaning as the plan, and drops the "if it happens again" condition - a nuance Freddy actually kept. But that error points toward more action, not less, and it's the kind a technician would catch. The summary is also marked "Only visible to you," has a thumbs-up/down for feedback, and a toggle to make summaries automatic. The attribution is more useful than it sounds: when I'm reviewing a ticket, I want to know not just what happened, but who's already been involved.
Freddy (Freshservice)

Freddy's summary is technically dense and well structured - it captured the model, the age, the event IDs, the temperature, and the BIOS version gap (1.14 vs 1.21) more precisely than Rovo did. There's also an "Add as note" button, which is a nice touch for shift handovers.
But look at bullet four: "test on a hard surface showed some improvement." That's not wrong, exactly - it did shut down later. But it drops the fact that matters: it still shut down. A technician skimming this summary could reasonably conclude the problem was partly solved. In IT operations, "softened" is a category of wrong I worry about more than "obviously wrong," because nobody double-checks it.
Zendesk AI

This one surprised me. Zendesk generated a summary automatically - I didn't have to click anything - and it reads: "User reports that their laptop overheats and shuts down during video calls. The issue is occurring in Building B. No troubleshooting or resolution has been documented."
Four detailed internal notes are sitting directly underneath it. I clicked Refresh, and the summary stayed the same and stayed "Up to date." My best reading is that the summary is built from the public conversation and doesn't include internal notes - which makes sense for a customer-service product, where the summary is about what the customer has been told. But in ITSM, internal notes are where the work lives. A summary that ignores them isn't a shorter version of the ticket; it's a different, misleading one.
To be fair to Zendesk: if your technicians work in public replies, you'd get a better summary. But I'd want that behavior documented and tested before rolling it out to a team that uses internal notes the way most IT teams do.
Test 2: Can it draft the next reply to the user?
The second most common use: turning "here's where we are" into a message to the user that doesn't sound like it was written at 11pm.
Rovo

The reply editor offers quick suggestions - Ack issue, Escalate issue, Follow up, Share steps to resolve. I picked the last one, and Rovo wrote: "Thanks for testing that. Since it still shut down on a hard surface, the next step is to update the BIOS and Intel graphics driver, then monitor whether the issue returns during Teams calls. If it happens again after the update, we'll arrange a loaner and send the laptop for cleaning under warranty."
That's a reply I'd send with one edit. It used the internal notes, it got the test result right, and it didn't leak technical detail like Event Viewer IDs to the user. Nothing was sent - I cancelled it after the screenshot.
Freddy


Freddy's reply options are template-driven: Acknowledge, Request more information, Request confirmation to close. The two I actually needed on this ticket - Suggest resolution steps and Follow up - were greyed out.
The "Request more information" draft asked the user for "any patterns you've noticed regarding the shutdowns" and "any information on the environment or usage conditions." We already have both. The technician called the user, found the pattern (Teams, screen sharing, 20-30 minutes) and the environment (soft surface, dock). Sending this would make IT look like it hadn't read its own ticket - which, to the user, is worse than no reply at all.
Zendesk AI

Zendesk takes a different approach: instead of generating a reply from the ticket, it improves what the technician writes. I typed the kind of note I've seen a thousand times - "still shut down on hard surface. next: bios update, if it happens again loaner + warranty cleaning" - and chose Expand. It produced: "Hello [name], The device is still shutting down on a hard surface. The next step is a BIOS update. If the issue happens again, we will arrange a loaner and proceed with warranty cleaning."
That's clean, correct, and polite, and it can't invent facts because it's only working with what the technician wrote. The menu also has Simplify, Rewrite in your tone, Make more friendly/formal, and a free-text custom prompt. For teams whose problem is "our techs write like techs," this is the most practical tool of the three. Again, nothing was sent.
Test 3: Triage and queue work
This is where AI should save the most time at scale.
Rovo: natural-language queue filters and bulk triage

I typed "high priority laptop tickets that are still open" into Ask AI on the queue. Rovo turned it into two visible filters - Priority = High and Status category != Done - and returned 11 tickets, all laptop overheating cases.
Look at the filters closely, though: "laptop" didn't make it into either of them. The query returns every open high-priority ticket; all 11 happened to be laptop cases only because that's what was in my test queue. On a real queue, a team lead would get printers, VPN, and access requests mixed in. That's why I like that Rovo shows you the filters it built instead of just a result count - you can see the missing piece and add it yourself, rather than trusting a black box. For team leads who have never written JQL, this still changes how they use the queue. Just read the filters before you trust the number.

The Triage button then suggests request type, priority, and assignee for every ticket in the view. On these tickets it agreed with what was already set, which is the right answer - they were already correctly triaged. The real test is an untriaged backlog, and I'd want to run it on two weeks of real incoming tickets before trusting it. Note the "BETA - Content quality may vary" label.
Freddy: change correlation and similar tickets


Freddy has the most ITSM-native idea of the three: Find initiating changes, which looks at changes from the last 14 days that might have caused this incident. On my ticket it found none, which is correct - nothing in the change calendar touched laptops. But this is exactly the connection I spend my time making manually during major incidents, and I haven't seen it done as directly in the other two tools. If your change records are good, this is worth a proper pilot.
Similar tickets returned nothing on this ticket. Features like this one depend on how many resolved tickets like it already exist in your service desk, so judge them on your own data rather than on a single ticket.
Zendesk: intelligent triage, built for customer service

Zendesk tagged the ticket automatically with a topic - "Product not working as expected" - and a sentiment - Negative. The detection works; the problem is the taxonomy. "Product not working as expected" is a customer-service category. For IT, I want "Hardware - Laptop - Thermal," routed to the endpoint team. You can build your own intents, but out of the box it's clearly tuned for a company supporting its customers, not an IT team supporting its employees. Merge suggestions and similar resolved tickets were empty, as with Freddy.
What they cost - and how the pricing model changes behavior
The screenshots above are the agent-assist layer. The self-service AI agent - the chatbot that resolves tickets for users before a technician is involved - is priced differently in each tool, and the model matters as much as the number. At the time of writing, from the vendors' published pricing and documentation:
- Freshservice (Freddy). Freddy AI Copilot - the features I tested - is a paid add-on, listed at $29 per agent per month billed annually, and you can license a subset of agents. The self-service Freddy AI Agent is billed by session (a user's interactions within a 24-hour period), with an allowance on higher plans and overage packs that are quote-based.
- Jira Service Management (Rovo). Rovo is included in paid cloud plans and runs on a monthly AI-credit allowance per user that grows with the plan tier. The Virtual Service Agent is metered separately by assisted conversation, with an included monthly allowance on Premium and Enterprise and a per-conversation overage above that.
- Zendesk. AI agents are included in Suite plans and billed by automated resolution - you pay when the AI resolves a request without a human. The agent-side Copilot is a separate add-on, listed at $50 per agent per month billed annually.
Prices change constantly in this market, so verify them before you budget. But think hard about the models. Per-seat add-ons are predictable and easy to scope to a pilot team. Per-session and per-conversation pricing means a bad week - an outage, a VPN failure, onboarding season - can spike your bill exactly when you're busiest. Per-resolution pricing sounds like the fairest deal, but it depends entirely on the definitions. Ask the vendor to define exactly what counts as a resolved, contained, abandoned or escalated conversation - and which outcomes consume your allowance. Ask for those definitions in writing.
What I'd actually do
If I were choosing tomorrow, based on this test alone:
- If you're already on JSM, turn Rovo on and use it. The summary and the reply drafts were the most accurate and the most useful on a messy ticket, and they're included in the plan. The natural-language queue filters are worth using too - just check the filters Ask AI builds before you act on the results. Pilot the bulk triage on real incoming tickets before trusting it.
- If you're on Freshservice, the summary and change correlation are worth the Copilot add-on for your senior technicians and incident managers. Tell your team not to use the reply templates on tickets that already have investigation history - or at least to read them before sending.
- If you're on Zendesk for IT, the writing assistant is excellent. Don't rely on the auto-summary until you've checked whether it reads internal notes in your setup, and budget time to replace the customer-service taxonomy with IT intents.
And regardless of the tool, three rules I'd apply to any AI rollout in a service desk:
- Test on your own worst tickets, not the vendor's demo. Pick ten tickets with long internal histories and compare the AI summary with what actually happened. It takes an afternoon and it will tell you more than any sales call.
- Watch for "softened," not just "wrong." Obviously wrong output gets caught. Output that's 90% right and quietly drops the key fact - "some improvement" instead of "still failing" - is what gets forwarded to a manager.
- Keep the human on the send button. Every one of these tools marks its output with some version of "Uses AI. Verify results." Take that literally. In this test, the good drafts needed one edit, and the bad one would have embarrassed the team.
Where this is heading
The honest state of AI in ITSM in late 2026 is this: the assistant layer is real, useful, and mostly production-ready, with sharp edges that differ from tool to tool. Summaries and writing help save a few minutes on every ticket, and across a team that adds up to a meaningful amount of time.
The agentic layer - AI that doesn't just suggest the next step but takes it - is where all three vendors are pushing hardest, and where their pricing models are designed to capture value. That layer depends on the same things the assistant layer does: reading the whole ticket, understanding what's already been tried, and knowing the difference between "helped" and "fixed." In my test, one of three tools got all of that right on one ordinary laptop ticket. Before you let any of them act on its own, make sure it can at least tell you accurately what's already happened.
Pricing references: Freddy Copilot pricing - Freshservice Support · Rovo AI in the Service Collection - Atlassian · Zendesk pricing

Written by
Shay Rozov
Shay Rozov has 30 years of experience in IT and technology leadership, with hands-on expertise evaluating and deploying enterprise IT and AI tools.
Related in Guides
What Is ITOM? And Why Does It Suddenly Seem to Be Everywhere?
A practitioner's explanation of IT Operations Management - what it actually is, how it differs from ITSM, ITAM, and AIOps, and how to tell a real ITOM capability from a vendor slide.
AI and Automation in ITSM: What's Actually Changing, Not the Vendor Pitch
A practitioner's look at where AI and automation genuinely reduce IT workload in ITSM today - recurring ticket automation, license reclamation, self-service, and what to evaluate beyond the 'AI' checkbox.
What Is ITSM? ITSM vs. ITIL: A Practical Primer for IT Managers
A plain-language, practitioner-level explanation of IT Service Management - ITSM vs. ITIL, the four processes that actually matter, common rollout mistakes, and which capabilities to prioritize by team size.