The Problem with AI Agents No One is Talking About | Yutori, Abhishek Das

EO12:20Added Aug 31, 2026

Most AI agents don't actually work. Abhishek Das, co-founder & co-CEO of Yutori, argues that the agent industry has quietly normalized unreliability. Even at...

Watch on YouTube →
Contributed by 刘嘉琪

Transcript

Transcript format
  1. Find the Dopamine of Building

  2. 00:00There's basically a hundred different agent products out there that are saying that like this can do anything on the web and you try it once and it doesn't really work. If we think of a 10step, 20
  3. 00:08step or 50step workflow, even if the accuracy at each step is like 90%, the 10% error rate compounds very quickly and so the overall success rate of a
  4. 00:17task of a workflow is like quite low. So that's one of the reasons why the technology is not there yet to do long horizon workflows. It feels like we have
  5. 00:26started normalizing and developed a tolerance for non-determinism and low reliability in shipping products. The product builder in me is like really
  6. 00:35annoyed that how is this okay for someone to ship an agentic product where they say that this can do anything but you try it the first time and like it doesn't work. I push back on that
  7. 00:44getting normalized especially with aentic products. If it's not good enough to work on the first try it's not good enough. The true differentiator is in.
  8. 00:53My name is Abhishek Das. I'm the co-founder and co-CEO of Ytorii. So with Ytori, we're building agents that can take actions and complete tasks on users
  9. 01:02behalf on the web so that you can focus on whatever is most meaningful to you.
  10. 01:05The three co-founders, we're all AI researchers by background. It is a bigger bet than just another agent company. So Faith Lee and Jeff and they were excited to support us.
  11. 01:25I come from a family of doctors and medical practitioners. So it was a bit of a an irony that I'm scared of blood.
  12. 01:30And so like the choice was pretty clear that like yeah I'm not going to pursue medicine. That's when I ended up deciding to to pursue engineering.
  13. IIT Roorkee and SDS Labs

  14. 01:37Science and the scientific process and method really appeals to me like the whole life cycle of coming up with hypothesis then designing experiments to
  15. 01:45validate or invalidate those hypothesis drawing conclusions and then coming up with the next set of hypothesis. I think that is a very neat sort of process and method. I found that really inspiring.
  16. 01:55And then I went to IIT Riy for my undergrad. That was a an amazing sort of learning experience. Met some of the smartest people I know. And so first
  17. 02:03year of college I was getting good grades but very quickly I realized that like electrical engineering especially this was more geared towards like power
  18. 02:10systems etc was not where my interest was and so at the end of first year that was sort of the first major sort of
  19. 02:17rebellious streak in me where I decided that okay like I don't see a future in electrical engineering I'm going to stop paying as much attention to it. So that was like a fairly sign significant fork
  20. 02:26in my in my life in some sense and instead I ended up spending a lot of my time learning programming and how to build software. That's when I got into
  21. 02:34building software and applications very seriously. What also helped was that IIIT Rurki had even at the time this is like almost 13 14 years back had a
  22. 02:44really strong programming club and culture. In particular, there was this group called SDS Labs, which was a group of like 10 15 coders from from every
  23. 02:52year who were just tinkering and hack like building a ton of applications for the internet for the rest of the campus.
  24. 02:59Seeing how users use it, building something from scratch and putting it out there and seeing how users interact with I think that was like a dopamine hit that kept like sort of fueling this.
  25. 03:09And I would stay up nights to to build this, to add features, to like improve it and and so on, right? Especially because I was surrounded by people who
  26. 03:17were as obsessed with this stuff as I was. And that was like extremely extremely motivating.
  27. The Last Generation to Use a Browser

  28. 03:26To be honest, I wanted to start something of my own for as long as I can remember. I had strongly considered it at the end of my undergrad, at the end
  29. 03:34of my PhD, and for various reasons didn't end up doing it. So, it was just a matter of time. And the main reason for it is that like there's lots of
  30. 03:41interesting problems in the world to solve to go after. Like I did want to push on that vision that I care about as opposed to working on somebody else's
  31. 03:50vision. Over the last two or three decades, web browsers by and large have stayed the same, right? Like we open a browser, we open a web page, we click around, scroll, type stuff, etc. There
  32. 03:58is an opportunity now to reimagine what that experience looks and feels like.
  33. 04:02We're going to be talking to our AI assistants that take actions and complete tasks on the web. And a lot of it is going to be agents that work in
  34. 04:11the background in a proactive manner for you. That's what the future looks like.
  35. 04:15And that is how we approached it. It felt like before physical agents become a reality um digital agents will become
  36. 04:22a reality like the timeline for digital agents is shorter than for physical agents. If you think about interacting with the web maybe like 5 to 10 years in
  37. 04:31the future, it is going to be at a slightly higher level of abstraction.
  38. 04:34Instead of us having to do digital chores ourselves manually, it lets us focus on tasks and stuff that's more
  39. 04:42meaningful, that's more interesting to us. Like if we can delegate all the mundane stuff to AI assistants, AI agents on our behalf, it lets us focus
  40. 04:49on stuff that's more more interesting to us. So it's more like humans and agents working together to overall improve productivity less so that like these
  41. 04:58agents are going to like replace humans and then humans won't have anything to do. But part of it is also just making it accessible to more people. Like my
  42. 05:05parents for example no longer have to learn every new website and how to operate it, right? Like if they can just tell an assistant that this is what I
  43. 05:12want to do on this particular website and it does it for them reliably, then that's awesome, right? So it makes it more accessible for more people.
  44. Stop Normalizing Broken Agents - What Separates Real Agents From Demos

  45. 05:23In this day and age, there's basically a 100 different agent products out there that are saying that like this can do anything on the web and you try it once and it doesn't really work. And there's
  46. 05:31also this notion of that like if you usually works right like if you try it 10 times then maybe like three times or like five times it it does the right
  47. 05:39thing I push back on that getting normalized. agents are basically making a sequence of decisions. Like if we think of a 10step, 20 step or 50step
  48. 05:48workflow, even if the accuracy at each step is like 90%, the 10% error rate compounds very quickly. And so the overall success rate of a task of a
  49. 05:56workflow is like quite low, right? And so that's one of the reasons why the technology is not there yet to do long horizon workflows. Being able to
  50. 06:04recognize when it makes mistakes and backtrack from that to then uh go down a different branch is really really important. We put in a lot of effort
  51. 06:14into building evals and guardrails. Like every single production query that a user runs goes through a fairly comprehensive set of evals that lets us
  52. 06:22quickly identify where these agents are doing well versus not, which domains need more work and so on. That's one aspect of it. And because we're in this
  53. 06:31space of web agents, right? Like agents that can do actions and tasks on the web. It will never be the case that we will be able to train on every single website that's out there. Like there's
  54. 06:40new websites coming up all the time. The number of websites that exist in the world is already pretty large. So we will always be training on a finite set
  55. 06:47of websites and improving these models there. Like people make mistakes on new website, click on the wrong buttons, etc. all the time, right? Like so it is very natural to expect models to also
  56. 06:55make mistakes. But when it makes a mistake, is it able to recognize and then backtrack and correct itself to do the right thing is a fairly important
  57. 07:03ingredient in the recipe of like how we train and build and ship these models.
  58. 07:08But the other part is is more ecosystemwide where like it feels like we have started normalizing and
  59. 07:15developed a tolerance for non-determinism and low reliability in shipping products. I don't like the normalization of slop and
  60. 07:23non-determinism and poor reliability especially with agentic products. Yeah, if it's not good enough to work on the first try, it's not good enough. We take
  61. The 80/20 rule

  62. 07:31sort of an 80/20 approach to it. Like there is always the prioritization question of like okay there are 100 features that we could be building. What are the top 10 that we need to focus on?
  63. 07:41Like some of those are informed by users and what what users are are telling us what they're asking for. But very often there are ways to build product that
  64. 07:49users may not be asking for. But if you built it and a lot of intuition goes into identifying what those features might be. Then users feel seen and they
  65. 07:58feel like oh this is someone who is listening to us. Even though that's not exactly what they asked for initially.
  66. 08:04I'll give you an example. The feature on iOS or Android that like anytime you get a two-factor authentication SMS, it auto
  67. 08:12reads your SMS and fills it into whichever app asked for it. It is hard to imagine like a user asking for that feature. But it saves a few seconds
  68. 08:22multiple times a day for people all across the world. But it's like a tiny thing that makes users feel seen like oh someone is actually giving thought how
  69. 08:29to reduce these tiny paper cuts in our in our day-to-day life. That's really important. So like it is a marriage of intuition with what users are actually
  70. 08:36asking for. In a world where it's very easy to come up with first prototypes using these coding LMS the true
  71. How to build taste - the weekly dogfooding ritual

  72. 08:44differentiator is in taste and craft in how intuitive and welldesigned the product is. One thing we do in the team
  73. 08:51that helps with that I think is we take uh dog fooding our own product very seriously like every single week we have
  74. 08:58an hour hour and a half docked out for dog fooding new features in the product at any given point of time we're running like tens of experiments internally and
  75. 09:07maybe like one of them will ship to the um production version of the product that external users will see. So like constantly dog fooding our our own
  76. 09:15product is a way to refine our own taste for like okay what is good versus bad what awesome or magical feels like. Like
  77. 09:22anything else um a lot of reps uh to build that muscle is like one way to go about it.
  78. Why Reliability Matters More Than Raw Performance

  79. 09:30The grad cam project was led by one of my labmates. I was sort of a supporting author on that paper. I was 25 when we did that paper. It's been extremely
  80. 09:38wellreceived. I think 20 30,000 citations is quite non-trivial at the time. Interpretability in like around deep learning models was like a big area
  81. 09:46of focus still is to this day. And so it was motivated from that that like okay like these models especially classification models to start with that
  82. 09:55go from like images to classifying it in one of thousand or 10,000 categories.
  83. 09:59What part of the image are they looking at to make those predictions? Right?
  84. 10:03There is clearly some signal coming from the image itself and then some signal that may be coming from the label that the classification model is predicting
  85. 10:12and how can we combine the two develop better intuition for what part of the image the model is looking at. To this day, it seems to work quite effectively
  86. 10:20across a bunch of tasks and models. Like with AI models, it is important for models to be able to convey not just the
  87. 10:27final prediction or the final answer, but also the proof of work. like what are the steps that went into coming up with this final prediction or the final
  88. 10:35answer. And so Grad Cam is like one manifestation of that. But even in how we build the the scouts product today, like you can set up these scouts and
  89. 10:44agents to monitor the web for something and they will generate these reports and notify you when they find something that's of value to you. But there is a
  90. 10:51button in the UI that lets you inspect the work that went in in behind the scenes like which websites were were visited, what did the agent actually
  91. 10:59look at to pull out this piece of information and that gives you a glimpse into the work that went in behind the scenes to put this together. It is very
  92. 11:07very important for trust building for users to be able to trust that yes this is a reliable product.
  93. 11:14A lot of our time and attention in how we're building our product at UTI goes into thinking about how should we build the product so that we don't make the
  94. 11:22same mistake. Whenever we ship something, we have to get it right. It has to really work. It has to be reliable. Users have to trust that it
  95. 11:29works well. If we put attention to detail into parts of the product that users can see, then the user is more
  96. 11:36likely to trust the parts of the product that they cannot see. Right? like everything awesome that we see around us, it's like individuals or groups who
  97. 11:45put in a lot of hard work and attention to detail to build that. So I think we should approach everything that we are building with that kind of philosophy.
  98. 11:54It takes time to build something meaningful, to build something right, to bring a vision of the future to life and building like delight delightful and
  99. 12:01reliable product experiences. It doesn't just appear out of nowhere.