TLCNLP logo

28 March 2024

Episode 15: The role of AI in transfer pricing and benchmarking, with Borys Ulanenko

Paul Sutton and Borys Ulanenko discuss the role of AI in benchmarking for transfer pricing, including the practical problems that AI can address, how AI can help TP professionals to demonstrate that their benchmark analyses are robust, and some common misconceptions around the use of AI.

Borys is the founder of ArmsLength.AI. This platform introduces AI solutions to complex tax challenges, streamlining data analysis and enhancing decision-making accuracy. Before ArmsLength.AI, he worked at Aibidia, focusing on digital solutions in the same field.

Paul and Borys discuss:

  • The problems inherent in the benchmarking process
  • The challenges that TP professionals face when creating benchmarking analyses
  • The role of AI in addressing these issues
  • Real-world examples of how AI can make benchmarking analyses that are more robust, and demonstrably so
  • Examples of what ArmsLength.AI does for its clients, and its pricing models
  • Common misconceptions around the use of AI, and what people should consider when assessing any AI tool on the market.

Transcript

The following transcript has been lightly edited for clarity. Borys Ulanenko can be contacted via his LinkedIn page .

 

Intro: Hello and welcome to The LCN Legal Podcast, bringing you expert views and analysis of the legal aspects of transfer pricing compliance. Our focus is always on real-world practical insights that you can apply in your everyday work. In this episode, LCN Legal’s co founder, Paul Sutton, talks to Borys Ulanenko about the role of AI in transfer pricing, particularly in benchmarking. Boris is the founder of ArmsLength AI, which offers a tool that greatly enhances both the process and the outcome. Paul and Borys look at, among other things, the problems and challenges inherent in the benchmarking process, the role of AI in addressing these, some real-world examples of what ArmsLength AI can do for its clients, and the things that people should consider when assessing any AI tool. We hope you enjoy the discussion.

Paul Sutton: Hi Borys, well, great to have you here on the podcast. Many thanks for joining.

Borys Ulanenko: Hello. Hi Paul, thank you for having me.

PS: So we’re here to talk about AI – artificial intelligence – and its role in transfer pricing benchmarking. So just by way of background, it would be great to hear about: what was your route to transfer pricing in the first place, what kind of experience did you have there, and what made you start to look at AI?

BU: Yes, so I am in transfer pricing for around eleven years now, and I started pretty typically. I started as an intern in Big Four firm, worked there for a couple of years, transfer pricing. The reason why I’ve got into transfer pricing was pretty random. So it was not my plan: I was just searching for the first entry-level job and Big Four seemed to be a good option, and then tax advisory seemed to be pretty interesting option, and then transfer pricing within it was just the opening that was available. Back then I was in Ukraine, the transfer pricing law was just enacted. It was 2012, and that’s how I got there.

And I was pretty lucky because I’ve got in this early stage of the transfer pricing teambuilding, and it was a lot of also business development activities, first conferences, first webinars that we were running there. Actually live events, much more live events than people are doing now. After a couple of years, I moved to work for in-house role, also in transfer pricing. Worked for Royal Dutch Shell for seven years. It was a very good in-house experience building this internal function. And that’s a pretty different perspective. Everyone talks about that – this is a different perspective – but it is indeed. And especially with companies like Royal Dutch Shell where you have… our team was around 30 people in transfer pricing, just doing transfer pricing full time, which is a huge team. I now know that there are teams that are bigger, but at that time, I think it was probably one of the biggest or the biggest in-house transfer pricing team in the world. Very exciting times.

And then I started gradually shifting my focus on technology, and that’s where I moved to a company called Aibidia. This is, I would say, a little bit more traditional technology in the sense that its basis is focusing on multinational companies and serving the largest companies with end-to-end transfer pricing platform from compliance to operational transfer pricing and other areas. Very good company, where I had this first full-time technology experience, basically, but no AI back then yet.

And I was a little bit of a coder myself, I would say. I didn’t use those skills in my job there, but it was just more of on a hobby side. But it allowed me to understand: how do you actually apply technology for transfer pricing? And then we’ve got all the ChatGPT buzz starting, right? So the ChatGPT goes live, everyone goes crazy. Like, ‘that will change our world, and how do we apply this technology everywhere we can?’ And then people are saying we all will be left without jobs in six months, in one year, whatever timeline throwing there. I was pretty excited about that part. So I started using ChatGPT myself in every area I could: in my hobby, in my daily life, in my work, in marketing. And then also gradually I was doing a little bit of transfer pricing – traditional work like writing local files and master files for some of my old clients. And then I was actually using it quite a lot and I realised that it is actually a big productivity boost as soon as you know how to use it. So that was pretty exciting. And at the same time, while working for Aibidia, I was talking to a lot of advisory firms who provide transfer pricing advice to their clients. And I quickly realised that that part of our industry is fundamentally underserved in terms of technology.

So while most of the companies like Aibidia, they build software for multinationals and in-house teams, if you look at advisers, advisers are almost exclusively are using traditional Excel, Word, whatever. They don’t really use a lot of specialised technology. And then I’ve seen this potential in: can I build something for them for solving their problems? And can I use AI for that? In the process of actually building – so can I actually build technology using AI? – and can I build technology that has AI in its heart to help them with some problems? And that’s where I started ArmsLength AI. It was eight months ago.

PS: Right. Really interesting. And I know that a lot of listeners will be familiar with some of the materials that you share, and clearly there’s no point in looking at tools and solutions unless there’s a genuine problem or issue to be fixed. So would you say that benchmarking in TP is broken? Would you say that the industry or practises are broken? And if so, what are some of the issues that you encounter?

BU: Yeah, I think benchmarking is probably the most broken transfer pricing problem that we have. I think overall transfer pricing is pretty complicated in terms of processes and how it actually works. But if we look at benchmarking studies, it’s broken at several levels. The first level is that companies… we do, globally, too many benchmarking studies. That’s my strong belief. So there are a lot of repetitive benchmarking studies that we do. There are a lot of benchmarking studies that are done where they probably are not required. And tax authorities and international organisations also recognise that. And the first trend we have seen, starting in 2015, was implementation of low-value-adding services. For example, safe harbour, right? Like 5% markup. Here you are, you don’t need to do benchmarks for that type of services because it’s obvious it’s 5%.

And now we have Amount B, which is kind of the same thing, or was about to be the same thing, for distribution activities. (Now it’s getting more complicated than that.) And then we also see some tax authorities locally introducing something similar, or along the lines of that – in Australia or New Zealand, for example. In Australia, the ATO was publishing the suggested ranges, basically, like low-risk ranges for inbound distributors, for example, which was a great initiative. So my strong belief is that when you look at benchmarking studies that are being done, there are so many similar benchmarks. And if you ask any TP professional who was working in transfer pricing for a few years, they will tell you that the final range will always be in those values. And then why do we do all of those benchmarks every time? So that’s one thing, right?

The other thing is: let’s assume that, OK, we do know that we needed to do the benchmarking study and we probably eliminated most of repetitive benchmarks and we introduced proper safe harbours (which still hasn’t happened, hopefully will happen in the future). There is still a lot of benchmarks to do. Then the process of benchmarking is pretty problematic too, in many cases. So, for example: databases is one big problem, right? There are only a limited number of databases that are available. Some of them are having the characteristics of a monopoly or oligopoly. The prices for data are extremely high, especially for companies that are not big force and not have a huge scale and need to do just 10, 20, 30 benchmarks per year. The price for data is prohibitive.

Then if we’re looking at the process itself, then there are a lot of different approaches. And each firm does benchmarks differently. For example, they apply different filters, right? Let’s talk about comparable companies searches. We are using traditional Bureau van Dijk database (now Moody’s analytics). You apply certain filters, and each company would apply their own filters and would use their own logic for applying those filters. And some companies will use one filter, the other companies will use the other filter. So there is no standard or even guidance on what those filters should be. And then everyone does it as they like.

And, moreover, when I was still in KPMG, I was coming from a more econometrics and statistical background, and I was like: OK, do those filters even make sense? Like, for example, the traditional filter is revenue, right? So we would eliminate companies that are too small, for example, or too big. And then I was running a lot of regression models, based on the data available in the databases, and it was evident that revenue actually doesn’t affect margins as significantly as advisers would claim to apply those filters. So the real reason why those filters are applied, for example, is just to narrow down the list of companies that needs to be manually reviewed, because that’s time-consuming part. But many of those filters really don’t have good substantiation behind that, and there is no guidance.

And then even if we solve that problem, then we are coming to the next problem, which is manual review, right? So your databases are still not having a lot of good information about the companies, about the nature of their business. So the descriptions of those comparables are pretty weak. And that’s why most of advisers around the world, what they are doing, they are going and manually checking each website of the company. If we are talking about, like, private companies and for public companies, they would be also checking their annual reports, trying to read through that, understand what those companies are doing, and make some decisions about if those companies are comparable or not. This is the most time-consuming process. It takes about 80% to 90% of time totally spent on preparation of the benchmarking study.

And this process is problematic because it’s subjective. It takes a lot of time. You don’t want your best resources to be allocated to that work. And since it’s very time-consuming, you may also have a bottleneck, where your team is overloaded with this type of work. And it doesn’t bring as much value as traditional, let’s say, transfer pricing advice, right? So those are like a map of all possible things that are broken in the benchmarking space, which is probably just a subset of everything that is broken.

PS: Right, OK, so we’re talking about two fundamental things. One is filtering and whether that makes sense: the filters that are adopted, are they rational or not? And then there’s the process of the manual review of the comparables, or potential comparables, that come out. So what’s the role of AI in all of this? Is it just about the second step in terms of replacing the manual review? Is that the focus of your thinking right now?

BU: Yes. So the focus of AI at this point of time is to eliminate manual reviews, or augment humans in manual reviews to the degree where the time spent in this part of the work will be very limited, and really focus on adding value and verifying what AI does. But if you think about that, the reason why we apply all the filters to arrive to the set that we want to review… the real reason is because we want to narrow down the list of companies to be manually reviewed in reality, right? So if you have a reliable algorithm and reliable AI that can do benchmarks, you don’t need to apply as many subjective filters, or filters that are not supported and can be disputed, right. You can apply your manual review, which will be AI review, to a much broader scale. Because things that we apply filters to are sometimes pretty arbitrary, and everyone knows about that.

Like, for example, industry codes, SEC codes: companies registering under them often do not mean that they are going to do business in this area. And many companies that are actually maybe good comparables for you are registered on completely other codes and you are ignoring them basically because you applied the filter that eliminates that code under which they’re registered. So using AI can allow you to apply to find better comparables at a much higher scale.

PS: Interesting. OK, when we step back and think: why are we doing this? What is the outcome that we’re trying to achieve? I would say, well, it’s not that we’re trying to achieve the theoretically highest quality comparables analysis. The practical task is that we’re trying to adopt price-setting, TP, policies which are least likely to be challenged by tax authorities, and to avoid the risks and the time and the cost involved in dealing with those objections. So in terms of that process, it makes complete sense what you’re saying, but how do we address that kind of practical end result that we’re trying to achieve, which presumably means getting clearer in terms of overall principles to be adopted not just by taxpayers but also by tax authorities?

BU: Yes, great question. When we do benchmarking studies, there are actually two type of goals that we are trying to achieve, and they are often contradicting each other.

So the first case is where we are doing a benchmark to support what has already happened or support something that we want to achieve, right? So for example, we know that we want for this jurisdiction to have this margin. And then we go to our advisor and we tell them: ‘Can you do us a benchmark? But we really want that benchmark to be in this range’. That’s one story.

And the other story is where we want to build something robust. We don’t care much about the outcome – there is some reasonable understanding of what should be there, but we want to build as robust benchmarking study as possible because, for example, we want to apply it proactively going forward in operational transfer pricing process, or document that in an intercompany agreement in one way or another. And then we know that we will try to apply it at scale. For example, for all Europe, a lot of European subsidiaries we have, we want to apply it and then we may face a lot of challenge in the future just because of the scale of application of that model.

So they’re a little bit contradicting each other.

But a robust process, which includes AI in it, can help both strategies, right? So in the case where you are targeting some particular range, or you are doing the benchmark trying to achieve a particular range, AI can help you because of the fact that you can iterate much faster through the companies and comparables: you can tweak the strategy, and rerun the search strategy and accept/reject criteria, to arrive to your desired set quicker, right? Because what is happening in reality is that a manager would give a consultant a task to do a benchmarking study. They would go through the comparables, they would have accept/rejects, they would calculate the range, and then the range is not what is expected. And then they would go back and think: how can we tweak the criteria to get us to the range? So this process is pretty lengthy and requires a lot of iteration between different people. With AI, you can do that faster, basically. That’s one thing.

And in terms of robustness, AI applies criteria much more consistently and often will find some peculiarities about those comparables which humans often overlook. And that way you can build more robust benchmark that will be more difficult to challenge by tax authorities.

PS: Makes complete sense. And maybe it’s just me, but I’m guessing that a lot of people think of AI as just this black box, and you kind of put in these prompts and something comes out and you’ve got no idea about how it got to that answer. Or people worry about hallucinations and so on. And if we talk about the end result that corporates are trying to achieve – whether trying to justify a previously adopted TP position or actually genuinely being open in not caring much about the range, but just wanting it to be genuinely defendable – it’s not just about the end result, is it? It’s also about being able to clearly articulate and demonstrate the process, so that that sequence of logic can be tested and hopefully accepted. How does that work in terms of the tools that you use?

BU: Yeah, that’s a great question too. I actually spoke to one of Big Four firms yesterday in the US. They were testing my algorithm for their benchmark, and they said that specifically what they liked about my way of thinking is that I try to break down the process as much as possible and to demonstrate all the decision tree and decision logic that AI does. So that it’s as ‘not black box’ as possible, basically.

And the idea is the following: you don’t want to give AI too much freedom. You don’t want it to give it too big of a task – to try to get AI do all from company name to final decision, without human intervention and human review at critical steps. And to give a human a chance to review what AI does, you need AI to explain each decision and each step it does. So how this works is then we’re breaking down all the manual review into smaller steps. And the first step would be: we need to access and go to the company website, and we need to read all the content from the company website. And if the company website is not available, we want AI to find the official company website and scan through that website, basically.

And that’s where we put a pause, and we give a human information about what company website did it identify if there was no website? And does human agree that this website should be used, right? So we’re not telling AI, like, just do it. We’re adding human decision there. And then we document that, right? So like how did we actually find that? And then we ask AI to write descriptions of each comparable based on the available information. But we also do that in very particular way where each piece of information, each fact, needs to be supported by supporting documentation, by a screenshot, by a link to a particular website highlighted there, so that again a human reviewer can verify that. But then by building the algorithm this way, each step is controlled by human. And AI works much better in this way too, where this is like kind of self-control feedback loop algorithm built at each step. And therefore as a result, if you look at all the documentation and all the proof points generated by this process, with human input at some critical moments, then the full process is very robust and you have very detailed proof points for every step, which is much more detailed than when you just give a task to a human.

Because a human, we don’t have time to document every step that we are doing in benchmarking process. I mean just if you don’t use any tools, what you would be doing is you would just visit the company website, you would read through it, and then you would say accept or reject, and maybe you will write like one sentence of description why. That’s all. If you are using AI, then you can break this down and give AI much bigger task of analysing much more information and documenting it much better, showing all the decision, basically. So as a result, those benchmarks tend to be also much more robust than just human-made. Of course you can do what AI does, and probably, like, if you put ten people doing that, you may get better results than AI, but that’s just unrealistic. Those benchmarks would cost hundreds of thousands of dollars.

PS: Yeah, that’s so interesting. When I think about our experience of all of this, obviously we don’t advise on benchmarking or any of these things, but we do read a lot of benchmarking reports. And I have to say the quality, even amongst Big Four firms, is typically really poor. Because generally speaking, there’s just like one or two lines about selection criteria, and that’s it. So going back to the original focus, so it sounds like the output is not just like ‘this is the answer, and this is the database that we’ve referred to’. It’s the sequence of steps. It’s like the audit trail for the decision, which makes a lot of sense.

BU: Yes. And just to add to that, when you articulate the strategy of the companies that you want to accept, the criteria that you want to accept… if you are doing that for AI, you are required to do that in a particular form and you need to be pretty precise. And then you know that the strategy is applied consistently. And that forces you almost to think about the strategy that you’re applying in much more detailed way. And that is another advantage of that. Because what I tend to see in human-made benchmarks is that there is a very vague definition of what we want to accept and what we want to reject. And then each reviewer interprets that themselves – and they are able to interpret it, but interpretations are very subjective and not documented, right? And then you end up with a set of accepts and rejects, but you don’t really know what were the actual criteria. They were kind of implicit criteria that the reviewer had. And then using AI forces you to do much more concrete strategies.

PS: Yeah. And for me it immediately raises the whole rationale for Amount B. This premise that developing countries might not have the resources for detailed functional analysis or benchmarking analysis and so on. And this is a totally different approach in terms of tools which potentially are much lower-cost and which could avoid all the additional complexity that Amount B seems to be creating here.

BU: Yes, I think to a degree. Though, again, we have another problem, which is database, for example, right? So even if you have a good algorithm for analysing this data very fast, of course you’re reducing the price of the benchmarking study. But still, in the total price of the benchmark, the portion of data is still very significant cost for anyone who does benchmarking studies. And there are obviously developments on that front too, which, I mean, I’m not directly participating in, but I’m talking to a lot of databases, who are either traditional players or some new entries. And I think that in the next five to ten years the situation will change drastically. I think all the data will become extremely cheap. But for now, yeah, I think it still makes sense to have some safe harbours. I’m not sure Amount B, as it stands now, is a good example of that, just because of the complexity they added in the last paper. But generally I think it’s the right direction still.

PS: Yeah, really interesting. OK, let’s move on and talk about what specifically you do. So my understanding is that you create tools that help TP advisers to achieve increased efficiency and accuracy, increased robustness in the benchmarking reports that they produce. So how does that work? How does that work in terms of them as users?

BU: Yes. So there are a couple of scenarios there. And when we started ArmsLength AI, we actually didn’t start with benchmarking studies. We started with other tools. And other tools were things like drafting functional interview questions. So when you go to a discussion with a client, it helps you prepare the questions that you need to ask. You give it some basic facts about the discussion you’re going to have, you get the questions to ask. And then as the next step you can get your meeting notes, and you can get intercompany agreement, and you can put that all into functional analysis tool that will use that data and write a detailed functional analysis for this model or customer, the client that you’re talking to. That’s where we started.

And then most of the users were transfer pricing advisers that just wanted to try how AI can help them in those particular tasks. So we basically took the standard AI available models like GPT-4 for example, and then optimised them for those particular smaller tasks. What we discovered quickly is that while those tasks make sense, people don’t do them as often as we hoped, basically.

So yes, you do functional interview questions, but you do it like once per month, maybe once per couple of weeks maximum. The same about functional analysis. Of course, when it’s high season and you prepare a lot of new local files and documentation, then functional analysis tool is helpful. But again, you don’t use it as often. And then in discussions with customers, we realised that benchmarking studies is something where they spend a lot of time. And they themselves believe that that’s the part that should be automated. And then we started looking actively into benchmarking.

And then basically how it works for benchmarks: advisers are still using their databases that they have now. Typical example would be TP Catalyst from BvD/Moody’s Analytics. They would still apply all the filters they want to apply on the data that they have in the database: like industry codes, revenues, independence, whatever other filters they want. And then they arrive to a set of companies that they need to manually review. And then instead of taking these 200, 300 companies and then giving it to their juniors or consultants to review manually, and then manager needs to review what those guys have been doing, they come to our AI benchmarks tool.

They just import the list of companies and what we need is just company name, country and company website. And then they work within the tool for AI to do the rest of the job. (So it’s not just about AI, right: there is a lot of different algorithms and different scripts that are tied together for this to work.)

But the first step will be for our algorithm to access each company website, analyse all the information and write detailed descriptions of each company. That’s the first step. You apply it to hundreds of companies at the same time. You do this very fast. You get the result from 100 companies in about 10 to 15 minutes. So it would take you like half an hour maximum to get very good, actual, relevant, current, accurate business description of each comparable in your set.

And then the next step is where you are setting this accept/reject strategy. It’s your task. And then you ask AI to apply those criteria to this set of companies that it already knows from the web search. And then it creates for you the accept/reject metrics, where each decision is explained and you can track back to the website description or to the information that AI was able to find about the company to understand how AI made each decision.

And then we have additional step where we guide humans how to actually review AI results too. So we indicate which of those companies that AI analysed are really important for a human to review, based on certain criteria, and which of them you probably don’t need to review because the chance that you will change the ultimate accept/reject decision is extremely low. So we’re just also guiding and giving humans information about what they actually need to do with what AI generated, because we don’t want them to just take AI results and then paste them in their final deliverable. We want them to do the review in a way that would maximise the quality of the overall benchmarking study. And then they basically use these accept/reject metrics and then they recalculate the ranges based on that, and that’s basically it, right? So we automate actually all the manual steps in the process.

PS: OK makes complete sense. And if I can ask, what’s your current pricing model in terms of offering these tools?

BU: Yes, currently there are two scenarios. There is a subscription model. That’s the first standard model, which is now around $220 per user per month, which gives users access to all the tools, including functional analysis and other things that I have mentioned and benchmarking studies. But there is a limitation on number of companies you can process per month. So you can analyse maximum 300 companies with this algorithm per month under this fee. And then for extra companies, you are paying extra, basically fifty cents per each additional company.

And we have several customers which are doing benchmarking studies at a big scale. So one case where the company does more than 150 benchmarking studies per year, the other company which does more than 1,000 benchmarking studies per year, because they are doing that for other advisers. And with them we are negotiating separate terms: there is a special enterprise plan on our website where, based on those parameters and based on the scale, we can arrive to a smaller fee per company, basically just because of the scale. So those are two models.

But ultimately in most of the cases you pay per use, right. So there is rather small subscription fee and then the rest is like how much you actually use it, how many companies you have analysed with that?

PS: OK, got it. And we touched on one of the fears that people have in relation to AI, which is ‘It’s just a black box which produces the results without the accountability’. So you’ve talked about that in terms of the audit trail or the documentation of the sequence of decision making. Are there other misconceptions or fears that you think people have or that you’ve encountered in relation to these kind of tools?

BU: Yeah, for sure. So I think the hallucination problem is the big one, which everyone is concerned about. Like how do we know that AI is not making up some facts? Taking benchmarking study as an example, how do we know that the description that the AI has written about the website is accurate and it doesn’t include some facts which are not on the website, for example, just made up some facts. And unfortunately, that’s an inherent feature of the large language model as a technology. So you never can give 100% guarantee that that’s something that will never happen. It’s like saying that Mercedes Benz on average would go 500,000 miles without breaking, but at some point it will break, or it can happen a little bit earlier too, sometimes. But it’s all about the design of the system, right? And it’s about us ensuring that that is not happening, by designing how we’re asking AI to do certain tasks and double-checking the results with, for example, other large language model or what one model generates.

And that substantially reduces the chance of that error. And we haven’t so far identified any hallucinations in our data set (because we’re now doing a lot of QA quality reviews of what AI does). And actually our typical way we start the conversation with any transfer pricing firm would be that they will send us their benchmarking study that they have already done, and then we will run it through our system and give them back the results. And then they compare. So it much more often happens that AI discovers something that human has missed. That’s around 20% of the cases. Then there are some cases – and we need to be obviously honest about that – where AI misses something and doesn’t do it right, though it happens now in less than 10% of the cases. And we have a way to identify those cases basically and highlight that for humans to review. So where it’s not sure or the mistake can lead to significant outcomes. So that’s what we can identify and highlight at least. And during those reviews we haven’t discovered any cases where it hallucinates, basically, though it’s still theoretical possibility.

And then there are some other issues with this. Like, for example, your algorithm cannot analyse all 100% of the companies and sometimes it will not be able to analyse this company or the other company, which happens, for example for Asian companies. And the main reason for that is, often like South Korean companies, Japanese companies, what they do is they will write texts on their website, but they will make them as pictures, not as actual texts. So they make pictures of text, which is difficult for AI to analyse, right? And then the percentage of companies that can be successfully analysed is lower. That’s another scenario. And then, well, our answer is like ‘yes, OK, it reduces the success rate from 95% to 85%, but you still have 85% of companies analysed successfully, right? So it still makes sense’.

And I think one other big one, which actually less relates to benchmarking studies, is data security and information security aspect. That’s more prevalent for our other tools like functional analysis, where you need to disclose some facts about the business model, intercompany agreement, etc. And you need to be sure that this information is secure, is not used for training AI, etc. And there, what you need to look in is what technology does this company that provides you AI service use, and how they explain how they use your data. So in our case, we’re only using Microsoft Azure, and Microsoft-provided AI, and therefore the terms and conditions of the use of information under this scenario is very different from what you would see in like ChatGPT, for example. So Microsoft explicitly gives a guarantee that they are not going to use your data for training purposes, they are not going to store your data, etc. So it went through AI, AI gave you the result, and then AI forgot about that basically. So that’s what you need to look into when you are assessing any AI tool on the market.

PS: Makes complete sense, thank you. And I guess going back to the hallucination point, one way of looking at it is: AI tools are a bit like hiring a new member of staff. You want to test them out for a bit and build up trust with them before you maybe trust them with bigger projects. And it’s going to be a natural leaning in, given that the potential benefits in robustness, in scale, in efficiency and productivity are so huge.

BU: Yes, indeed. And that’s why I think you need to start small. So the idea that we had is this. We do benchmark for you and then you compare the results, and then you have already a feeling of like, does it make sense? And then we give you a trial access. You try yourself. And then you need to build the confidence and comfort and know how to work with it before you actually start applying it to your real-life scenarios, obviously.

PS: Fantastic. Great. So final question: where can people get more information if they’re interested in finding out more about AI and the solutions that you offer?

BU: You can find us on armslength.ai. That’s where you’ll find most of information about the product and how it works. And you can also reach out to me on LinkedIn. You will see my name and just find me on LinkedIn. I’m also happy to have conversations. It’s very interesting times and everyone wants to talk about where we’re heading next with this.

PS: Totally. Well, thank you so much for sharing time with us. It’s been fascinating and I’ve learned a lot. So thank you again.

BU: Thank you. Thank you, Paul.

Outro: Thanks for listening to The LCN Legal Podcast. We’d love to hear what you think. You’ll find the details on our website: lcnlegal.com. And in the blog section, you’ll find a transcript of this episode, with contact details for Borys. If you enjoyed this episode, please subscribe. Go to your podcast provider and search for ‘The LCN Legal Podcast’. Until next time, thank you and goodbye.

More podcast episodes

Listen on-demand to previous episodes of the LCN Legal Podcast as they get released.

Join our inner circle

Sign up for our free email newsletter

Receive weekly practical insights on how to keep intercompany agreements and cross-border corporate structures tax-audit ready and transaction ready.