This is a follow up question to a previous one that I did some days ago, in ehich I asked the community if they knew an AI that could do the following task
I have many emails that pretty much answered a question that I formulated to the addressees. I predict that these replies gave a specific type of answer to that question: So imagine that the question overall was "do you think that ice cream is the best dessert that exists?" and I want to see how many if them answered something like "yes it is!", so that no matter how the reply is formulated, it basically answers something along these lines
I would like to use an AI to see the degree of accuracy of this prediction, but there are a few emails that I don't want to see their actual content under any circumstances, to be unbiased
So, many people here told me that it was a trivial task. I was thinking in using perplexity, which is one of the most reliable AI models that I have been able to use. Do you think that the free version could be enough? Or paying for a better model could be crucial for reliability? And if you think that Perplexity is not the best option, which AI would you suggest?
Also, I have transformed all emails into a big pdf document with many pages (although I have not seen the contents of these pdfs of course) and I have joined them into a single pdf (1000 emails was an exaggeration, there are actually about 100 in total).
I was thinking about two possible methods:
One is to give the pdf with all the emails to the AI and ask it to make a percentage or "score" of all replies that pretty much accomodate to the answer that I am expecting, and the same for the ones that are neutral, unrelated or don't give an answer and as well as the negative ones. However, I fear that this may not be very reliable, and since I want to avoid looking at some of the emails (at least for now), I don't want to go see the actual emails to test if the AI has gotten this right
So I was thinking about another option: I did another pdf document of "expected" answers. In this document I posted the original question that I asked to all the addressees (the questions are overall the same, but the details change in each case, so there is pretty much a unique question by email) and then I actually wrote the type of answer that I expect. Then I would ask the model to check the degree of accuracy or similarity that my written "expected" answers have with the actual ones, and then ask it to give me a number like a percentage or score. Do you think this would be a good idea?
A third option is AILYZE, which is a model that is more or less suitable for what I am looking for, last year I used it for a similar task, but the problem is that it directly shows you why does it think that the texts that you have fed to it accomodate or not to your expectations giving you actual examples. Since I don't want to see some of the emails, I don't know how could I use it without seeing the actual replies. Another thing is that it didn't gave actual percentages or scores in numbers, but the output was something like "the majority of answers are negative and the minority of them are affirmative..." whatever, but it didn't give any actual numbers that could give you a better idea of what was going on
And finally, another big problem is what prompt could I use so that the reliability would be maximal, since I am pretty much noob, I am a bit lost on this as well...Any advice or ideas?
[link] [comments]