"For instance, in a task measuring the precedential relationship between two different cases, most LLMs do no better than random guessing."
"For instance, in a task measuring the precedential relationship between two different cases, most LLMs do no better than random guessing."