Role Overview Evaluate and benchmark how two AI language models answer real questions of Spanish law by reviewing paired outputs, scoring them against a provided rubric, and delivering concise written feedback for each evaluation. This is a short-term, remote engagement…