Skip to main content
Language feature score helps analyze how well the language used in a response conveys the intended message, whether it addresses the question or issue comprehensively, and if it is free from ambiguity or confusion. Columns required:
  • response: The response given by the model

How to use it?

By default, we are using GPT 3.5 Turbo for evaluations. If you want to use a different model, check out this tutorial.
Sample Response:
Higher language features scores reflects a good response.
The reponse generated does not seem good, it has innapropriate words like “dummy”, there are some grammatical errors and uses unnecessary slangs like: “I will guide you”, “it is pretty straightforward” Resulting in low language feature scores.

How it works?

We evaluate language features by determining which of the following three cases apply for the given task data across features such as fluent, polite, grammatically correct, and coherent:
  • The response is highly rated on these features.
  • The response is moderately rated on these features.
  • The response is poorly rated on these features.

Tutorial

Open this tutorial in GitHub

Have Questions?

Join our community for any questions or requests