Validation of AI models for ITCZ Detection from Climate Data

Serra, J.; Fortes, S; Tellez, N.; Allaico, A; Landaverde, E.; Quezada, R.; Kumar, Y.; Li, J. J.; Morreale, P.

doi:10.1109/DSIT55514.2022.9943879

Citation Details

Validation of AI models for ITCZ Detection from Climate Data

This paper presents an innovative testing framework, testFAILS, designed for the rigorous evaluation of AI Linguistic Systems, with a particular emphasis on various iterations of ChatGPT. Leveraging orthogonal array coverage, this framework provides a robust mechanism for assessing AI systems, addressing the critical question, "How should we evaluate AI?" While the Turing test has traditionally been the benchmark for AI evaluation, we argue that current publicly available chatbots, despite their rapid advancements, have yet to meet this standard. However, the pace of progress suggests that achieving Turing test-level performance may be imminent. In the interim, the need for effective AI evaluation and testing methodologies remains paramount. Our research, which is ongoing, has already validated several versions of ChatGPT, and we are currently conducting comprehensive testing on the latest models, including ChatGPT-4, Bard and Bing Bot, and the LLaMA model. The testFAILS framework is designed to be adaptable, ready to evaluate new bot versions as they are released. Additionally, we have tested available chatbot APIs and developed our own application, AIDoctor, utilizing the ChatGPT-4 model and Microsoft Azure AI technologies. more »

Award ID(s):: 2034030

PAR ID:: 10430159

Author(s) / Creator(s):: Serra, J.; Fortes, S; Tellez, N.; Allaico, A; Landaverde, E.; Quezada, R.; Kumar, Y.; Li, J. J.; Morreale, P.

Date Published:: 2022-07-01

Journal Name:: Proceedings of 2022 5th International Conference on Data Science and Information

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.1109/DSIT55514.2022.9943879

More Like this