The provided bar charts compares the performance between GPT-4o-mini and GPT-o1 in terms of metrics and datasets in a particular long-text generation work.
Overall, this assessment involves three types of tasks, which is labelled as ROUGE-L, METEOR and ROUGE-L respectively and repeated four times. GPT-4o-mini’s performance surpasses GPT-o1 at almost every testing round, particularly at task 1 named meteor at which their score reaches the highest figure among all comparisons.
With regard to task 1 and task 2, it is noticeable that they achieve a substantially higher mark compared with the latter tasks. During the task 1 phase, METEOR accounts for the highest number of scores at approximately 0.27 and exactly 0.25 for GPT-4o-mini and GPT-o1, respectively. On task 2, while GPT-o1 decreases to around 0.22, GPT-4o-mini also experiences the familiar trend, consequently dropping to 0.2. The ROUGH-L still remain the smallest figure, however, has a slight difference in pattern where both GPT-4o-mini and GPT-o1 level off to exactly 0.15 and under 0.13, correspondingly.
Turning to the final rounds, task 3 and task 4, a distinct tendency can be witnessed that METEOR is no longer the dominant candidate in this benchmark, leaving its position for ROUGH-L and ROUGH-1 separately for each mentioned task. During task 3, although GPT-4o-mini scores at 0.2, GPT-o1 is a 0.03-lower in comparison with the former in ROUGH-L, followed by a total score of 0.25 for two GPT engines, with the exception of at METEOR for GPT-4o-mini at over 0.16. For the final task phase, ROUGH-1’s total score increases to 0.2 for GPT-4o-mini and 0.17 for GPT-o1 , while METEOR’s candidates maintain the same score at above 0.17 . In contrast, ROUGH-L’s respondents also remain the same tendency with a modest figure where they account for 0.15 and 0.13 in the end.
