Consider something like this: after extraction has been completed. Go through every extracted feature (maybe even every non-extracted feature?) and ask the model to provide supporting evidence in the extraction text. Come up with a rating system of confidence based on that (maybe even a self-fix option)
Ideas from discussion:
Ask the model to differentiate evidence at multiple level along the lines of "mentions the evidence keywords," "evidence can be inferred," "direct evidence in the text" with additional levels specific to questions.
Consider something like this: after extraction has been completed. Go through every extracted feature (maybe even every non-extracted feature?) and ask the model to provide supporting evidence in the extraction text. Come up with a rating system of confidence based on that (maybe even a self-fix option)
Ideas from discussion:
Ask the model to differentiate evidence at multiple level along the lines of "mentions the evidence keywords," "evidence can be inferred," "direct evidence in the text" with additional levels specific to questions.