Internal Coherence Maximization for Unsupervised Persona-Conditioned Value Specification in Pluralistic Alignment
Open Access DepositedAligning AI systems with human values requires more than broad principles alone. In pluralistic settings, different groups may reasonably prefer different answers, so models often need concrete examples that show how values apply in practice. This thesis studies whether such group-specific examples can be constructed automatically and then reused for downstream prediction. It uses Internal Coherence Maximization (ICM), an unsupervised method that infers labels for logically related claims and refines them to improve internal coherence while preserving support from model scores. On OpinionQA, the few-shot condition built from ICM-inferred labels reaches 81.93% accuracy, slightly above the gold-label few-shot reference at 81.03% and well above zero-shot baselines. The same pattern transfers to Persona-Tailoring, where the ICM-label few-shot condition reaches 77.37% and remains close to the gold-label reference at 77.97%. Perturbation and stability analyses further indicate that coherence improves generalization and logical stability. These results suggest that coherent unsupervised examples can serve as practical value-specification artifacts. A separate local OpinionQA intervention analysis also suggests that Human-Guided Correction may support bounded post hoc revision, though it is evaluated on a different local slice from the thesis’s main aggregate benchmark.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.
| Thumbnail | Title | Date Uploaded | Visibility | Actions |
|---|---|---|---|---|
|
|
Pei_gwu_0075M_17759.pdf | 2026-06-24 | Open Access |
|
|
icm_for_pluralistic_alignment.zip | 2026-06-24 | Open Access |
|
