Self-Supervised
Frontier models have improved their math and coding capabilities at a startling rate. This is due to the success of reinforcement learning with verifiable feedback: where correctness can be easily verified, environments can be hillclimbed to superhuman performance. Soft properties, like sycophancy, honesty, research taste, and writing quality, cannot be hillclimbed this way. This is the fundamental problem of soft evaluation: soft environments require judges, and training models against a judge learns the judge, not the property.
Capability makes this problem worse. Persuading the judge is itself a capability which scales with everything else. The gap between being good and seeming good widens as models grow, and a model that has learned to exploit its judges is difficult to distinguish from one that is robustly aligned. This presents an existential challenge for alignment as it is currently practiced.
Self-Supervised is a research company that builds alignment environments: training environments that sidestep the problem of soft evaluation by using model internals in the evaluation process itself. Our research is on which internal signals survive being optimized against. Our current focus is on sycophancy and honesty, but our long-term vision is to scale self-supervised learning past judges by grounding reward in models’ native representations.