← Back to the trend map

Safety · Concept

Beneficial Trait RL for Broad Alignment Generalization

Post-training domain (helpfulness vs. coding) differentially degrades mid-trained compassion values, with helpfulness training causing 35.7% vs. 65.2% retention on animal harm benchmark.

Trend strength 4/10
Momentum +4/q
Confidence medium
Status new
Forecast horizon

Helpfulness training degrades general moral reasoning by 25.5 pp, raising concerns about standard RLHF pipelines.

Connections

Connections · 2

How this node ties into the rest of the map, and the evidence behind each link.

Signal sources

Signal sources

Dated facts from primary sources in this direction.