Analyze the impact of system prompts on LLM performance by: 1) Comparing Claude's Opus 4.8 vs 5.0 system prompts from GitHub history 2) Testing response quality with/without system prompts using smol framework 3) Measuring token usage efficiency between prompt-heavy vs task-heavy contexts 4) Documenting how model self-perception affects output (e.g. hierarchy mentions in prompts)