Interpretable Steering for LLM Decision-Making using Activation Engineering
Built an interpretable framework to analyze and steer LLM decision-making using activation engineering and representation analysis. Conducted large-scale evaluation by integrating over 73,000 World Values Survey samples and running 2,000 persona-based prompts to test demographic and alignment biases. Applied semantic vector steering and regression analysis to evaluate and neutralize LLM biases in environment versus economy decision tasks. • Utilized LLMs for social simulation and bias analysis using text data inputs. • Designed and executed quantitative evaluation prompts and ratings. • Used PCA and vector steering to assess model decisions. • Generated insights on alignment between AI output and human patterns.