When Perspective Becomes Control: Verifying Role-Conditioned Image Generation
Abstract
Role prompts are increasingly used in image generation as a soft form of control, allowing users to render scenes from perspectives such as a newcomer, visitor, local resident, or urban planner. Unlike explicit controls such as masks, poses, layouts, or object instructions, role prompts do not define a fixed visual target, making their effects difficult to verify. We frame role-conditioned image generation as a verification problem for open-ended semantic controllability and pluralistic visual exploration. Using preliminary urban-scene examples across multiple roles and models, we show that role prompts may materialise through foreground embodiment, local salience, spatial emphasis, demographic cues, occupational proxies, or broader scene reconstruction, while caption-level descriptions often collapse these differences into generic summaries. We propose an evaluation lens based on emergence, consistency, recoverability, and semantic appropriateness, arguing that such generation requires both computational analysis and interpretive engagement with humanities and social sciences.