OutputExtracts descriptive keywords and tags from visual scenes or video frames for indexing, search, and content summarization.