
Understanding Multimodal AI
- What is Multimodal AI?
- Supported Modalities
- Model Requirements
- Text (natural language)
- Images and video
- Audio and speech
- Charts and diagrams
- Structured data
Working with Images
Uploading Images
Access image upload

- PNG
- JPEG/JPG
- GIF (static)
- WebP
- BMP
- SVG (as an image; code parsing may be limited)
Add context (optional)
Submit for analysis
Types of Image Analysis
General Image Description
Text Extraction (OCR)
Chart and Graph Analysis
Technical Diagram Interpretation
Document Analysis
UI/UX Analysis
Content Categorization
Object and Entity Recognition
Example Prompts for Image Analysis

Advanced Image Interactions
Reference Specific Parts of Images
Reference Specific Parts of Images
Compare Multiple Images
Compare Multiple Images
Sequential Image Analysis
Sequential Image Analysis
Specific Use Cases
Text Extraction (OCR)
Extract and work with text from images:
Upload an image containing text
- Scanned documents
- Photos of printed materials
- Screenshots with text
- Whiteboards and handwritten notes (with limitations)
Request text extraction
Work with the extracted text
- Summarize the content
- Answer questions about the text
- Format or structure the information
- Translate the extracted text
- Find specific information within it
- Text clarity and image quality
- Font type and size
- Background contrast
- Image resolution
Chart and Graph Analysis
Get insights from data visualizations:Upload a chart or graph
- Bar charts and histograms
- Line graphs
- Pie and donut charts
- Scatter plots
- Area charts
- Combined visualizations
Ask for analysis
Explore specific aspects
Technical Diagram Interpretation
Understand complex visual information:Upload a technical diagram
- Flowcharts and process diagrams
- Network and system architectures
- UML diagrams
- Circuit diagrams
- Engineering schematics
- Entity-relationship diagrams
Request explanation
Ask for specific details
UI/UX Analysis
Evaluate and improve user interfaces:Upload UI screenshots
- Website pages
- Mobile app screens
- Software interfaces
- Design mockups
- Forms and interactive elements
Request design analysis
Focus on specific aspects
Working with Audio
AI SecureChat can also process audio content with compatible multimodal models:Uploading Audio
Access audio upload
- MP3
- WAV
- M4A
- OGG
- FLAC
Add context (optional)
Submit for processing
Audio Analysis Capabilities
Transcription
Meeting Summarization
Translation
Speaker Identification
Content Analysis
Q&A on Audio Content
Example Prompts for Audio Analysis
Try these prompts after uploading an audio file: Transcribe this audio recording. Summarize the key points from this meeting. What action items were mentioned in this recording? Translate this speech to French. Identify the main topics discussed in this conversation. Create a timeline of events mentioned in this recording. What was the sentiment of the speakers in this discussion? Extract all the numbers and statistics mentioned. CopyAudio Transcription and Processing
- Basic Transcription
- Meeting Summarization
- Content Analysis
- Verbatim transcription (including filler words, pauses)
- Clean transcription (removing stutters, false starts)
- Timestamped transcription
- Speaker-attributed transcription (where possible)
Audio Generation
Some multimodal models may offer limited audio generation capabilities:- Are typically more limited than image generation
- May only be available with specific models
- Often have restrictions on duration and complexity
- May be in experimental phases depending on your organization’s Prisme.ai version
Best Practices for Multimodal Work
Use High-Quality Media
Be Specific in Prompts
Combine Modalities Strategically
Verify Critical Information
Consider Privacy and Sensitivity
Use Canvas for Complex Work
Save Intermediate Results
Provide Context
Troubleshooting Multimodal Issues
Image not being processed
Image not being processed
- Check that you’re using a multimodal-capable model
- Verify the image format is supported
- Ensure the image isn’t too large (try compressing)
- Check that the image uploaded completely
- Try describing what’s in the image as context
- For complex images, try focusing on specific parts
Poor image analysis quality
Poor image analysis quality
- Improve image quality (resolution, lighting, focus)
- Try a different multimodal model if available
- Be more specific in your prompts
- For text extraction, ensure text is clear and readable
- For charts, make sure data points and labels are visible
- Try cropping the image to focus on the relevant part
Audio processing issues
Audio processing issues
- Check audio quality and reduce background noise if possible
- Verify the audio format is supported
- Try shorter audio segments for complex recordings
- Provide context about speakers, topic, or terminology
- For non-English audio, specify the language
- Try a model specifically optimized for audio if available
Image generation not working
Image generation not working
- Verify your model supports image generation
- Check if generation features are enabled in your instance
- Be more specific and detailed in your description
- Break complex images into simpler requests
- Try different styles or approaches
- Be aware of content policy restrictions
