Monday, 7 July 2025
31.9 C
Singapore
35.9 C
Thailand
22.6 C
Indonesia
30.1 C
Philippines

ChatGPT could soon gain the ability to see

ChatGPT’s Advanced Voice Mode might soon include a live camera feature, enabling AI to identify objects and interact visually.

ChatGPT’s Advanced Voice Mode, known for enabling real-time conversations with a chatbot, might soon include visual capabilities. Code uncovered in the latest beta version of the app hints at introducing a “live camera” feature. This discovery in ChatGPT v1.2024.317, as reported by Android Authority, suggests that the rollout of this exciting feature could be just around the corner. However, OpenAI has yet to confirm an official release date.

A glimpse into the feature’s early tests

The idea of ChatGPT having a visual edge has been introduced previously. During the initial alpha testing phase of Advanced Voice Mode in May, OpenAI demonstrated its potential visual capabilities. In one example, the chatbot used a phone’s camera to identify a dog, recognise its ball, and associate the two in the context of playing fetch. This ability to observe, understand, and link objects to real-world scenarios was widely praised by early testers.

Alpha testers were quick to explore the feature’s uses. A notable example came from a user on X (formerly Twitter), Manuel Sainsily, who utilised the camera to ask questions about his new kitten. This interactive capability showcased how the feature could provide fun and practical benefits.

When Advanced Voice Mode entered beta testing in September for ChatGPT Plus and Enterprise users, its visual functionality was notably absent. Despite this, the voice feature gained immense popularity for enabling natural, dynamic conversations. According to OpenAI, users could interrupt the chatbot at any moment, and it could even pick up on the speaker’s emotional tone.

What sets it apart from competitors?

ChatGPT could have a unique edge over rivals like Google and Meta if the live camera feature is introduced. Google’s conversational AI, Gemini Live, may speak over 40 languages but lacks visual processing capabilities. Similarly, Meta’s Natural Voice Interactions, showcased at the Connect 2024 event in September, cannot use camera inputs. While these systems are competent in their ways, OpenAI’s visual feature could redefine how AI assistants interact with the world.

Desktop users can now enjoy Advanced Voice Mode

In a related update, OpenAI announced that Advanced Voice Mode is now available to paid ChatGPT Plus users on desktop. Previously limited to mobile devices, this update means users can now access this feature directly on their laptops or PCs.

The introduction of the live camera could mark a significant leap forward, combining the ability to see and hear into one seamless AI experience. While the exact timing remains uncertain, the potential impact of this development is already generating excitement among users and industry experts alike.

Hot this week

Mainland investment boom lifts Hong Kong’s market

Chinese firms turn to Hong Kong listings after mainland investors spend US$93B on stocks, eyeing global growth and fresh funding sources.

SBF and CapitaLand Investment unite business leaders to reaffirm support for national defence on SAF Day

SBF and CapitaLand Investment host SAF Day event, reaffirming business community’s commitment to national defence and support for NSmen.

Secretlab teams up with Genshin Impact for first Liyue-inspired chair and desk collection

Secretlab reveals its first Genshin Impact collection, which includes Liyue-themed chairs and a desk inspired by Xiao, Ningguang, and the Lantern Rite.

WWE 2K25 confirmed for Nintendo Switch 2 launch on 23 July

WWE 2K25 will launch on Nintendo Switch 2 on 23 July, offering full game features, new content, and multiple special editions.

Apple hits key milestone in foldable iPhone development

Apple’s foldable iPhone has reached a key milestone with a working prototype, and the company is eyeing a potential launch in the second half of 2026.

Embedded LLM and AMD launch TokenVisor to boost AI monetisation for GPU neoclouds

Embedded LLM and AMD launch TokenVisor, a platform enabling monetisation and management of AMD GPU clusters for LLM workloads.

Kahoot! teams up with Tour de France to deliver interactive learning experiences

Kahoot! partners with Tour de France to bring interactive cycling-themed learning to classrooms, fan parks, and homes worldwide.

How will AI integration transform industries in 2025?

AI is transforming industries in 2025 through innovation, efficiency, and new business models. Explore key tech investments, sector impacts, and future trends.

vivo introduces X200 FE, its first compact telephoto flagship smartphone

vivo launches the X200 FE in Singapore, a compact flagship with telephoto imaging, ZEISS optics, and powerful performance in a lightweight body.

Related Articles

Popular Categories