Results

These overlooked AWS re:Invent launches could solve pain points

These overlooked AWS re:Invent launches could solve pain points

AWS re:Invent is an overwhelming barrage of features, services and launches that fly by so fast you can miss a lot of things that could drive real business value.

To that end, here are a some of the announcements that team Constellation Research thought were interesting even if they didn't get all the attention that Amazon Q, SageMaker, Graviton, Trainium and Inferentia get. These items were dropped by AWS CEO Adam Selipsky in passing while others hit the wires ahead of the lead keynote.

Data sovereignty requirements. Before AWS re:Invent kicked off at scale, AWS outlined AWS Control Tower, a set of 65 controls to meet data sovereignty requirements, which requires enterprises to control where data resides and flows.  AWS Control Tower offers a consolidated view of the controls enabled, your cmpliance status, and controls evidence across your multiple accounts. 

Constellation Research analyst Dion Hinchcliffe said:

"This announcement is very significant for large enterprises operating across different countries. Cloud is a real challenge for large multinational organizations. These new controls are vital for them to stay in compliance with data residency and other requirements. The support for multi-account controls is particularly noteworthy."

More from re:Invent:

Zero-ETL integrations. Selipsky quipped that the audience winced in unison when he said ETL (extract, transform and load), but in the future zero-ETL will be a reality. ETL is a major enterprise pain point and while zero-ETL integrations will garner yawns over headlines what AWS announced could be promising for enterprises. I'm betting that there enough enterprise buyers that'll care about zero ETL to throw you a few links about integrations across AWS data stores.

Constellation Research analyst Doug Henschen knows the ETL pain. He said:

"The idea of Zero-ETL is compelling because it promises considerable savings in time, effort and administrative headaches over ETL development work. It promotes low-latency insight while also reducing ETL processing and development costs. The Zero-ETL service will clearly introduce its own costs, but the time and labor savings are compelling. As for the DynamoDB to OpenSearch integration, this will enable data from massive, customer-facing DynamoDB-based transactional deployments to be quickly available to OpenSearch full-text search, fuzzy search, auto-complete, and vector search for machine learning (ML) capabilities. Talking to AWS executives it’s pretty clear a future step might be using the Zero-ETL capability to do reverse ETL from Redshift back into operational databases such as the various flavors of Aurora, RDS and DynamoDB."

    Amazon DataZone. Like ETL, data cataloging isn't a lot of fun either. Anyone in the data trenches knows that it's difficult to provide context around organizational data. The process of data cataloguing matters and anything that cuts down on the labor will be welcomed by enterprises.

    Henschen, who penned a report on the importance of data cataloging, added:

    "The use of ML/AI for augmented cataloging is pervasive among metadata management, cataloging and governance platforms, with examples including Alation, Collibra, Microsoft Purview and Google Dataplex. What's novel here is application of GenAI, which is an obvious next step that multiple vendors are either previewing or adding to their roadmaps. Given that everything is in preview, it's hard to say whether anybody has an edge in using GenAI at this point. DataZone is in the early days of its adoption by AWS customers, so anything it can do to remove friction from using the service will help to promote wider adoption." 

    Amazon Q Code Transformation. This announcement was dropped during the keynote but may have been lost. Amazon Q, a generative AI assistant that runs horizontally across AWS' portfolio, can be used to upgrade Java applications quickly. Amazon Q Code Transformation will analyze existing code, generate a transformation plan and complete tasks. Given how much enterprises need to update and transform code, Amazon Q Code Transformation is worth a look. In a blog post, AWS said:

    "Previously, developers could spend two to three days upgrading each application. Our internal testing shows that the transformation capability can upgrade an application in minutes compared to the days or weeks typically required for manual upgrades, freeing up time to focus on new business requirements. For example, an internal Amazon team of five people successfully upgraded one thousand production applications from Java 8 to 17 in 2 days. It took, on average, 10 minutes to upgrade applications, and the longest one took less than an hour."

    Amazon launches WorkSpaces Thin Client. This announcement received some press play but felt very retro. Thin clients?!? The economics of thin clients have made sense for a while. Adoption has been another story. WorkSpaces Thin Client will cost $195, be centrally managed and give access to Amazon WorkSpaces, WorkSpaces Web or Amazon AppStream 2.0, which provides wider access to applications. Thin clients solve a pain point and it'll be interesting to see if AWS gets traction.

    Amazon EC2 high memory U7i instances for in-memory databases. These instances are in preview and designed to support large, in-memory databases including SAP HANA, Oracle, and SQL Server. Given that many enterprises are moving to SAP HANA, these instances are worth a look.

     

    Data to Decisions Tech Optimization amazon Chief Information Officer

    HPE sees Q4 strength in AI, edge, high performance computing

    HPE sees Q4 strength in AI, edge, high performance computing

    Hewlett Packard Enterprise saw strong intelligent edge and high-performance computing and AI revenue growth  in the fourth quarter, but its legacy compute and storage businesses struggled.

    In the fourth quarter, HPE reported earnings of 49 cents a share on revenue of $7.4 billion, down 7% from a year ago. Non-GAAP earnings for the fourth quarter were 52 cents a share. Wall Street was expecting fourth quarter earnings of 50 cents a share on revenue of $7.55 billion.

    For fiscal 2023, HPE delivered earnings of $1.54 a share on revenue of $29.1 billion, up 2% from a year ago. Non-GAAP earnings of $2.15 a share for fiscal 2023 were at the high range of guidance given at HPE's annual analyst meeting in October.

    HPE CEO Antonio Neri said: "As we continue to capitalize on growing market opportunities – particularly as customer interest in AI continues to explode – I am confident in our ability to deliver substantial returns to our shareholders."

    "CFO Jeremy Cox said HPE was seeing "promising indicators of continued demand in the areas of the market we are prioritizing, especially in AI."

    On a conference call, Neri said:

    "Even against an uncertain macroeconomic backdrop, we saw continued though uneven, demand across our HPE portfolio with a significant acceleration in AI orders. Demand in our AI solutions is exploding. We saw a significant uptick in customer demand in recent quarters for accelerated computing infrastructure and services. In Q4, orders for servers that include accelerated processing units or APUs represented 32% of our total server order mix, up more than 250% from the beginning of fiscal year 2023. APUs, which includes GPU-based servers orders across our business, represented 25% of our total server order mix in fiscal year 2023."

    As for the outlook, HPE said first quarter revenue will be between $6.9 billion and $7.3 billion and reiterated fiscal 2024 sales growth between 2% to 4% in constant currency. First quarter non-GAAP earnings will be in the range of 42 cents a share to 50 cents a share.

    Fiscal 2024 non-GAAP earnings will be between $1.82 a share to $2.02 a share.

    HPE also reiterated annual recurring revenue growth of 35% to 45% from fiscal 2022 to fiscal 2026.

    Data to Decisions Tech Optimization HPE greenlake SaaS PaaS IaaS Cloud Digital Transformation Disruptive Technology Enterprise IT Enterprise Acceleration Enterprise Software Next Gen Apps IoT Blockchain CRM ERP CCaaS UCaaS Collaboration Enterprise Service Chief Information Officer

    Workday Q3 shows strength, raises outlook

    Workday Q3 shows strength, raises outlook

    Workday reported better-than-expected third quarter earnings and raised its outlook for the fiscal year.

    The cloud HR and finance application company reported third quarter earnings of 43 cents a share on revenue of $1.87 billion, up 16.7% from a year ago. Subscription revenue for the quarter was up 18.1% from a year ago. Non-GAAP earnings were $1.53 a share.

    Wall Street was expecting Workday to report third quarter non-GAAP earnings of $1.41 a share on revenue of $1.85 billion.

    Carl Eschenbach, co-CEO of Workday, said the company was seeing momentum from "AI innovation, strength in full platform deals, expanding partner ecosystem, and international growth." Aneel Bhusri, co-CEO of Workday, said the company's strategy to build AI into its core products is resonating with customers.

    Workday's approach to generative AI revolves around building it into its core products instead of going the add-on route that has been popular with vendors. Workday passed the 5,000 core HCM customers in the third quarter. Company executives say Workday is playing the long game and aiming to take advantage as enterprises move to consolidate vendors. 

    On a conference call with analysts, Eschenbach said:

    "Generative AI is becoming a business imperative. As a trusted partner and a market leader with over 65 million users under contract we can uniquely drive efficiencies and improve the employee experience. What we are doing and not just saying is resonating with our customers. Simply put, our value proposition has never been so relevant and powerful."

    As for the outlook, Workday projected fiscal 2024 subscription revenue of $6.598 billion, up 19% from a year ago. Non-GAAP operating margins will come in at 23.8%, which is higher than expectations.

    Bhusri said:

    "We're infusing generative AI into our platform is through our investment in conversational AI. While we are still in the exploratory phase with this technology. We believe conversational AI will fundamentally change how users interact with Workday. By enabling them to easily surface information they need and interact with data through simple conversation. We're also leveraging generative AI to create a conversational experience for Workday Adaptive Planning customers. The use of conversational text will simplify the process of surfacing key planning insights. Enabling users to make quicker, more strategic decisions about their businesses."

    Eschenbach had some interesting comments on how customers are thinking about AI and factoring it into their evaluations. 

    At this point, I don't think people are making decisions yet, just purely on AI. I think it's something that every customer looks at to make sure that they're going to be covered with a new deployment or a customer knowing that Workday has them in a strong place, but they're still looking first and foremost at running their business and moving off of crappy legacy applications into the cloud. And we're unmatched in that category. And then when we add the AI stuff, I think it just checks that AI box.

    But, I would say that despite all the hype, it's still in the early days of actual large scale deployments of AI in HR and finance, we're ready."

    More:

    Future of Work Data to Decisions Tech Optimization Innovation & Product-led Growth Next-Generation Customer Experience Digital Safety, Privacy & Cybersecurity workday AI Analytics Automation CX EX Employee Experience HCM Machine Learning ML SaaS PaaS Cloud Digital Transformation Enterprise Software Enterprise IT Leadership HR GenerativeAI LLMs Agentic AI Disruptive Technology Chief Financial Officer Chief People Officer Chief Information Officer Chief Customer Officer Chief Human Resources Officer Chief Executive Officer Chief Technology Officer Chief AI Officer Chief Data Officer Chief Analytics Officer Chief Information Security Officer Chief Product Officer

    AWS launches Amazon Q, makes its case to be your generative AI stack

    AWS launches Amazon Q, makes its case to be your generative AI stack

    Amazon Web Services made the case at re:Invent that it should be your complete AI stack with Amazon Q, a horizontal generative AI tool that will be embedded throughout AWS and backed up with Amazon Bedrock and infrastructure for model training and inference powered by Trainium and Inferentia processors.

    The catch? AWS' strategy rhymes with Microsoft's copilot everywhere plans as well as Google Cloud's Duet AI plans. The interesting twist for AWS is that it's more horizontal across the platform instead of focused on one application, use case or role. AWS also had a big developer spin on Amazon Q, which is seen as an ally to developers because it can connect code and infrastructure.  

    For enterprises, the big question is whether they will go with one AI stack or multiple. And is that decision made by CXOs or developers? Time will tell, but for now the biggest takeaway from re:Invent is that AWS is taking its entire portfolio generative AI with a bevy of additions that are in preview.

    Adam Selipsky, CEO of AWS, said during his re:Invent keynote AWS has been leveraging AI for decades and now plans to reinvent generative AI. "We're ready to help you reinvent with generative AI," said Selipsky, who said experimentation is turning into real world productivity gains.

    Selipsky said AWS is investing in all layers of the generative AI stack. "It's still early days and customers are finding different models work better than others," said Selipsky. "Things are moving so fast that the ability to adapt is best capability you can have. There will be multiple models and switch between them and combine them. You need a real choice."

    In a veiled reference to OpenAI, Selipsky said "the events of the last few days illustrate why model choice matters."

    Here's a look at the moving parts of AWS' generative AI strategy as outlined by Selipsky and other executives.

    Top layer of AWS' AI stack is Q

    Amazon Q is billed as "a new type of generative AI assistant that's an expert in your business."

    AWS' answer to Microsoft's Azure's copilot everywhere theme is Amazon Q. According to AWS, Q will engage in conversation to solve problems, generate content and then take action. It'll know your code, personalize interactions by role and permissions and be built to be private.

    The vision for Amazon Q is to do everything from answering software development questions to be a resource for HR to monitor and enhance customer experiences. The most interesting theme is that AWS is playing to its strengths--code and infrastructure--and using Q to connect the dots.

    "AWS had built a bevy of cross services capabilities and Amazon Q is the next evolution," said Constellation Research analyst Holger Mueller.

    Mueller said:

    "What sets Q apart is that it has a shot to provide one single assistant across all of AWS services. Q reduces one of the main challenges of AWS--the complexity introduced by the thousands of services offered. But it is only possible because Amazon has been working on the integration layer across its services, starting from access and security over an insights layer, one foundation with SageMaker for all AI and e.g. Data Zone. It can now collect the benefit and has a shot at redefining the future of work. Partnerships with SAP, Salesforce, Workday and more ERP vendors make it easily the Switzerland of GenAI assisted work in the multi-vendor enterprise.”

    Doug Henschen, Constellation Research analyst, crystallized the significance of Amazon Q. He said:

    "Amazon Q was the most broadly compelling and exciting GenAI announcement during Adam Selipsky’s keynote. Amazon Q is an AI assistant that will engage in conversations based on understanding of  company-specific information, code, and technology systems. The promise is personalized interactions based on user-specific roles and permissions. AWS previously offered QuickSight Q for natural language querying within its QuickSight analytics platform, but Amazon Q is single, all-purpose GenAI assistant. The idea is to deliver a AI assistant that will understand the context of data, text, and code. At this point they have some 40 connectors to popular enterprise systems outside of AWS, such as Office 365, Google cloud apps, Dropbox and more. The promise is nuanced, contextual understanding of what users are seeking when they ask questions in natural language."  

    Here's how Amazon Q will be leveraged:

    • Developers and builders will use Amazon Q to architect, troubleshoot and optimize code, develop features and transform code based on 17 years of AWS knowledge.
    • For lines of business Amazon Q is about getting answers to questions with access controls and complete actions.
    • Specialists will get Amazon Q in QuickSight, Connect and Supply Chain.

    Selipsky said AI chat applications are falling short because they're in silos by use cases and applications. "They don't know your business or how you operate securely," said Selipsky. "Amazon Q is designed to work for you at work. We've designed Q to meet your stringent enterprise requirements from day one."

    What's interesting for enterprise buyers is the approach of Amazon Q. Does a horizontal approach to generative AI across multiple functions curb model sprawl?

    Amazon Bedrock: More choices, more relevancy after OpenAI fiasco?

    Amazon added new models to Bedrock, added secure customization and fine tuning of models, agents to complete tasks, automated model evaluation tools as well as knowledge bases.

    Specifically, new models added to Bedrock include:

    • Claude 2.1;
    • Llama 2 Chat with 13 billion and 70 billion parameters;
    • Stable Diffusion XL;
    • Titan models focused on generating images and multimodel embeddings;
    • Command Light & Embed.
    • Fine tuning will be available for Meta Llama 2, Cohere Command and Titan with Claude 2 coming soon.

    Selipsky said enterprises need to orchestrate between models and get foundational models to take actions.

    To support this level, AWS has integrated AI capabilities into its data services such as Redshift and Aurora. Redshift Serverless will get a bevy of AI optimizations. AI will also be used to simplify application management on AWS.

    The breakdown of the data pieces underpinning the broader AI strategy for AWS include:

    • Amazon Aurora Limitless Database, which supports automated horizontal scaling of write capacity. Constellation Research analyst Doug Henschen noted "Aurora Limitless Database is an important step forward in potentially matching rivals such as Oracle on automated scalability."
    • Redshift Serverless AI Optimizations, which brings machine learning scaling and optimization to AWS analytical database for data warehousing. Henschen said "matching rivals by adding sophisticated, ML-based scaling and optimization capabilities with cost guardrails will make Redshift more efficient and performant as well as even more cost competitive." 
    • AWS Introduces Two Important Database Upgrades at Re:Invent 2023

    Trainium, Inferentia with a dash of SageMaker

    AWS launched the latest versions of its Trainium and Inferentia processors, two GPUs that may be able to bring the price of model training down. Today, AI workloads are dominated by Nvidia and AMD is entering the market.

    However, AWS has had GPUs in the market and has been able to acquire workloads for enterprises that may not need Nvidia's horsepower. Here's the breakdown:

    • For training generative AI models, AWS launched Trainium2, which is 4x faster than its previous version and operates at 65 exaflops.
    • For using generative AI models and inference, AWS launched Inferentia2, which has 4x the throughput of the previous version and 10x lower latency.
    • Riding on top of these processors is SageMaker HyperPod, which can reduce the time to train foundation models by up to 40%. AWS said SageMaker HyperPod can distribute model training in parallel with 1000s of accelerators, automatic checkpointing and resiliency.

    AWS also added Inference Optimization to SageMaker and said the new service can reduce foundation model deployment cost by an average of 50% with intelligent routing, scaling policies and better efficiency by deploying multiple models to the same instance.

    Bottom line: AWS made its case to be the generative AI stack for enterprises, but Henschen noted there were a lot of things in preview. One thing is clear: Selipsky's not-so-veiled references to Microsoft Azure and OpenAI illustrate that the narrative gloves are coming off.

    Data to Decisions Innovation & Product-led Growth Future of Work Tech Optimization Next-Generation Customer Experience Digital Safety, Privacy & Cybersecurity amazon AI GenerativeAI ML Machine Learning LLMs Agentic AI Analytics Automation Disruptive Technology Chief Information Officer Chief Executive Officer Chief Technology Officer Chief AI Officer Chief Data Officer Chief Analytics Officer Chief Information Security Officer Chief Product Officer

    AWS presses custom silicon edge with Graviton4, Trainium2 and Inferentia2

    AWS presses custom silicon edge with Graviton4, Trainium2 and Inferentia2

    Amazon Web Services launched Graviton4, its custom chip for multiple workloads, with big improvements over last year's Graviton3. AWS also launched the latest versions of its Trainium and Inferentia processors, two GPUs that may be able to bring the price of model training down.

    The takeaway: AWS plans to push its custom silicon cadence to gain more workloads even as it partners with big guns such as Nvidia, Intel and AMD.

    Graviton4 is billed as AWS' "most powerful and energy efficient chip that we have ever built. The launch of the chip also shows a faster cadence for AWS' processors. Graviton launched in 2018 with Graviton2 following up two years later. Graviton3 launched last year.

    According to AWS, Graviton4 is 30% faster than Graviton3, 30% faster for web applications and 40% faster for database applications. AWS has 150 different Graviton-powered Amazon EC2 instance types globally, has built more than 2 million Graviton processors, and has more than 50,000 customers including Datadog, DirecTV, Discovery, Formula 1 (F1), NextRoll, Nielsen, Pinterest, SAP, Snowflake, Sprinklr, Stripe and Zendesk.

    Hyperscalers are racing to create custom processors that offer an option for enterprises to cut compute costs. While most of the focus is on model training and inferencing, AWS custom processor strategy is wider. In addition to Graviton, AWS launched Trainium and Inferentia aimed at AI workloads.

    Adam Selipsky, CEO of AWS, said during his re:Invent keynote that Graviton is an effort to lower the cost of cloud compute. "We have more than 50,000 customers for Graviton," said Selipsky, who cited SAP as a key customer. "Other cloud providers have not delivered on their first server processors," he said.

    Graviton4 will power R8g instances for EC2 with more instances planned. R8g instances are in preview.

    Juergen Mueller, CTO of SAP, said Graviton-based EC2 instances have provided a 35% bump in price performance for analytical workloads. SAP will be validating Graviton4 performance. 

    In addition, AWS launched the latest versions of its Trainium and Inferentia processors, two GPUs that may be able to bring the price of model training down. Today, AI workloads are dominated by Nvidia and AMD is entering the market.

    However, AWS has had GPUs in the market and has been able to acquire workloads for enterprises that may not need Nvidia's horsepower. Here's the breakdown:

    • For training generative AI models, AWS launched Trainium2, which is 4x faster than its previous version and operates at 65 exaflops.
    • For using generative AI models and inference, AWS launched Inferentia2, which has 4x the throughput of the previous version and 10x lower latency.

    "We need to keep pushing on price/performance on training and inference," said Selipky, who again referenced that other cloud providers were behind on custom silicon. Microsoft announced its AI processors at Ignite 2023

    Constellation Research analyst Holger Mueller said:

    "AWS pushes its custom siliiicon with version 2 on Trainium and Inferentia chips, as well as its fundamental Graviton chip. When a large ISV like SAP moves 4M lines of code and sees cost savings it is something to take note of. Equally the sustainability aspect is key for the next version of custom silicon--it saves real money."

    During the keynote Selipsky was sure to note that Nvidia remains a key partner for AWS, which has a bevy of Nvidia GPU instances. Nvidia CEO Jensen Huang appeared on stage to outline the next phase of the partnership with AWS.

    Selipsky said AWS will add Nvidia DGX Cloud and latest GPUs to its platform. Huang said it will build its largest AI Foundry on AWS. Note that Huang also appeared with Microsoft CEO Satya Nadella to tout Nvidia’s partnership on Azure.

    However, AWS is looking to broaden GPU workloads and the bet is that it'll handle a lot of those on its own silicon.

     

    Data to Decisions Tech Optimization Innovation & Product-led Growth Future of Work Next-Generation Customer Experience Digital Safety, Privacy & Cybersecurity amazon SaaS PaaS IaaS Cloud Digital Transformation Disruptive Technology Enterprise IT Enterprise Acceleration Enterprise Software Next Gen Apps IoT Blockchain CRM ERP CCaaS UCaaS Collaboration Enterprise Service AI GenerativeAI ML Machine Learning LLMs Agentic AI Analytics Automation Chief Information Officer Chief Technology Officer Chief Information Security Officer Chief Data Officer Chief Executive Officer Chief AI Officer Chief Analytics Officer Chief Product Officer

    AWS Introduces Two Important Database Upgrades at Re:Invent 2023

    AWS Introduces Two Important Database Upgrades at Re:Invent 2023

    Monday night, Nov. 27, at Re:Invent 2023, AWS’ Peter DeSantis, SVP of Utility Computing, announced two important database features: Amazon Aurora Limitless Database and Redshift Serverless AI Optimizations. Here's my analysis.



    Amazon Aurora Limitless Database: Announced in private preview, This is an automated sharding feature for Aurora PostgreSQL that will enable customers to horizontally scale database write capacity via sharding. Aurora previously supported automated horizontal scaling of read capacity, but write capacity could only be automatically scaled vertically, by implementing larger and more powerful compute instances via the Aurora Serverless V2 feature. When customers reached the limit of vertical scaling, meaning they’ve already employed that most powerful compute instances, they would have to resort to sharding data across multiple database instances. This has been a common practice for large database deployments, but manual sharding at the application layer introduces complexity and administrative burdens. Aurora Limitless Database does away with these burdens by automating the sharding of data across database instances behind the scenes while ensuring transactional consistency.

    Doug’s take: Aurora Limitless Database will step up competition with the world’s number-one database and Aurora’s biggest competitive target, Oracle Database. AWS is actually playing catch up with this feature, as Oracle introduced automated sharding back in 2017. Nonetheless, given that Aurora is so cost competitive, touted as one tenth the cost of its rivals, Aurora Limitless Database is an important step forward in potentially matching rivals such as Oracle on automated scalability.



    Redshift Serverless AI Optimizations: In another move to match competitors, AWS introduced Amazon Redshift Serverless AI Optimizations. This feature brings ML-based scaling and optimization to AWS’s flagship analytical database for data warehousing. AWS introduced Redshift Serverless in 2021 in order to automate database scaling, but the capability was reactive. Given the time it takes to add new instances up in running, there were sometimes penalties in performance. The AI Optimizations feature introduces a new, machine learning-powered forecasting model that does a better job of forecasting the capacity requirements of existing as well as new and unfamiliar queries. A simple slider controls is said to enable administrators to set the balance between maximizing performance and minimizing cost.

    Doug’s take: ML-based forecasting and optimization is old hat in the world of data warehousing, implemented by the likes of Oracle and Snowflake. Here, too, AWS is playing catchup, but the appeal of Redshift is as a cost-competitive data warehousing option within the AWS ecosystem. Matching rivals by adding sophisticated, ML-based scaling and optimization capabilities with cost guardrails will make Redshift more efficient and performant as well as more cost competitive.

    Related resources:
    Google Sets BigQuery Apart With GenAI, Open Choices, and Cross-Cloud Querying
    Salesforce Data Cloud Emerges as an Obvious Choice for CRM Customers
    How Data Catalogs Will Benefit From and Accelerate Generative AI
     

    Data to Decisions Tech Optimization Big Data ML Machine Learning LLMs Agentic AI Generative AI AI Analytics Automation business Marketing SaaS PaaS IaaS Digital Transformation Disruptive Technology Enterprise IT Enterprise Acceleration Enterprise Software Next Gen Apps IoT Blockchain CRM ERP finance Healthcare Customer Service Content Management Collaboration Chief Information Officer Chief Analytics Officer Chief Data Officer Chief Technology Officer Chief Information Security Officer

    Alianza, AWS team up for cloud communication services

    Alianza, AWS team up for cloud communication services

    Alianza and Amazon Web Services (AWS) signed a multi-year partnership to enable traditional communication service providers to deliver and monetize voice and cloud communications services.

    The deal, outlined at AWS re:Invent, also highlights how AWS partners to take workloads in key markets and verticals. Alianza's Dag Peak, Chief Product Officer said the joint Alianza-AWS combination has been deployed in more than 100 communications service providers (CSPs). These CSPs, which include Lumen, Brightspeed and Viasat, are moving from traditional voice networks to more nimble cloud platforms.

    Alianza, which raised $61 million in financing Oct. 31, offers a cloud communications platform for service providers that replaces legacy systems so providers can provide cloud meetings, collaboration and other digital services.

    According to the companies the combination of Alianza and AWS will include the following:

    • Lower costs and simplified operations by replacing soft switch voice over IP networks and legacy hardware with unified communications as a service platform.
    • Improved customer service via digital automation and control over customer experiences.
    • The ability to launch new services built on Alianza and AWS.
    • A unified view into operations via a software-as-a-service interface.
    • Upcoming tools so CSPs can offer new generative AI services via Amazon Bedrock and other AI services. 

    I caught up with Peak to talk about the CSP market and how it's migrating to the cloud.

    The market. Peak said the traditional telephony market is still large, but often forgotten as vendors have moved up the stack to communication and collaboration apps (think RingCentral, Zoom, Cisco Webex, Microsoft Teams). "In markets where we play well, we don't have many competitors. There's very little interest in smaller service providers," said Peak.

    The need for cloud platforms. CSPs need to move as they transform from phone companies to focusing more on communications. "These CSPs can offer a full stack of services all living in the cloud," said Peak. "AWS is interested in Alianza because our platform is allowing CSPs to offer a full stack of services in the cloud and it can pull in the workloads."

    Transformation. CSPs don't want to rebuild the traditional services, but modernize everything they are doing, said Peak. However, Peak noted that transformation for CSPs starts with traditional voice services with AI because many customers are mobile and don't want to communicate via apps. Alianza's services are delivered via CSPs to customers via an eSIM. "Not everyone is sitting at a computer. Small businesses rely on telephones," said Peak. "We want to enable AI for the hair salons, insurance agencies, and flower shops."

    Use cases. Peak said it's possible to bring tools like sentiment analysis and experience tracking to small businesses. "Contact centers get all the AI love, but we can use AI far down market with AWS integrations," said Peak. The goal is to bring enterprise grade AI services to mainstream businesses without an app.

    Next-Generation Customer Experience Data to Decisions Future of Work Innovation & Product-led Growth New C-Suite Marketing Transformation Digital Safety, Privacy & Cybersecurity Chief Information Officer

    AWS bets palm reading will come to an enterprise near you

    AWS bets palm reading will come to an enterprise near you

    Amazon Web Services launched Amazon One Enterprise, a palm-based identity service that aims to make palm-reading a mainstream way to enter buildings, improve security and verify credentials.

    AWS said Amazon One Enterprise is being used by Boon Edam, IHG Hotels and Resorts, Paznic, and KONE. The service is in preview in the US and pricing wasn't immediately available. Amazon One Enterprise's FAQ is worth checking out for various details on enrollment, security and device setup. 

    The company announced the launch at AWS re:Invent. AWS sees palm reading as a way to better secure and authorize access to physical locations such as data centers, offices, buildings, airports and hotels as well as a way to restrict software and document access.

    More from re:Invent:

    As for the potential returns on investment, the argument for Amazon One Enterprise is straightforward. Enterprises wouldn't have to create and manage badges, fobs and PINs and IT departments could install Amazon One devices. The help desk hours for lost security devices and PINs could justify a look at palm-screening methods.

    Here are the key points about Amazon One Enterprise:

    • It is a fully managed service via the AWS management console and a biometric identification device.
    • Security controls are built in to every stage of the service from the Amazon One device to data in transit and in the cloud. Palm images, metadata and user credentials are immediately encrypted. Each palm has its unique key.
    • AWS said the accuracy rate of the palm and vein imagery is 99.9999% and better than scanning two irises.
    • Amazon One uses AI and machine learning to associate a palm signature with credentials such as badge ID, employee ID or PIN.
    • Authentications, status and software updates and enrollment as well as analytics.
    • Amazon One Enterprise offers two options: A standalone device and a pedestal, where the Amazon One device is mounted on a pedestal.
    Future of Work Next-Generation Customer Experience Digital Safety, Privacy & Cybersecurity amazon Security Zero Trust Chief Information Officer Chief Information Security Officer Chief Privacy Officer

    AWS launches Braket Direct with dedicated quantum computing instances, access to experts

    AWS launches Braket Direct with dedicated quantum computing instances, access to experts

    Amazon Web Services rolled out Braket Direct, a service that allows researchers to procure dedicated private access to quantum processing units from providers such as IonQ, Oxford Quantum Circuits, QuEra, Rigetti, or Amazon Quantum Solutions Lab.

    The effort is part of Amazon Braket, AWS' quantum computing marketplace launched in 2020. Announced at AWS re:Invent, Braket Direct also has experts at the ready to give guidance on workloads and get access to features and devices. These experts offer free office hours and one-on-one reservation prep sessions.

    At a keynote Monday night, Peter DeSantis, Senior Vice President of AWS Utility Computing, said the issue AWS is trying to solve is that qubits are too noisy for workloads. In his keynote coverage, Constellation Research analyst Dion Hinchcliffe said:

    "This tech looks five years out at least, but they are clearly gearing up because the tech will be critical to tackle issues in scientific research, cryptography, pharmacology, and other domains."

    With Braket Direct, customers can reserve an entire quantum machine on IonQ Aria, QuEra Aquila, and Rigetti Aspen-M-3 devices for a period of time. Since the machines are completely dedicated, customers can run complex and time sensitive workloads or use the systems for training.

    Constellation Research analyst Holger Mueller said:

    "AWS kept its quantum plans under tight lid - saying it was only research and now DeSantis shows a brand new quantum chip. AWS is focussing on a new approach to qubit error correction - which will be promising for both for cost and performance of quantum machines." 

    Braket Direct customers can also access systems that have reduced or limited availability. AWS cited IonQ's 30-qubit Forte system as one of those devices. It remains to be seen whether Braket Direct can boost the revenue bases of pure play quantum computing vendors. For instance, Rigetti Computing's revenue for the nine months ended Sept. 30 was $8.63 million. IonQ revenue for the same time frame was $15.94 million. 

    Pricing varies, but an IonQ system will run you $7,000 an hour, Rigetti goes for $3,000 an hour and QuEra is $2,500 an hour. Expert advice offerings are billed separately. Hybrid jobs have a different pricing setup. Here's a look at the Braket Direct QPU pricing.

    Tech Optimization Data to Decisions Innovation & Product-led Growth amazon Quantum Computing Chief Information Officer Chief Technology Officer

    AWS re:Invent 2023: Perspectives for the CIO | Live Blog

    AWS re:Invent 2023: Perspectives for the CIO | Live Blog

    I'm excited to be live from AWS re:Invent 2023 in Las Vegas This year's event is packed with announcements about the leading-edge of cloud computing and the hot topic of the year, generative AI. It's also rife with opportunities for cloud professionals to learn and grow. From a CIO perspective, I'm particularly interested in the keynotes, innovation talks, and builder labs to show where the AWS as a platform is heading for IT leaders. I'm also looking forward to networking with CIOs and cloud experts from around the world to compare notes.

    Jump right to the Live Blog

    As arguably the cloud industy's pre-eminent event, I'm eager to explore the following topics at re:Invent this week, which I believe are the most vital to examine in our enterprise journey through cloud today:

    1. Cloud Economics and Cost Optimization

    With rising economic pressures, CIOs are increasingly focused on optimizing cloud costs without compromising performance or agility. AWS re:Invent 2023 is expected to showcase new "tools and strategies for managing cloud expenditures, including cloud cost management tools, FinOps frameworks, and cost optimization techniques.

    2. Embracing Hybrid, Private, and Multi-Cloud Environments

    Organizations are increasingly adopting hybrid and multi-cloud strategies to leverage the best of each cloud provider and ensure resilience. AWS re:Invent 2023 will explore advancements in hybrid cloud management, including solutions for managing multi-cloud environments, data portability, and application deployment across different cloud platforms including private cloud, the resurgence of which is one of my major research areas currently.

    3. Accelerating Innovation with AI and Machine Learning

    AI and machine learning (ML) are transforming businesses across industries. CIOs are eager to harness the power of these technologies to drive innovation and gain a competitive edge. AWS re:Invent 2023 will delve into the latest AI and ML services, including tools for building AI models, deploying ML applications, and automating IT operations with ML.

    4. Enhancing Cybersecurity and Data Protection

    Cybersecurity threats are becoming more sophisticated, and CIOs must prioritize protecting sensitive data and ensuring compliance. AWS re:Invent 2023 will feature sessions on cloud security best practices, identity and access management, data encryption, and threat detection and response.

    5. Building and Managing Sustainable Cloud Infrastructures

    Sustainability is becoming a critical factor for organizations, and CIOs are seeking ways to reduce their cloud footprint's environmental impact. AWS re:Invent 2023 will highlight cloud-based sustainability solutions, including carbon footprint tracking tools, energy optimization techniques, and green data center initiatives.

    6. Empowering Developers with Cloud-Native Technologies

    Developers are the backbone of cloud innovation, and CIOs must provide them with the tools and resources they need to succeed. AWS re:Invent 2023 will showcase cloud-native technologies, including containerization, serverless computing, and API management solutions, to empower developers to build and deploy applications rapidly and efficiently.

    7. Fostering a Culture of Cloud Agility and Innovation

    Cloud adoption is not just about technology; it's also about fostering a culture of agility and innovation within the organization. AWS re:Invent 2023 will explore strategies for driving organizational change, empowering employees to embrace cloud technologies, and creating a culture of continuous learning and experimentation.

    8. Leveraging Cloud for Industry-Specific Solutions

    Cloud adoption is transforming industries across the board. AWS re:Invent 2023 will feature sessions tailored to specific industries, showcasing how cloud solutions are being used to address unique challenges and opportunities in healthcare, finance, manufacturing, retail, and other sectors. This is a key reason I find that the hyperscalars are key platorms for digital transformation, if they have the blueprints and templates for bringing the cloud directly into how businesses work.

    9. Exploring the Future of Cloud Computing

    As cloud computing continues to evolve, CIOs are looking to the future for insights into emerging trends and technologies. AWS re:Invent 2023 will provide a glimpse into the future of cloud computing, with sessions on quantum computing, edge computing, and the next generation of cloud infrastructure.

    AWS re:Invent 2023 Live Blog

    First up is the Monday evening keynote with Peter DeSantis at 7:30pm PT. Peter is Senior Vice President of AWS Utility Computing. He will discusses how AWS is pushing the envelope of what’s possible. He will describe the engineering that powers AWS services and illustrates how their unique approach and approach to innovation helps create leading-edge solutions across the spectrum of silicon, networking, storage, and compute, with the goal of uncompromising performance and cost.

    https://aws.amazon.com/rds/aurora/

    7:30pm PT: Peter DeSantis comes on stage and begins to make a sophisticated argument for highly scalable cloud computing and serverless.

    7:40pm PT: "You can get serverless scaling of your database storage because your database has access to robust distributed storage service [in tthe cloud] that can scale seamlessly and efficiently as a single table from massive database. And when the database gets smaller, that's taken care of to drop a large index to stop paying for the index. That's how the service works. With our launch of Aurora, we took a big step forward on our journey to making the relational database less, more server less."

    7:45pm PT: After walking through many scenarios of how to scale databases, DeSantis conclucdes that "we're still limited by the size of the physical server. And that's not serverless. However, database sharding is a well known technique for improving databases performance of a single server, given both horizontally partitioning your data into subsets and distributing it to a bunch of physically separate database servers called shards." Clearly, horizontal scaling is the answer, but how do it with the epic scale that today's multilmillion users systems require?

    /system/files/uploads/user-16818/Sharding%20in%20a%20Serverless%20World%20Peter%20DeSantis%20reInvent%202023.png

    7:50pm PT: "We've changed the database to use Wallclock to create a distributed database that is very high performance. And it's made possible by a very novel approach to synchronizing. Syncing the clock sounds like it should be as simple as one server or another server timings. But of course, because the time it takes to send a message from one server to another server and without knowing this propagation time, it's impossible from this box by passing protocols to calcualte by sending round trip messages and subtracting tropical clouds." Peter builds up to a big announcement by tipping the breakthrough research required to achieve it.

    7:55pm PT: Announces Amazon Aurora Limitless Database. "With a limitless database there's no need to worry about providing a new database, your application is just a single endpoint that has to be available." My take: This is a major reduction in complexity that will move the state-of-the-art in cloud data management forward.

    8:30pm PT: Now Peter DeSantis takes us on a fascinating exploration of their efforts in quantum computing, including AWS's work on making it real. The basic issue is that qubits today are far too noisy to tackle serious computation issues. DeSantis takes through the story of logical qubits and the work AWS is doing to make them commercially viable. This tech looks five years out at least, but they are clearly gearing up because the tech will be critical to tackle issues in scientific research, cryptography, pharmacology, and other domains. Overall, an impressive look at where AWS will take enterprises into the frontier of computing, all delivered through public cloud of course.

    Logical Qubits in Quantum Computing AWS reInvent Peter DeSantis 2023

    Adam Selipsky Keynote | November 28th, 2023 at 8:00am

    8:00am PT: Starting right on time, Adam comes out and makes the case that Amazon is the leading cloud platform in the world. "We are relentless about working backwards from our customers needs and from their pain points. And of course, we were the first by about five to seven years to have this broadest and deepest set of capabilities we still have and we're the most secure, the most reliable."

    8:10am PT: Continuing the trend to showing love to the core AWS platform, Selipsky talks about S3, their original cloud storage service. "So, 17 years ago, AWS reinvented storage by launching the first cloud service for AWS with S3. As you can see as we continue to how development use store, same simple interface, low cost, and high performance. It's another example of general purpose computing. We realized almost 10 years ago, that we wanted to continue to push the envelope on price performance for all of your workloads, which are reinfect general purpose computing, for the cloud era, all the way down."

    8:12am PT: The first big announcement of Day of re:Invent: "I'm excited to announce Amazon s3 Express was a new s3 storage class, purpose built. Purpose Built on a high performance and lowest latency Cloud Object Storage for your most frequently accessed data. Express One uses purpose built hardware and software to accelerate data processing also gives you the option to actually choose your performance for the first time. It can bring frequently accessed data next to your high performance compute resources to minimize latency." Express One supports millions of requests per minute with a claimed single digit millisecond latency with very fast object storage in the cloud. Selipsky says it is up to 10 times faster than S3 standard storage. While cloud costs continue to remain paramount with many IT departments, AWS is still focusing on overall performance as well, key to landing the largest online services.

    Amazon S3 Express One Storage

    8:15pm PT: Now Selipsky moves onto server compute, the workhorse of cloud platforms. Notes that 50,000 customers currently use their custom-designed Graviton chips today. "For example, SAP partners with AWS to power SAP HANA Cloud with the Graviton chip. SAP is seeing up to 35% better price performance for analytics workloads as it aims to reduce carbon footprints or carbon impact by an estimated 45%." This confirms the performance gains/cost savings of using chips optimized for cloud compute workloads.

    8:17pm PT: Announces the Graviton4, "the most powerful and the most energy efficient chip that we have ever built. With 50% more cores and 75% more memory bandwidth the Graviton, with 30% faster on average than the Graviton3, performs even better for certain workloads, like 40% faster for database applications." Notes that AWS is on its fourth generation of cloud servier processors, and says their competitions is often not even on their first. There is a new Graviton4 preview that customers can join if they like.

    Graviton4 server processor at re:Invent

    8:20am PT: Now Selipsky gets into how AWS things about artificial intelligence. Specifically, they see it having three layers

    Layer 1 - Infrastructure for training/inference
    Layer 2 - Tools to build with LLMs
    Layer 3 - Applications that leverage LLMs

    8:22am PT: Next up is GPUs, the workhorse of AI. Selipsky notes that "having the best chips is great and necessary. But to deliver the next level of performance you need more than just the best GPU. You also need really high performance clusters of servers that are running these that can be deployed and easy to use ultra clusters." Next he talks about AWS differentiation with GPUs, namely: "All of our GPU instances that have been released in the last six years are based on our breakthrough Nitro system, which reinvented virtualization by offloading, storage and networking specialized chip to all the service compute dedicated for running your workloads. This is not something that other cloud providers can offer." My take: This lightweight hypervisor is key to AWS's cost/performance advantage for many advanced workloads, and scale them up to 20,000 GPUs at once, AWS claims.

    8:23am PT: Brings Jensen Huang up on stage, co-founder and CEO of NVIDIA, who is certainly making the rounds at all the cloud events as the generative AI partner darling of the year, given their pre-eminence in the GPU industry, instrumental to training and running AI models. Their DGX platform has been a key differentiaor for them and has captured the leading marketshare in the industry.

    Jensen Huang CEO NVIDIA at re:Invent 2023

    8:28am PT: Huang makes a lot of interesting statements, mostly about scale, which is the signature challenge of generative AI. "We are incredibly excited to build the largest AI factory NVIDIA has ever built. We're going to announce inside our company we call it Project Siba(?). Siba, as you all probably know, is the largest, most magnificent tree in the Amazon. We call it project Siva. Siva is going to be 16,384 cores connected into one giant AI supercomputer. This is utterly incredible. We will be able to reduce the training time of the largest language models the next generation mo ease these large, extremely large mixture mixture of experts models, and be able to train it in just half the time. essentially reducing the cost of training in just one year and how now we're going to be able to train much much larger multimodal M models this next generation large language models." Another proof point that NVIDIA is at the forefront of helping build the largest AI models ever created.

    8:35am PT: Now the new Tranium2 chip is announced. Designed to quickly and inexpensive training AI models. "This second generation system is purpose-built for high performance training. It was designed to deliver four times faster performance compared to our first generation chips that makes it ideal for training front foundation models with multiple hundreds of billions or even trillions of parameters." This will very much help sustain AWS's claims that their Tranium chips offer unique performance and cost effective aids for organizations making major forays into generative AI. 

    Tranium 2 chips at re:Invent 2023

    8:37am PT: Now Selipsky talks about the software frameworks on top of their custom ML and AI chips.  "AWS Neuron is our software development kit that helps customers get maximum performance from our ML chips. Neuron supports our machine learning framework frameworks like TensorFlow so customers can use their existing knowledge both training and inference pipelines is just a few lines of code. Just as importantly, you also need the right tools to help train and deploy your models. And this is why we have Amazon Sagemaker. Our managed service makes it easy for developers to train to build machine learning and foundation balls. In the six years since the first Sagemaker, we've introduced many powerful innovations like automatic public tuning, a distributed training, flexible model deployment tools for ml ops and built in features like responsible AI." This vertical integration of hardware and software is a very compelling offering in the ML and AI spaces, and certainly we've seen Sagemaker climb the usage charts around the world since its release. Selipsky continually beats the drum today on public cloud being the most compelling price/performance for training. My verdict: It depends on the workload and how often is must be trained especially.

    The Price Performance Value Prop of AWS at reInvent 2023

    8:40am PT: Selipsky reaffirms AWS's stance on full model choice: "So you don't want a cloud provider who's beholden primarily to one model. You need to be trying out different models. You need to be able to switch between them randomly even combining them with the same use case. And you need a real choice model as you decide who's got the best technology, but also who has dependability that you need a business partner. I think the events of the past 10 days. We've been consistent about lead for choice for the whole history of AWS." So, after taking a swing at the turbulence at OpenAI, Selipsky then shows that they do have preference, just that they're open. Because of who is up next...

    8:42am PT: Now the CEO of Anthropic, Dario Amodei, is invited up on stage with Selipsky. They talk about Claude 2. Talks about their desire to be the leader in safe, reliable, steerable Generative AI. Claude 2 can handle 200K tokens and is "10 times" more resistant to halllucinations that other opic models. My take: It looks like Anthropic is the AI model that's first among equals in AWS AI model choice.

    AWS and Anthropic at re:Invent 2023

    9:03am: After some customers stories from Pfizer and others, Selipsky moves onto the vital enterprise topic of Responsible AI, something that AWS recenly reaffirmed its corporate commitment to. "We need generative AI to be deployed in a safe, trustworthy and responsible fashion.  That's because the capabilities to make generative AI such a promising tool for innovation also have the potential to deceive. We must find ways to unlock generaive AI's full potential while mitigating the risks. Dealing with this challenge is going to require unprecedented collaboration through a multi-stakeholder effort across technology companies, policymakers, community groups, scientific communities, and academics.

    We've been actively participating in a lot of the groups have come together to discuss these issues. We've also made a voluntary commitment to promoting safe, secure and transparent AI development technology [to build] applications that are safe, but avoid harmful outputs, and that stay within your company. And the easiest way to do this is actually placing limits on what information we can or can't return. We've been working very hard here and so today we're announcing Guardrails for Amazon bedrock." 

    This is maybe the most impactful announcement regarding AI at re:Invent so far. Making enterprise safe, transparent, and risk managed is one of the highest priorities for organizations as they develop generative AI policies -- see my roadmap for AI at work here -- and build applications for it.

    Guardrails for Amazon Bedrock at reInvent 2023

    9:16am PT: Now Selipsky talks about cloud talent, is vital subject and urgent situation holding back innovation and digitat transformation in many organizaitons today. He notes that that AWS plans to "provide the cloud skills that we are going to be needed across the world for years to come. AWS has committed to training 29 million people for free with cloud computing skills by the year 2025. We're well on our way to 21 million already."

    9:18am PT: Selipsky moves the conversation onstage to "generative AI chat applications. So these days, what the early [generative AI] providers in the space have done is really exciting, and it's genuinely super useful for consumers. But a lot of ways these applications don't really work at work. With their general knowledge and their capabilities are great But they don't know your company. They don't know your data or your customers or your operations. And this limits how useful their suggestions can be.

    They also don't know much about who you are at work. They don't know your all your preferences. What information you use, what you do and don't have access to. So critically other providers from the launch tools, they launched out their privacy and security capabilities that virtually every enterprise requires. So many CIOs actually banned the use of a lot of these [Dion: 22% of orgs in my most recent CIO survey] most popular AI systems inside their organization that has been well publicized. Just ask any Chief Information Security Officer CISO. They'll tell you the full time security fact and expect it to work as well much much better to build security, fundamental design technology."

    The Challenges of Chat AI Apps at Work - re:Invent 2023

    9:20am PT: In what is probably the biggest enerprise AI announcement at re:Invent, Selipsky tips Amazon Q, a new enterprise-grade generative AI chat system, akin to ChatGPT, but designed businesses. Amazon Q is especially designed for enterprises to "understand your systems, your data repositories, your operations. And of course we know how important rock solid security privacy are that you understand respect your existing identities for roles in your permissions that the user does not have permission to access something without you they cannot access it with you either. We've designed to meet enterprise requirements. Enterprise customers have stringent requirements from day one." Says they will never use customer data in their models, ever, which will be absolutely key. Amazon Q is in preview today (see link above.)

    The pricing page for Amazon Q puts a premium investment level on AWS's new business chat app service:

    - $20/month per user for Q Business
    - $25/month per user for Q Builder

    My take: Given that finding information to carry out knowledge work is still one of the biggest unmet needs in the digital workplace, if Amazon Q can deliver the goods, it has the potential to be worth the price.

    Amazon Q - AWS's Chat AI for Enterprises at re:Inventt 2023

    9:25am PT: Amazon Q is also an expert on the AWS platform, and can actively help developers and operations staff (DevOps teams) get far more from he platform. "Amazon Q is your expert assistant for building on AWS how to supercharge developers and IT pros. We've trained as on two and a half years worth of AWS knowledge. So I can transform the way you think, optimize and operate application workloads on AWS. And we put Amazon Q where you work, so it's ready to assist you in the AWS Management Console as your code whisperer and in your team chat apps like Slack, Amazon Q is an expert in AWS tech pattern and practices and solution implementations."

    Swami Sivasubramanian Keynote | November 29th, 2023 at 8:30am

    8:30am PT: Sivasubramanian comes out and talks about Ada Lovelace, and her conjecure that computers could only do what they're programmed to do, not come up with ideas that are enirely new. Suggests the AI era will change that. Then dives right into a discussion of the generative AI stack as AWS sees it:

    The Generative AI Stack reInvent 2023

    8:41pm PT: Now Sivasubramanian gets to the most important topic perhaps of all: Model choice. He underscores the AWS position: "No one model will rule them all." Large language models and foundation models will come in many flavors for many needs. Indeed, our estimate is that most orgs will soon have dozens or even hundreds of them.

    Model Choice AWS reInventt 2023

    8:43am PT: Next, Sivasubramanian explores the specific foundation models (FMs) and large language models (LLMs) that Amazon Bedrock supports. Bedrock stands at the core of AWS's generative AI services in the cloud. AWS describes Bedrock as a "managed service that offers a choice of high-performing foundation models (FMs) via a single API." It supports FMs and LLMs from leading AI companies including AI21 Labs (Jurassic), Anthropic (Claude), Cohere (Command + Embed), Meta (Llama 2), Stability AI (Stable Diffusion), and Amazon (Titan).

    My take: Model choice is one of the most important dimensions of generative AI to properly realize and enable. Cloud vendors are now in a race to provision and make safe as many FMs and LLMs as possible, with Google at the head of the pack, AWS catching up, and Microsoft betting heavily mostly on OpenAI. Choice is likely to win the day and will be a critical dimension to track over the next couple of years.

    Model Choice in Bedrock reInvent 2023

    8:51am PT: Interestingly -- and continuing a trend throughout re:Invent -- gives special attention to Anthropic's Claude 2.1 LLM, citing key advantages like a comparatively large 200K token context window, 2x less model hallucination, and 25% lower cost of prompts/completions.

    Anthropic Claude 2.1 at AWS re:Invent 2023

    8:53am PT: Multimodal applications that use different types of daa are complex and hard to build says Sivasubramanian, correctly. "Developers need to spend time piecing together multiple models. Not only does this increase the complexity of your daily tasks, but it also decreases the efficiency and impacts customer experience. We wanted to make these applications even easier."

    Multimodal AI Apps at AWS reInvent 2023

    8:54am PT: "That's why today I'm excited to announce the general availability of Titan Multimodal Embeddings. This model enables you to create richer, multimodal search and recommendation options, we can quickly generate store and retrieve embeddings more accurately and contextually relevant, in one type of search. Companies like Opera are using pattern multimodal embeddings as well as inline which is using this morning to revolutionize this document search experience for their customers." A key feature is that it can adapt to unique and proprietary business data. And it has built-in bias reduction.

    Amazon Titan Multimodal Embeddings AWS reInvent

    8:55am PT: Continuing a focus on price/performance in AI -- something key as FMs and LLMs can be very costly to run in daily opeations -- Sivasubramanian announces Titan Text Lite and Titan Text Express (see video demo here). "These next models help you optimize for accuracy, performance and cost depending on your use cases. These are really, really small models that are extremely cost effective model that supports use cases like text summarization. It is ideal for fine tuning, offering your highly customizable model for your use case. Express can be used for a wide range of tasks such as open ended text generation and conversational chat. These two models provides a sweet spot for cost and performance compared to other really big [foundation/large language] models."

    Titan Text Lite and Titan Text Express at AWS reInvent 2023

    8:56am PM: The new Amazon Titan Image Generator service is announced, now available in preview.

    • Studio-quality images from natural language prompts
    • Customize with enterprise data/brand​​​​​​
    • High alignment of text to image

    Amazon Titan Image Generator AWS re:Invent

    9:05am PT: After a deep-dive into how Intuit is using AWS's AI platforms, Sivasubramanian shifts the conversation to the absolute centrality of data to generative AI. The key here is that AWS is prepared to indemnify orgs against what's generated on their platform, including saying it will guard them against legal and reputational damage. They also promise enterprise data will never flow back into their models, but don't offer the same sort of indemnification. 

    Data as the Differentiator in Generative AI at AWS reInvent

    9:07am PT: So, how can organizations "enable your model to understand your business over time." Sivasubramanian explains how fine-tuning is the process for customizing models with enterprise data. Titan supports both fine-tuning with labelled and unlabelled raw data, and is key to getting AI adapted to a business. Sivasubramanian re-emphasizes that fine-tuning always stays with the customer, and never flows back into AWS's base AI models.

    Fine Tuning AI Models with Enterprise Data AWS reInvent

    9:09am PT: There are two major ways to customize AWS's Titan model to adapt to a given business:

    1. Small, labelled data -> Fine-tuning
    2. Large amount of unlabeled data -> Continued pre-training

    These two approaches are key to making generative AI adapt to a business.

    Sivasubramanian gives an example: "A healthcare company can pre-train the model that using medical journals, articles or research papers to make it more knowledgeable on the evolving industries are not. And you can leverage both of these techniques, to act on Titan Lite and Titan Express. These approaches complement each other and will enable your model to understand your business over time. But no matter which method you use, the output model is accessible only to you and it never goes back to the base model."

    Fine Tuning and Continued Pre-Training of AI Models at AWS reInvent

    9:12am PT: Then Sivasubramanian takes us through an an example of how a generative AI model in Bedrock can be fine-tuned w/itha copy of Meta's Llama2 LLM. Then labelled enterprise data is added in via the S3 storage service. The resulting fine-tuned model is used to generate customized, business-relevant content as needed. This is one way to adapt a model in the AWS platform to a business. But it's not the only way, and may not be suitable for large amounts of data or when domain accuracy is strictly needed.

    Retrieval Augmented Generation at AWS reInvent

    9:15am PT: But there are other ways to integrate enterprise data with an LLM, even than these two more strategic options. There is also an approach known as retrieval augmented generatation (RAG), which Sivasubramanian walks through as well, and is what many people already do by hand with LLMs today. Understanding these options in terms of pros/cons so that enterprises can evaluate how best to mix in their valuable data with an AI model's data. In the RAG approach "the prompt that is sent to your foundation model contains contextual information such as product details, which it draws from your private data sources. These are in the context within the prompt text itself. Hence the model provides more accurate and relevant response due to the use of specific prompt-supplied business details." This is a 3rd way, beyond fine-tuning and continued pre-training, to get an AI model to generate using business-specific private data.

    Retrieval Augmented Generation at AWS reInvent

    9:18am PT: The next announcement that Sivasubramanian makes is Knowledge Bases for Amazon Bedrock, which is an off-the-shelf way to implement the RAG workflow described above to give foundation models contextual information from an organization's private data sources. Specifically Knowledge Bases performs the following in a proven workflow:

    • Converts text docs into embeddings (vector representations)
    • Stores them in a vector database
    • Retrieves them + augment prompts Avoids custom integrations.

    Put simply, the service can look at Amazon S3 and then it automatically fetches the documents, divides them into blocks of text, converts the text into embeddings, and stores the embeddings in a vector database. Then this information can be retrireved and used in a variety of generative activities. Interestingly and very usefully, the service can provide source attribution if needed as well.

    Knowledge Bases for Amazon Bedrock at AWS reInvent

    Related Research

    My AWS re:Invent 2023 "mega thread" with all the ongoing details this week

    A Roadmap to Generative AI at Work

    Last year's AWS reInvent Live Blog for 2022 

    My current Digital Transformation Target Platforms ShortList

    Private Cloud a Compelling Option for CIOs: Insights from New Research

    The Future of Money: Digital Assets in the Cloud for Public Sector CIOs

    New C-Suite Tech Optimization Data to Decisions Innovation & Product-led Growth Future of Work Next-Generation Customer Experience Digital Safety, Privacy & Cybersecurity AWS reInvent aws amazon ML Machine Learning LLMs Agentic AI Generative AI Robotics AI Analytics Automation Quantum Computing Cloud Digital Transformation Disruptive Technology Enterprise IT Enterprise Acceleration Enterprise Software Next Gen Apps IoT Blockchain Leadership VR SaaS PaaS IaaS CRM ERP CCaaS UCaaS Collaboration Enterprise Service business Marketing finance Healthcare Customer Service Content Management Chief Analytics Officer Chief Data Officer Chief Digital Officer Chief Financial Officer Chief Information Officer Chief Information Security Officer Chief Procurement Officer Chief Supply Chain Officer Chief Sustainability Officer Chief Technology Officer Chief Executive Officer Chief AI Officer Chief Product Officer Chief Operating Officer