The infusion of #AI into enterprise applications is creating a need for the continuous delivery and automation of AI workloads. Register for these #MLOps sessions at #GTC24 to learn best practices for generating enterprise business value with AI. https://
nvda.ws/3SFuGXQ
SOFTWARE
-

MLOps Best Practices for Enterprise AI Workloads
By
–
-

IBM WatsonX Foundation Models LangChain Integration for Business
By
–
@IBM WatsonX Foundation Models for Business Dive into IBM WatsonX's powerful foundation models with our LangChain integration! Leverage WatsonX's AI and data platform, built specifically for business applications. Begin your journey today with the `langchain-ibm` python
-

SnapLogic Launches GenAI Builder for Enterprise Workflows
By
–
Watch now on-demand the recording of the @SnapLogic GenAI Builder launch event: https://
snaplogic.com/resources/webc
asts/introducing-genai-builder?utm_source=TW&utm_medium=SOC&utm_campaign=2024_0117_ONL_CORP_GenAIBuilder-Webinar
…
…Other panelists and I discuss the impacts of #GenerativeAI and #LLMs on workflows, processes, products & services for multiple enterprise use cases and lines of business. -
Faster Model Serving Supercharges AI Reflection Capabilities
By
–
Nice! Reflection is going to be super charged when models are served much faster / groqed
-

Cascade Speculative Drafting Accelerates LLM Inference Speed
By
–
Cascade Speculative Drafting for Even Faster LLM Inference Chen et al.: https://
arxiv.org/abs/2312.11462 #ArtificialIntelligence #DeepLearning #MachineLearning -

Google Gemma 2B 7B Models Optimized NVIDIA TensorRT-LLM
By
–
Google’s newly announced Gemma 2B and 7B models, optimized with NVIDIA TensorRT-LLM – allows developers the ability to optimize inference performance across NVIDIA AI platforms, from the datacenter to local PCs with RTX GPUs: https://
nvda.ws/48psSb5 -

New HF Space for visualizing chunk splitting methods in RAG
By
–
I've built a new HF Space to let you visualize how different splitting methods affect the chunks you fet for RAG! Try it out here: https://
huggingface.co/spaces/m-ric/c
hunk_visualizer
… It's heavily inspired from @GregKamradt 's http://
chunkviz.com – all credits to him for the idea! -
AI Tools Need Better Mobile UI for Real Usage
By
–
Agreed. I’ve used both and they are the right type of answer. I just am waiting for a better UI and real usage on mobile.
-

OpenAI Doubles GPT-4 Turbo Rate Limits to 1.5M Tokens
By
–
OpenAI also doubled rate limits for GPT-4 Turbo now reaching a maximum of 1.5M tokens per minute while also removing daily limits. Good news for developers. But when Sora @openai
?? -
Nous-Hermes-2 Quantized to 4-bit for MLX Apple Silicon
By
–
I just quantized this amazing model to 4-bit, with support for the MLX platform so you can run it super fast on Apple Silicon https://
huggingface.co/mlx-community/
Nous-Hermes-2-Mistral-7B-DPO-4bit-MLX
…