"We really feel Responsible AI is about explainability. It's also about how we are responsible to our customers and to our community." -Nidhi Sinha, General Manager, Analytics, @CommBank #H2OWorldIndia #CustomerObsessed #AI4Good
SAFETY
-
Crunchbase LLM Hallucination Issues and Data Accuracy
By
–
Maybe crunchbase's LLM hallucinated that info…
-
AI image creator rejects photo award over authenticity concerns
By
–
The camera never lies? Creator of #AI image rejects prestigious photo award:
-
AI Model Limitations with BPD Topic Mentions
By
–
Ah, so if you explicitly mention BPD it says something like “as an AI model, I can’t answer that”?
-

Hierarchical Partitioning for Differentially Private Heatmap Computation
By
–
Today on the blog, read all about how we use a hierarchical partitioning procedure to build an efficient differentially private algorithm for computing heatmaps with provable guarantees #DifferentialPrivacy → https://
goo.gle/3MV5zyT -
Discussion on Prompt Engineering, RLHF, AI Safety, and AGI
By
–
Got to chat about prompt engineering, RLHF, LLM red teaming, AI safety, and AGI with @labenz on the @CogRev_Podcast — thanks so much for having me, Nathan!
-
Test Benchmark Rerun and Jailbreak Submissions Review
By
–
reran the test benchmark after a lot of complaints it wasn't working as well anymore & those other ideas are good and I have thought abt them somewhat in regard to the pending jailbreaks, there are a LOT of junk jailbreaks that are submitted to me in regard to the comments, I
-
GPT-4 Intelligence Assessment Requires New Evaluation Methods
By
–
GPT-4 often "seems" remarkably clever, and some believe it exhibits features of more general intelligence. Experts say we need new methods for probing the model's intelligence, and more transparency from @OpenAI about how it works to truly understand it.
-
Claude upgrades maintain capabilities while improving safety standards
By
–
For businesses using Claude, capabilities in all domains should stay the same or improve as you upgrade from previous versions. We always work to improve safety and performance in tandem.
-
Claude-v1.3 Safer Model Released With Improved Adversarial Robustness
By
–
We are offering a new version of our model, Claude-v1.3, that is safer and less susceptible to adversarial attacks.
