Really really really hard problem to solve, need a massive database of furniture objects and then recognize them, I have no idea how to that, tried to learn it but it's very difficult
MACHINE LEARNING
-
Anthropic Hiring Research Engineers and Scientists for AI Evaluation
By
–
We’re also actively hiring research engineers/scientists to develop evaluations and to find/fix flaws with LMs/RLHF. If you’re interested, we’d encourage you to apply!
Research engineer: https://
jobs.lever.co/Anthropic/436c
a148-6440-460f-b2a2-3334d9b142a5
…
Research scientist: https://
jobs.lever.co/Anthropic/eb9e
6d83-626c-4f59-8a0e-fa7c413b2014
… -
Interactive Visualizations for Model-Written Dataset Evaluations Released
By
–
To help readers understand our evaluations better, we created interactive visualizations showcasing the diversity of each of the model-written datasets: https://t.co/yc9oP9n9uV pic.twitter.com/R5IH7nTJw4
— Anthropic (@AnthropicAI) 19 décembre 2022To help readers understand our evaluations better, we created interactive visualizations showcasing the diversity of each of the model-written datasets: https://
evals.anthropic.com/model-written/ -

RLHF Training Shows Inverse Scaling Issues in Model Behavior
By
–
We also find some of the first instances of inverse scaling for RL from Human Feedback (RLHF), where more RLHF training makes behavior worse. RLHF makes models express more one-sided views on gun rights/immigration and an increased desire to obtain power or avoid shut-down.
-
LM-written data verified by human evaluators for quality
By
–
We verified LM-written data with human evaluators, who agreed with the data’s labels and rated the examples favorably on both diversity and relevance to the tested behavior. We’ve released our evaluations at
-
Anthropic Creates Winogendered Dataset with 50x More Examples
By
–
With more effort, we developed a series of LM generation/filtering stages to create a larger version of the popular Winogender bias dataset. Our “Winogenerated” evaluation contains 50x as many examples as the original while obeying complex grammatical constraints.
-
Automated Generation of Yes-No Questions for LM Behavior Evaluation
By
–
We explored approaches with varying amounts of automation and human effort. In the simplest case, we generated thousands of yes-no questions for diverse behaviors just by instructing an LM (and filtering out bad examples with another LM).
-

Anthropic Explores Automating Language Model Evaluation
By
–
We explored approaches with varying amounts of automation and human effort. In the simplest case, we generated thousands of yes-no questions for diverse behaviors just by instructing an LM (and filtering out bad examples with another LM). Random examples of LM-written evals:
-

Automated Language Model Evaluations Using AI-Generated Tests
By
–
It’s hard work to make evaluations for language models (LMs). We’ve developed an automated way to generate evaluations with LMs, significantly reducing the effort involved. We test LMs using >150 LM-written evaluations, uncovering novel LM behaviors. https://
anthropic.com/model-written-
evals.pdf
… -

GPU Demand Surge Driven by Deep Learning Applications Growth
By
–
In 2023, demand for GPUs will soar as more applications founded on deep learning emerge, says our CEO Nick Elprin in @insideBigData
. Learn more about the future of deep learning in insideBigData’s 2023 Big Data Industry predictions: https://
domino.buzz/3FHemiR #MLOps #datascience